When I first heard that statistics had a long-running dispute between “Bayesians” and “frequentists”, I was confused: this is math; how can there be long-running disputes? In math, things are proven; the losing side loses definitively and moves on. So what gives?

It turns out the dispute is about the social role of the statistician, not the math itself. In a nutshell, if you’re the person in charge, you take a Bayesian approach, while if you’re a subordinate with a mere advisory role, frequentist is more appropriate. As an example, a frequentist might say (or might have said in the days of feudalism):

“Prithee, my lord, please kindly deign to notice that the last 37 people who tried your idea died horribly.”

whereas a Bayesian might say:

“Besides being obviously idiotic, this idea of yours killed the last 37 people who tried it. I don’t have time for any more of this shit; you’re fired.”

To explain in detail:

A common misconception about statistics relates to P values. In medicine and biology, a P value less than or equal to 0.05 is generally considered “statistically significant”, while P values greater than that are considered “statistically insignificant”. The misconception is to say that a P value of 0.05 means that “there is only a 1 in 20 chance that this result is just random noise”, 0.05 being equal to 1/20.

(There’s no magic about the number 0.05; it’s just the convention that the business has standardized on, as a compromise between the risk of ignoring real results and the risk of wrongly accepting false results. Alternative thresholds might well be better, but it’s not a ridiculous convention, and with any alternative threshold there’d also be the same misconception, just with another number rather than 20.)

I’ve seen this misconception even from some very sophisticated people, but it’s not what P values actually mean. The actual meaning is: if this were just random noise, then there’d only be a 1 in 20 chance of it happening. (Or to be really precise, a 1 in 20 chance of it or a result even more extreme happening.)

This is a much less useful and more complicated meaning. We use it because it’s what we can calculate easily and objectively – we know the math of random noise – but you have to know the context to interpret it. If, for instance, this is a “hot” research topic, and many researchers have jumped in, there will be a lot of achievements of P<0.05 even if the whole field is just measuring random noise. Monkeys typing randomly at keyboards cannot actually produce the works of Shakespeare in any realistic sense, but ten researchers measuring noise can easily produce a result with P=0.05. Of course 19 of 20 of their experiments (on average) will produce null results, but there’s a fair chance that they won’t even try to publish those, perhaps even thinking that they must have messed up their experiments rather than the new hot theory being wrong. (There are always lots of ways to mess up experiments, especially when they’re cutting-edge experiments at the frontiers of science.)

So interpreting scientific results is hard; it would be much nicer if that misconception were actually true. Bridging the gap between the misconception and the reality is what Bayes’ Theorem is for.

Frequentists don’t deny the truth of Bayes’ theorem, nor could they, since it’s exceedingly simple and easy to prove. They’d just prefer not to use it. And with good reason, since using it demands that you come up with a “prior probability”. The theorem is written as:

\[ P(A|B) = \frac{P(B|A)P(A)}{P(B)}. \]

The thing we want to know, and that the misconception says we really know, is the quantity on the left, P(A|B), “the probability of A given B”, where here we’re taking A to be “the theory is correct” and B to be “the experimental results came out the way they did.” In other words, assuming the experimental results are correct – they might not be; there might be some monkey business going on; but assuming that isn’t the case – what is the chance of the theory being correct? That’s the meaning of P(A|B). Or, well, to be precise, not “the theory is correct” – that’s way too strong a statement – but rather “the results agree with this theory”; there are always other theories that would predict the same thing in this case but different things in other cases. And usually in biology / medicine, the theory in question is the null hypothesis (“nothing here but random noise”), and we’re trying to disprove it – to make the rather modest statement that “there’s something here even if we can’t say quite what.” (The rest of the paper should address what that something might be.)

Bayes’ Theorem tells us that there are three quantities which go into computing the number that we really want to know. The first is P(B|A), which is the aforementioned “P value”. We know the math of random noise; with a few relatively modest assumptions to precisely define which sort of random noise we’re dealing with, we can easily compute that value. (Those assumptions are of the “never exactly correct but generally close enough” sort.)

The second number is P(B), the probability of B, irrespective of whether A is true. We can get it from background statistics of that particular outcome, or from a control group. So it’s also not a big problem.

But unfortunately the third quantity, P(A), is a total wild card. In Bayesian lingo, it’s the “prior probability” of our hypothesis – what we considered its chance to be, before we did the experiment. Bayes’ theorem is not just giving us an opportunity to sneak in our prejudices; it’s demanding them from us.

To illustrate, suppose someone does an experiment showing that a law of physics is wrong – and not some recent, speculative law of physics, but one of those basic laws about which Rutherford is said to have stated that you should never bet against them at odds of less than 1012 to one (a quadrillion to one). He might not actually have said that, but regardless, suppose you agree with that quote and 10-12 is your prior probability of this law of physics being wrong. Then you look into the details of the experiment, and if it were mere random noise there’d be a one percent chance of the results coming out that way. Then your posterior probability is 10-10, even if you totally believe the experiment was done correctly and the results reported accurately. (The number is not quite that, since I’m ignoring P(B), but that’s the general drift.)

In this case the difference between the misconceived and the real meaning of the P value is particularly stark: rather than believing there’s now only a 1% chance the law of physics is correct, you still believe there’s still something like a 99.99999999% percent chance.

Islamic extremists are also Bayesian in their attitudes: their prior probability that Islam is incorrect is zero, and thus any evidence that it might be incorrect gets multiplied by zero in their heads. However many times they run into such evidence, zero multiplied by anything is still zero.

The now-deceased owner of the yacht Bayesian was also presumably a Bayesian, as is appropriate for someone like him who was the head of a financial company. Michael Lewis, in Liar’s Poker, describes Wall Street traders calling themselves “Big Swinging Dicks”; in accordance with that theme, the world-record 72-meter mast on the boat was quite a large erection, and in a sudden squall it exerted so much leverage that it swung down to the water, capsizing the boat and killing most on board, including its owner. But while we may laugh at such people’s excesses, still, as the ones who call the shots, they should be Bayesians and use all their prior knowledge, not be shy about it.

Their subordinates, though, responsible for feeding them facts they rely on, should take a different tack. Their prejudices are not supposed to override their boss’s. They may be in training for his position, but they’re not there yet. (Well, one of them may be now.) It’s best to give him only the information they can prove, and let him make the wild extrapolations, or perhaps inject their opinions in a mild, tentative, take-it-or-leave-it type of way.

In this dispute, there’s an attempt to split the baby by using a so-called “non-informative prior”. This is a contradiction in terms; the whole point of the prior is to provide information. If you have a prior belief, use it; if I disagree with that belief, I can substitute my own and get a different result. There can be a place for a mild, compromise prior, that satisfies neither of us but which both of us can accept as the price of agreement on other things. But it should be called a compromise, rather than insulting the world’s intelligence by calling it “non-informative”.

In any case, as a dispute about social role rather than about math itself, this can go on forever, and probably will. Of course the disputants themselves are almost all in the same social role: math professor. But they are coming up with theorems and formulas to be used by others, and the question is which sorts of others they prefer catering to. It is not, in any case, a vicious dispute; neither side is trying to exterminate the other or drive them out of academia. Thomas Dormandy has written that medical disputes are always particularly vicious when both sides are wrong. This is rather the reverse: both sides are right, at least in appropriate contexts.