Pandas and Space Teapots Why extraordinary claims require extraordinary evidence

 

Pandas and Space Teapots

Why extraordinary claims require extraordinary evidence



There is a phrase popularized by Carl Sagan that often comes up in scientific and pseudoscientific discussions: “extraordinary claims require extraordinary evidence.” The phrase sounds reasonable enough, but what exactly does it mean? What makes a claim “extraordinary”? And why should we demand more evidence for some claims than for others?

There is actually nothing particularly mysterious about this principle. It is something we all apply constantly in everyday life, even if we do not normally bother to do the math. Let us look at a simple example.

Suppose a friend comes to visit me and, when he arrives, says:

—On my way here, I came across a cat in the street.

I will probably accept his claim without giving it much thought. I will not ask him for a photograph of the cat, search for security-camera footage, or demand independent testimony from other passers-by. His word is enough.

Now imagine that he says:

—On my way here, I came across a giant panda in the street.

My reaction will be rather different. I will probably start by asking whether he is serious. Perhaps I will think he is joking, that he mistook some other animal for a panda or, depending on the friend, that he is simply pulling my leg. If he insists, I will probably ask whether he took a picture. And if he keeps trying to convince me that he really did see a panda calmly walking down the street, I will start demanding evidence.

The curious thing is that the witness is exactly the same in both cases. If I previously considered my friend an honest person, I have no reason to suddenly think that he has stopped being one. But the fact remains that his testimony alone seemed perfectly sufficient in the case of the cat and completely insufficient in the case of the panda. Why?

Because I wasn't born yesterday.

I have accumulated an enormous amount of information about the world, and my friend's claims have to pass through the filter of that experience. If I knew absolutely nothing about reality beyond the fact that my friend is someone whose testimony I consider reliable, then the two claims, cat or panda, would be equivalent to me and I would regard them as equally plausible. But that is not the case. I know that cats are everywhere. I also know that pandas are very rare animals and that, outside a handful of very specific places, encountering one walking down the street is extraordinarily unlikely.

Just to put some orders of magnitude on that intuition, there are estimated to be around 600 million domestic cats in the world, whereas the known population of giant pandas in the wild is around 1,800 individuals, concentrated in a small region of China. This comparison should not be taken as a literal calculation of the probability of encountering either animal in the street—for that we would have to consider geographical distribution, captivity, behavior, and so on—but it illustrates the enormous difference in our starting point.

And that difference is precisely what matters.

I have prior information.

Bayesian mathematics

This intuitive way of reasoning is also reflected in the way we do science. When we want to compare how plausible our hypotheses are in light of the available data, we turn to Bayes' theorem. Let us see how to apply it to our previous example.

We want to compare how probable it is that my friend really did see a giant panda—let us call this HPH_P, the panda hypothesis—with the hypothesis that my friend was mistaken and there was actually no panda—let us call this HNPH_{NP}, the no-panda hypothesis.

We will denote by DD, the data, my friend's testimony that he saw a panda. The reliability of our friend as a witness enters through what is called the likelihood, which tells us how probable it is that he gives a correct or incorrect account. Technically, we are talking about the probability of DD conditional on HPH_P or HNPH_{NP}, denoted by:

P(DHP)\boxed{P(D\mid H_P)}

and

P(DHNP)\boxed{P(D\mid H_{NP})}

Bayes' theorem tells us how to obtain the probability we are looking for, P(HPD)P(H_P\mid D), which is simply the probability of the panda hypothesis conditional on the data DD:

P(HPD)P(DHP)P(HP)\boxed{ P(H_P\mid D)\propto P(D\mid H_P)\,P(H_P) }

Remember that P(DHP)P(D\mid H_P) represents the reliability of my friend—the probability that he gets it right. Our friend is very reliable, which is why we trust him. Let us say he is right 99% of the time. But he is also human and occasionally makes mistakes, say 1% of the time. Thus,

P(DHP)=0.99\boxed{ P(D\mid H_P)=0.99 }

and

P(DHNP)=0.01\boxed{ P(D\mid H_{NP})=0.01 }

The other term, P(HP)P(H_P), is what is called the prior. It represents the probability of the hypothesis before we have the data—which is why DD does not appear in that factor. And this is where the key lies: the “I wasn't born yesterday” part is incorporated into this term.

We can combine what we know about panda and no-panda simply by dividing the probabilities:

P(HPD)P(HNPD)=P(DHP)P(DHNP)P(HP)P(HNP)\boxed{ \frac{P(H_P\mid D)} {P(H_{NP}\mid D)} = \frac{P(D\mid H_P)} {P(D\mid H_{NP})} \, \frac{P(H_P)} {P(H_{NP})} }

On the left-hand side we have the ratio between the probabilities of the two hypotheses once we have received the data. It tells us how probable the two competing hypotheses are after hearing our friend's testimony. That is precisely what we want to find out: how likely is it that my friend really did see a panda?

On the right-hand side we have two terms. The first is the reliability of our friend. As we said before, our friend is right 99% of the time, so this term is:

0.990.01=99\boxed{ \frac{0.99}{0.01}=99 }

The term on the far right contains the priors. This is where we put the information from our previous knowledge, the knowledge we already had before receiving the data. As we know, we live on planet Earth, where cats outnumber pandas by something of the order of a million to one. We can therefore conservatively assign a value of one million to this factor in favor of no-panda.

Different people might assign a somewhat different value here. We all have different experiences of the world and different prior knowledge. But I do not think any reasonable person would put the figure much below a million in favor of no-panda. Remember that this is the probability before our friend has told us anything. We have to imagine how we would answer if somebody asked us about the probability that our friend had encountered a panda before hearing anything from him.

So, putting in the numbers:

P(HPD)P(HNPD)=99×11000000\boxed{ \frac{P(H_P\mid D)} {P(H_{NP}\mid D)} = 99\times\frac{1}{1\,000\,000} }

Therefore,

P(HPD)P(HNPD)9.9×105\boxed{ \frac{P(H_P\mid D)} {P(H_{NP}\mid D)} \simeq 9.9\times10^{-5} }

Or, turning the ratio around:

P(HNPD)P(HPD)10000\boxed{ \frac{P(H_{NP}\mid D)} {P(H_P\mid D)} \simeq 10\,000 }

The result is still of the order of 10,000 to one in favor of the no-panda hypothesis, even though we chose a very conservative prior. My friend's testimony has changed my perception of the probability of encountering pandas in the street, but not enough for me to believe him.

What, then, is “extraordinary evidence”?

“Extraordinary evidence” does not mean exotic or spectacular evidence, nor evidence obtained using some particularly sophisticated instrument. It means evidence that produces a likelihood ratio large enough to compensate for prior knowledge that assigns an extremely low probability to the hypothesis.

If, before receiving the data, a hypothesis has a probability of one in a million, then to bring it even approximately up to 50% we need evidence with a likelihood ratio of the order of a million.

That is all.

An extraordinary claim is, in this sense, a claim with an extraordinarily low prior probability given our existing knowledge. And extraordinary evidence is evidence capable of overcoming that enormous initial handicap.

Note that this does not mean being irrationally biased and refusing to accept new evidence. It means that we were not born yesterday, that we have already accumulated a certain amount of knowledge, and that there are some things we know to be extremely uncommon. We are back to pandas walking down the street.

Of course, it does not have to be a single spectacular piece of evidence. Many moderately strong and independent pieces of evidence can accumulate. If three independent observations each contributed a factor of 100 in favor of a hypothesis, their combined effect would be one million:

100×100×100=1000000\boxed{ 100\times100\times100=1\,000\,000 }

Science works this way all the time. Initially implausible hypotheses are not forbidden, nor can they simply be dismissed by decree. They merely have to accumulate enough evidence to establish their validity.

And if the data are good enough, they will overcome that initial disadvantage.

This last point is important because “extraordinary claims require extraordinary evidence” is sometimes used incorrectly, as a kind of barrier intended to protect established ideas against anything new.

That is not what it means.

A small prior probability is not a probability of zero. Bayes allows a hypothesis that was initially extremely improbable to end up being virtually certain if sufficiently compelling data appear. In fact, doing otherwise would be rather unscientific.

Russell's teapot

In 1952, the philosopher and mathematician Bertrand Russell proposed roughly the following argument. Imagine that somebody claims that there is a small china teapot orbiting the Sun somewhere between Earth and Mars, but that it is too small to be observed by our telescopes. We could not prove that the teapot does not exist. But it would be absurd to conclude from our inability to disprove it that we should regard its existence as reasonable.

Here again we encounter a fairly common confusion:

“I cannot prove that it is false” does not mean “both possibilities are equally probable.”

I do not have an observation that absolutely rules out Russell's teapot. But I wasn't born yesterday either.

I know many things about teapots. I know that they are objects manufactured by human beings. I know roughly where they are made and what they are used for. I also know quite a lot about the objects we have sent into heliocentric orbits.

All that information forms part of the prior knowledge with which I evaluate the hypothesis. That is why my prior for a teapot spontaneously orbiting the Sun is extraordinarily small.

But now imagine a slightly different story.

In February 2018, SpaceX launched its Falcon Heavy rocket for the first time. As a test payload it carried a red Tesla Roadster, with a spacesuit-clad mannequin—Starman—sitting behind the wheel. The upper stage eventually sent the car into a heliocentric orbit that extends even beyond the orbit of Mars.

Now imagine that, before launch, SpaceX had publicly announced that Starman would be holding a china teapot in his hand.

Years later, somebody asks me:

—Do you think there could be a teapot orbiting the Sun somewhere in the region between the orbits of Earth and Mars?

My answer would now be very different.

I would no longer be so categorically certain that space teapots do not exist. Perhaps I do not know whether the teapot is still attached to the car. Perhaps I do not know whether it came loose during launch or sometime later. Perhaps I have no idea what its present trajectory might be. I have no recent photograph of it and I have not observed it through a telescope.

But I would no longer regard its existence as absurd.

What has changed?

Not the laws of physics. Nor my present ability to observe the teapot.

What has changed is my prior information.

In strictly Bayesian terminology, knowing that a teapot was launched into space is also evidence. But once that information has been incorporated into our knowledge, it becomes part of the background information on which we establish our prior before receiving new data.

The difference between the two teapots is not that one is falsifiable and the other is not. It is that our accumulated knowledge of the world gives them radically different prior probabilities.

We have different priors.

And this also shows why complete ignorance does not imply a probability of 50%.

If I had been born five minutes ago knowing absolutely nothing about teapots, planets, or rockets, I probably would not have enough information to assign a well-informed probability to the existence of a space teapot.

But it does not follow that I should assign a probability of 50% to its existence and 50% to its non-existence.

Not knowing something does not make all possibilities equally probable.

We weren't born yesterday

That, in fact, is perhaps the central idea behind all of this.

We never evaluate a claim in a vacuum. We evaluate it on top of an enormous mountain of accumulated knowledge: our personal experience, previous observations, experiments, well-tested scientific theories, and everything humanity has learned up to that point.

When we formalize the problem in Bayesian terms, we call that previous knowledge the prior.

If somebody tells me that he saw a cat, his testimony fits comfortably with everything else I know about the world. A small amount of additional evidence is enough.

If he tells me that he saw a panda, the same testimony has to struggle against a much lower prior probability.

And if he tells me that he saw a twenty-meter dragon flying over the city, my skepticism will be greater still. Not because I am dogmatically committed to the proposition that dragons do not exist, but because that hypothesis conflicts with an enormous body of previous knowledge, while there are many alternative explanations that are vastly more probable: a joke, a hoax, a hallucination, a fake, a misidentified object, or simply a lie.

Of course, the dragon could exist.

But then I want to see it.

(And even then I would need more evidence.)



Disclaimer: Generative AI has been used to assist in the writing of this article

Comments

Popular posts from this blog

Los OVNIs del New York Times

The fallacies behind the cult of Loeb

3I/ATLAS. A cat on my balcony