In August 2024, Google’s quantum computing group posted a preprint, Quantum error correction below the surface code threshold, describing the superconducting processor it would later name Willow, with 105 physical qubits. The number most readers carried away was 105. The result inside the paper was something else. As the team scaled its error-correcting code, the arrangement of many ordinary qubits into one protected qubit, from distance 3 to distance 5 to distance 7, the error rate of the protected qubit fell at each step instead of rising. Errors going down as the machine gets bigger is the behaviour error correction has needed since the 1990s, and it says more about where the field stands than the count of 105 does.
That gap, between the number in the headline and the result in the paper, is where most bad quantum procurement decisions begin. Carl Sagan built a “baloney detection kit” for exactly this problem, a short set of habits for telling a real claim from a dressed-up one. Murray Gell-Mann, who had watched physics vocabulary get borrowed by people with no interest in physics, called the dressed-up version quantum flapdoodle. Both were describing the same phenomenon we now see in vendor decks and funding announcements.
We teach this material because buyers keep asking us the same question in different words: how do we tell, without a physics degree, whether a claim in front of us is solid? The answer is more mechanical than most people expect.
Overstated Quantum Claims Rarely Lie
A quantum claim that gets a company into trouble is usually true as written. It misleads by substitution. Somewhere in the sentence, one quantity has been swapped for a similar-sounding quantity that would have been much harder to earn, and the reader supplies the stronger meaning without being asked to.
That’s good news for anyone doing evaluation, because there are only about six substitutions in common use. Once you can name them, you can test for them, and the test is usually a single question you can put in an email.
Physical qubits for logical qubits
A physical qubit is one piece of hardware, a superconducting circuit or a trapped ion, that holds quantum information for a short time and makes errors at a measurable rate. A logical qubit is one protected qubit assembled from many physical ones, using an error-correcting code that detects and repairs those errors as the computation runs. The ratio between them is not small. In September 2024, Microsoft and Quantinuum reported 12 logical qubits built from 56 trapped-ion physical qubits, and that ratio is unusually favourable because trapped ions have low error rates to begin with. On superconducting hardware the ratio is far worse, and the Google result above is the illustration: a 105-qubit chip running one distance-7 logical qubit.
So “500 qubits” and “500 logical qubits” describe machines separated by a decade or more of engineering. Almost nobody writes the second sentence falsely. They write the first one and let the reader round up.
The test: ask which kind, and if the answer is physical, ask for the two-qubit gate error rate and the coherence time, meaning how long a qubit holds its state before noise destroys it. A vendor who cannot produce those two numbers within a day is not measuring their own hardware carefully.
A roadmap date for a demonstrated result
Roadmaps are legitimate planning documents and every serious hardware company publishes one. They become a substitution when a date on a slide is quoted back as a capability. “By 2029 the system will support fault-tolerant workloads” is a plan. “The system supports fault-tolerant workloads” is a claim. Decks blur the two constantly, usually by putting the date in small type under a bold verb.
There’s a related move in the other direction that we think is honest and worth recognising: IBM announced Condor, a chip with 1,121 physical qubits, at its Quantum Summit on December 4, 2023, and then shifted its public emphasis to the smaller Heron line, which has fewer qubits and better error rates. Reading a company de-emphasise its own largest number tells you something real about what the field now considers a meaningful metric. The market rewards the count because the count fits in a headline; the engineering rewards the error rate.
The test: for every capability in a proposal, ask whether it exists on hardware today, exists in simulation, or exists on a roadmap. Three columns, one per claim.
A benchmark for a problem
Google’s 2019 Sycamore experiment ran random circuit sampling, a task that asks a quantum chip to produce output patterns from a randomly chosen circuit. The task was selected because it is hard to simulate classically and natural for the hardware to perform. It has no application. That was never hidden, and it doesn’t diminish the experiment, but it does mean the result answers a question about physics rather than a question about your business.
Quantum annealing, a different approach that searches for low-energy configurations of a physical system rather than running a general circuit, has produced a long series of advantage claims on optimisation-shaped tasks. Several have later been matched or beaten by improved classical algorithms written specifically to attack the same problem. The pattern repeats often enough to be predictable, which is why claims of annealing advantage on production-scale logistics or scheduling deserve slow reading.
The test: ask what the task was, in one sentence, without the word quantum in it. Then ask whether anyone outside the vendor needs that task done.
A comparison without a control
This is the most common substitution in commercial material, and the easiest to catch. “One hundred times faster” is a ratio, and a ratio needs a denominator. Faster than what classical method, tuned by whom, on what hardware, at what accuracy?
Sycamore again gives the clearest illustration. Google’s 2019 paper estimated that reproducing its 200-second run would take a classical supercomputer on the order of 10,000 years. IBM responded within days with an approach using a supercomputer’s disk storage that it argued would take about two and a half days. Later classical work narrowed the gap further. Nobody was lying at any stage. The quantum figure was measured and the classical figure was an estimate against a moving target, and the target moves because a public quantum claim is an invitation to every classical algorithms group on earth.
The test: ask who wrote the classical baseline and when. If the baseline is a textbook method or a five-year-old library, the comparison is decorative.
A press release for a result
Real quantum results carry a paper, either in a journal or as a preprint on arXiv, with methods detailed enough for another group to attempt the same thing. Companies have genuine trade secrets and will not publish everything, and that’s fine. What should worry us is a headline capability with no technical artefact behind it at all, anywhere, six months later.
Cloud access is the commercial equivalent of a methods section. IBM, IonQ, Rigetti, Quantinuum and others expose hardware to outside users, which means an exaggerated performance figure gets contradicted by customers rather than by critics. A closed device whose results only its owner has ever seen is a different proposition, whatever the results are.
The test: ask for the paper, the preprint, or the account credentials. One of the three should exist.
A word for a mechanism
The last substitution swaps vocabulary for explanation. “Our quantum-enhanced AI platform delivers resilient optimisation across the enterprise” contains no testable statement. Genuine technical description is narrower and duller: this modality, this qubit count, this gate fidelity, this problem class, this comparison.
The physics itself gets substituted too. The single most repeated misstatement in quantum marketing is that a quantum computer tries every possible answer at once and therefore returns the right one immediately. A quantum computer does hold many states simultaneously, but reading the machine collapses that to one outcome, so the whole craft of quantum algorithms is arranging interference so that wrong answers cancel and right ones survive. Only some problems have that structure. A pitch built on the parallel-universe story is either a pitch from someone who doesn’t know this or a pitch that assumes you don’t.
The test: ask the speaker to restate the core claim without using the word quantum. If nothing survives, nothing was there.
Running the Toolkit on a Real Case
In December 2022, a group of Chinese researchers posted a preprint titled “Factoring integers with sublinear resources on a superconducting quantum processor”. Its closing extrapolation, that RSA-2048, the encryption keystone of most of the internet, could be broken with 372 physical qubits, was the line that travelled. Hardware at that scale already existed. Several outlets, the Financial Times among them, ran the story as an imminent cryptographic emergency.
Take the substitutions in order.
The number 372 is a physical qubit count, and the paper’s own method depends on those qubits behaving well enough for a long optimisation to converge, which nothing in the paper establishes. The 372 figure is also an extrapolation rather than a demonstration, the same gap as a roadmap date quoted back as a capability: the experiment the team actually ran factored a 48-bit integer using 10 qubits, many orders of magnitude away from a 2048-bit key. The task was chosen to fit the method rather than the method built for the task. There was no classical control, because the classical component of the approach, a lattice-based factoring technique due to Claus Schnorr, was itself already disputed in the cryptographic community and had not been shown to scale. And the load-bearing element was the Quantum Approximate Optimization Algorithm, a heuristic that searches for good solutions without guaranteeing it will find them, described in the paper without any argument that it converges at the sizes required.
Only the fifth test came back clean. The preprint did exist and was public, which is to the authors’ credit and is precisely how the problem got found.
Five of the six substitutions, in one document, in a paper that never states a falsehood. Scott Aaronson’s public assessment at the time was blunt about the gap between what was shown and what was inferred, and nothing since has changed the position: RSA-2048 has not been factored, and the migration deadlines organisations are working to were set by standards bodies and regulators, not by that preprint.
We use this case in our own teaching because it demonstrates something a checklist alone cannot. Every red flag was visible to a careful reader in under an hour. The people who got it wrong were not fooled by fraud. They read a technical document at the speed of a press release.
The Failure Mode on the Other Side
A toolkit like this one has a predictable side effect, and we’d rather name it than watch it develop. Trained scepticism turns into reflex dismissal, and reflex dismissal is just as expensive.
Two things follow from that. The first is that criticism carries incentives exactly as promotion does. When Scorpion Capital published a short-seller report in 2022 alleging that IonQ’s achievements were overstated, the correct response was not to accept the report because it was critical, any more than to accept a company’s own deck because it was confident. A short-seller profits when the share price falls. The report contained specific technical assertions, and the way to handle it was to check the assertions. Apply the same six tests to the accusation.
The second is that the field is genuinely progressing, unevenly and slowly, and the progress is easy to miss if you have decided in advance that all of it is noise. Error rates are falling. Below-threshold error correction has been demonstrated on more than one hardware modality. Logical qubit counts have moved from zero to low double digits in three years. None of that supports a claim that your encryption breaks next year, and all of it supports a claim that migration planning should already be under way in organisations with long data confidentiality requirements. Scepticism is a method for finding out which claims are true, not a conclusion that none of them are.
Six Questions to Send Back
The whole toolkit compresses into an email. For any quantum capability claim in front of you:
- Physical or logical? If physical, what are the two-qubit gate error rate and the coherence time?
- Today, in simulation, or on the roadmap? One answer per claimed capability, in writing.
- What was the task, stated in one sentence without the word quantum?
- What is the classical baseline, who wrote it, and when?
- Where is the paper, the preprint, or the hardware access?
- Restate the core claim with no quantum vocabulary in it. What remains?
A vendor with real technology answers all six comfortably, because the answers are the specifications they work with daily. A vendor who cannot answer question one has told you what you need to know before you reach question two.
Building the Judgment Behind the Questions
The questions are the easy part. Knowing whether an answer is a good one requires knowing what a normal gate error rate looks like in 2026, which modalities need dilution refrigerators and which do not, why connectivity constrains circuit depth, and how the published benchmarks relate to each other. That knowledge is what turns a checklist into judgment, and it’s the difference between an evaluator who can be reassured by a confident answer and one who cannot.
Quantum Academy’s programs are built for people in exactly that position: professionals who have to make procurement, architecture, and risk decisions about quantum technology and need the technical grounding to do it without depending on the vendor’s own framing. You can see the current programs and what each one covers at quantumacademy.com/. For deeper technical context on quantum computing and post-quantum security, PostQuantum.com covers the ground at length.
Sagan’s advice was to keep an open mind without letting your brains fall out. In quantum technology, the open mind is warranted. So is the practice of asking, every single time, which of the six substitutions is being made.