Quantum Academy begins operations on September 15, 2026. Enrollment opens soon.
Skip to content

Market Reality

Evaluating a Quantum Company: The Claims

Marin Ivezic11 min read

A technology slide arrives with one number on it: 256 qubits. The number is accurate. The company really does have 256 qubits, the hardware exists, and anyone who visited could count them. You still know almost nothing, because that number sits on the slide with no unit attached, no sample described, no operating conditions given, and nobody outside the company willing to say they measured it.

Most quantum diligence goes wrong at exactly this point, long before anyone reaches the physics. The investor who gets burned is usually not the one who missed a subtlety in the error model, but the one who accepted a true number as the answer to a question nobody had specified.

This piece is about claims rather than companies. The engineering behind a claim and the business built on top of it are separate assessments, each with its own questions. Before either of those is worth starting, you need to know which of the company’s statements can be scored at all.

A quantum performance claim is scoreable when five things are present: the unit, the population, the conditions, the baseline, and the witness. A claim missing one of these isn’t necessarily false. It’s unscored, which means you can’t yet mark it right or wrong, and moving claims out of that state is the whole job.

The unit

The unit is the distinction that decides whether a four-figure qubit count means anything at all.

A physical qubit is the hardware object: a superconducting circuit on a chip, a single trapped ion, a photon in a waveguide. Left alone, it is noisy, and it holds its quantum state for a limited time. A logical qubit is one qubit of protected information, spread across many physical qubits, with error correction running the whole time it exists. Quantum error correction (QEC) is the machinery doing that protection: the system takes extra measurements, called syndrome measurements, which reveal that an error occurred and where, without measuring the encoded data itself and destroying it.

The exchange rate between the two is not a constant. It depends on the physical error rate, the code, and how long the computation needs to run. For a concrete sense of scale, take the most-modelled task in the field. In 2019, Craig Gidney and Martin Ekerå estimated that factoring RSA-2048 would take roughly 20 million physical qubits running for about eight hours. Gidney’s 2025 revision brought the requirement down to under a million, on better codes and better algorithms. Neither figure describes anything a vendor can sell you today, and both describe the same handful of logical qubits sitting on top of an enormous physical substrate.

So when a company states a qubit count, the count is a raw material, not a capability. The questions are short:

  • Are those physical or logical qubits?
  • If logical: which code, how many physical qubits per logical qubit, and what logical error rate did you measure?
  • Is the logical error rate below the physical error rate of the constituent qubits, and by how much?

That last question is the one that separates error correction from error mitigation. Mitigation techniques reduce the effect of noise in post-processing and do not produce a protected qubit. Both are legitimate engineering. Only one of them is QEC, and the vocabulary gets blurred in press releases far more often than in papers.

The population

The population is which qubits, and how many of them, a number was measured across – one best pair, or the whole device.

Gate fidelity is the accuracy of a single operation, stated as 100% minus the error rate. A 99.9% two-qubit gate fails roughly once in a thousand attempts. Vendors quote fidelity constantly, and the number is nearly useless without knowing which qubits it describes.

Three population questions cover most of it. Is this a single-qubit gate or a two-qubit entangling gate, since the two-qubit number is almost always the worse one and almost always the constraint? Is it the best pair on the device or the median across all pairs? And on what date was it measured, since devices drift and calibration is a daily activity rather than a one-time event?

The gap between best and median is not cosmetic. Take a device where the best pair achieves 99.9% and the median across twenty pairs is 98.5%. On a circuit with 500 two-qubit gates, the best-pair number predicts that about 61% of runs complete without an error. The median predicts about one run in 2,000. Same chip, same week, same honest measurements, and two conclusions that have nothing to do with each other.

Ask which method produced the number. Randomized benchmarking runs long sequences of random gates that should return the qubit to a known state, then infers an average error rate from how fast the result degrades. It’s the standard and it’s reasonable, and it also averages away certain correlated errors, so a benchmarked figure can be better than what an actual algorithm experiences. A company that volunteers this limitation without being asked has told you something useful about how it reports everything else.

The calibration table is the document you want: every qubit, every pair, with dates.

The conditions

Coherence time is how long a qubit holds usable quantum information before ambient noise degrades it. T1 is the time it takes to lose energy and relax into its ground state. T2 is the time it takes to lose phase relationships, which is what superposition and entanglement depend on, and it is usually the shorter and more relevant of the two.

A coherence figure measured on an idle, isolated qubit in a quiet fridge is a ceiling, not a working number. The figure worth having is the ratio: how many operations fit inside the coherence window. A 100-microsecond T2 with 200-nanosecond two-qubit gates leaves room for roughly 500 operations. The same 100 microseconds with 10-microsecond gates leaves room for ten. The headline is identical and the machines are not comparable.

This ratio also explains why cross-modality comparisons collapse so easily. Trapped-ion systems report coherence times orders of magnitude longer than superconducting systems, and they also run gates orders of magnitude slower. Neither number tells you which platform runs a deeper circuit. The quotient starts to.

So ask what the qubit was doing during the measurement, whether the figure holds when neighbouring qubits are being driven, whether it was measured on one qubit or across the register, and whether it survives a week of continuous operation rather than a good afternoon.

The baseline

A quantum advantage claim asserts that a quantum machine did something useful better than a classical computer could. Quantum supremacy, the phrase John Preskill introduced in 2012, is the weaker and more precise idea: a quantum machine doing something a classical machine practically cannot, whether or not the task has any use.

Advantage claims have a structural oddity. They are claims about somebody else’s software. The quantum side of the comparison is under the company’s control and the classical side is not, which means the claim can be falsified by people who never touch the hardware.

The best-documented example is public. In October 2019, Google reported in Nature that its 53-qubit Sycamore processor completed a random circuit sampling task that would take the leading classical supercomputer on the order of 10,000 years. IBM published a counter-analysis within days, arguing that with better use of disk storage the same task could be done classically in roughly two and a half days. Classical simulation methods have kept improving since. Nobody’s measurements were wrong. The baseline moved, because baselines are contested by people who are trying hard to move them.

The questions follow from that:

  • Which specific classical algorithm was the comparison run against, on what hardware, by whom?
  • Was the classical side tuned by someone who wanted it to win?
  • Is the problem one a customer would pay to solve, or one chosen because it is hard for classical machines?
  • Is the result on arXiv or in a journal, and if not, what is the reason?

A useful closing question in the meeting: name the classical researcher most likely to beat this result, and tell us whether they have seen the data. Teams with a real result usually have an answer, and sometimes a name and a date.

The witness

The last part is the cheapest to check and the most predictive. Who outside the company has seen the thing.

Witnesses come in grades. A preprint is weaker than peer review, peer review is weaker than independent replication, and all of them are weaker than a device sitting on a public cloud queue where anyone can run their own circuits and publish what they find. Google’s December 2024 surface-code result in Nature is a reasonable template for what a witnessed claim looks like: a named code, specified distances of 3, 5 and 7, a logical error rate falling by roughly half with each step up in distance, published data, and a field full of specialists arguing about the details in public.

The absence of a witness is not by itself a problem. A company two years old with an unpublished result is normal, and early teams have real reasons to hold data back. What separates the two cases is the response when you ask. A team that says “not yet, here’s what we can show under NDA, and here is who we have invited to test it” is early. A team that treats the question as hostile has told you where the risk sits.

Extraordinary claims raise the witness bar rather than lowering it. Room-temperature operation is the recurring one. Certain modalities, including nitrogen-vacancy centres in diamond and some photonic approaches, genuinely work at ambient temperature, and none of them has yet matched cryogenic systems on scale and gate fidelity together. A claim that skips cooling and delivers scale and delivers fidelity is a claim that several very well-funded laboratories have been beaten simultaneously, and it needs evidence in proportion.

Security claims need a sixth part

Post-quantum cryptography claims fail differently, because the unit is a published standard rather than a measurement. “Quantum-safe” with no algorithm named is a marketing phrase and nothing else.

The named algorithms are ML-KEM for key encapsulation (FIPS 203), ML-DSA for digital signatures (FIPS 204), and SLH-DSA as the hash-based signature alternative (FIPS 205), with FN-DSA following. Ask which algorithm and which parameter set, whether the implementation has been through validation testing, whether the product runs the new algorithm alongside a classical one in hybrid mode or on its own, and where key management actually sits. Quantum key distribution claims are a separate category again, and the questions there are about the trust model and the certification regime rather than about algorithms.

For the broader migration questions behind these claims, the methodology at pqcframework.org covers the inventory and sequencing work that a vendor answer has to fit into.

Claim triage in one table

The claimUsually missingWhat to ask for
“We have N qubits”UnitPhysical or logical, plus error rates and connectivity
“99.9% fidelity”PopulationFull calibration table, one-qubit and two-qubit, with dates
“Coherence of 100 µs”ConditionsGate duration, and whether it holds during computation
“Quantum advantage on X”BaselineThe classical algorithm, hardware, and who tuned it
“Error-corrected qubit”Unit and witnessCode, physical-to-logical ratio, logical error rate, publication
“Quantum-safe product”StandardNamed algorithm, parameter set, validation status, hybrid mode

What can be done in a week

Three checks fit into a normal diligence window and between them they resolve most unscored claims.

The paper trail. An afternoon of searching arXiv, Google Scholar and the patent record for the founders’ names tells you what the underlying work actually showed, and whether the pitch version has grown since. Most quantum companies come out of academic groups, so the trail exists.

One expert call. Thirty minutes with a specialist in the right subfield, not just any physicist, will usually settle whether a claimed number is plausible for that modality. Bring specific numbers to the call rather than the deck.

The classical rerun. For any advantage claim on a software product, have someone competent implement a properly tuned classical baseline on the same problem. This is the single most decisive week of work available in quantum diligence, and it’s often the cheapest.

What a well-formed claim sounds like

Here is the same performance statement with all five parts present:

Across 20 two-qubit pairs on our 32-qubit device, randomized benchmarking on March 3 gave a median error of 1.2% and a best pair of 0.4%. Gates run at 240 nanoseconds against a median T2 of 90 microseconds. The circuits and raw data are in our preprint, and the device has been on a public cloud queue since January, so the measurement can be repeated by anyone.

It’s longer than “world-class fidelity” and considerably less exciting. It’s also the version you can act on, argue with, and check against the same company’s numbers six months later. Teams that write sentences like this tend to reward more diligence rather than less, because everything else they say is built the same way.

The habit underneath all of this is small and portable. Take the claim apart before you take a position on it, and ask which of the five parts is absent. Most of the time the missing part is the answer.

Quantum Academy’s programs teach this vocabulary and this method to people who have to make decisions about quantum technology without a physics background, working from real vendor claims and real benchmark data rather than analogies. You can see the current programs at quantumacademy.com/. For the longer technical treatment this article draws on, including the wider diligence framework, see the original analysis at PostQuantum.com.