Atlas · Methodology

How the verdict is produced.

The claim "this circuit is classically tractable" is only worth the method that backs it. Atlas does not predict simulability from features — it measures it, four independent ways, and certifies the case only when independent methods agree. That is how "you don't need the QPU" becomes a constructive, checkable result instead of a guess. Here is the full mechanism, including where it abstains.

Atlas (Krenn·IQ) · the methodology behind every result on the Evidence page

1 Four independent estimators

Each estimator measures the same circuit from a different mathematical space. Independence is the point: agreement among methods built on different mathematics is corroboration a single tool cannot fake.

Stim · #T / stabilizer

Magic (non-stabilizerness)

Counts the non-Clifford (T-gate) budget. A pure stabilizer circuit is classically simulable at any size; magic is what can push a circuit past the classical frontier.

Space: the stabilizer polytope · Clifford algebra
quimb · MPS bond

Entanglement / bond dimension

Measures the matrix-product-state bond needed to represent the state. Area-law / structured circuits stay cheap; volume-law entanglement makes the bond blow up.

Space: tensor networks · Schmidt rank
cotengra · treewidth

Contraction complexity

Measures the treewidth of the circuit's interaction graph — the cost of contracting it as a tensor network. Low-treewidth topologies contract cheaply regardless of depth.

Space: graph theory · circuit topology
Pauli-spread

Operator locality

Tracks how a local operator spreads under the circuit (scrambling). Slow spread means locality is preserved and classical methods stay tractable.

Space: Heisenberg picture · operator growth

These are open SoTA engines (Stim, quimb, cotengra) plus a measured operator-spread signal. Atlas's contribution is not the engines — it is the layer that runs them as independent witnesses and adjudicates the result. Reproduce: route_adjudicator.py.


2 The Certificate — agreement as convergent validity

When ≥2 independent folds agree "cheap", that is corroboration, not a single prediction. Atlas turns that agreement into a signed, hash-stamped certificate with five levels.

LevelAgreementMeaning
STRONG≥2 folds cheap, 0 hardclassically simulable — convergent validity
FIRMclassical majority, named dissentersimulable, dissent documented
WEAKa single fold carries itprovisional, thin evidence
SPLITeven spliton the frontier — the intractable region made observable
NULLno fold cheapQPU-required

Each certificate carries a SHA-256 content hash of the circuit, the per-fold signers (axis · cost · vote), the level, and the engine's actual route — an archivable, citable audit artifact, not a black-box verdict. Soundness battery: as the compute budget tightens, the level degrades STRONG → SPLIT → WEAK exactly as the independent folds begin to disagree, and 0 false-STRONG is ever issued on a hard verdict.

Reproduce: physics_magnitude_lab.certificate.certificate(n, circuit, budget_log2=30) → level + signers + hash. Soundness: scripts/certificate_validate.pycertificate_validation.json. Full discussion: Evidence §10.


3 The Convergence Map — disagreement made observable

The same call that certifies agreement also names which fold dissents, and in which direction. A SPLIT is not noise — it is the classically-intractable frontier (Leone-region) becoming visible.

When the treewidth axis and the MPS-bond axis disagree, that divergence is not a failure of the tool — it is a measurement of where the cheap structural proxies and the entanglement proxy stop agreeing about the same circuit. Atlas surfaces that divergence as a map rather than hiding it inside a single score. The disagreement tells a researcher exactly where the circuit sits on the simulability frontier, and which resource (magic, entanglement, topology, locality) is the one pushing it across. Two faces, one call: agreement certifies; disagreement maps the frontier.

This is the publishable core — formalising the treewidth↔MPS divergence as a central metric, with concordance against the noise-bound theory of Shao et al. (arXiv:2606.00474). Roadmap-tracked, not yet a closed result.


4 Five honesty principles

The architecture is constrained, by design, against the failure modes that make confidence scores dangerous — fail-closed guards plus a permanent regression battery where each is measured at zero. Design + zero observed, not a proof of impossibility.

  1. No false-STRONG. A STRONG certificate on a hard circuit is the failure Atlas is built to never make — and the permanent regression battery measures it at 0 false-STRONG to date. That is a measured zero under fail-closed design, not a guarantee of impossibility.
  2. Honest abstention = MEDIUM. Deciding exact stabilizer-polytope membership is provably super-exponential (Leone et al., arXiv:2602.22330). When the evidence splits, Atlas returns a calibrated MEDIUM — the theoretical ceiling speaking, not a bug. An exact classifier cannot exist, so feigning one would be the lie.
  3. No black box — an Evidence Ledger. Every verdict ships with the per-estimator signers, costs, votes, and a content hash. The reasoning is re-derivable from named scripts and data; nothing is asserted that cannot be reproduced.
  4. No lying translation. The plain-language layer never states more than the estimators measured. Natural-language summaries are constrained to the ledger — no constructive hallucination beyond what the numbers license.
  5. No anonymity. Every certificate is signed and attributable; the operator and the engine version are on the record. A rating no one stands behind is worth zero.

And it prices the decision. Beyond the verdict, Atlas estimates what the circuit would cost on a real QPU (per-shot vs per-minute, device-calibrated mitigation) and contrasts it with the measured classical cost — timed on a dev Apple M4. Full formula, sources and honest caveats: Execution economics →

5 Field context — the papers it builds on, cited honestly

Pre-flight simulability triage went from folklore to an active research topic in mid-2026. Atlas is the measured, multi-engine, hardware-corroborated point in that space — complementary to the predict-from-features and pure-theory work. Precisely: the device noise model reproduces hardware behavior (TVD ≈ 0.03–0.10 on tested families); independent QPU routing corroboration was attempted and honestly abstained. We cite this work because honesty about the landscape is the point; Atlas is first as a running product, not first to ask the question.

WorkWhat it doesHow Atlas relates
Leone, Eisert & Oliviero — arXiv:2602.22330Proves deciding exact stabilizer membership is super-exponential (Ω(2^(n²)) under ETH)The theoretical basis for why MEDIUM is the honest answer — an exact classifier provably cannot exist
Xing et al. — arXiv:2606.11620Family-aware ML that predicts the MPS-bond threshold from static gate features (~50 ms)Atlas measures the exact bond/treewidth — a predicted threshold can be silently wrong (false-security risk); a measured one fails loudly — it ships truncation/exactness flags, and the 1 boundary false-safety we did observe is published, not hidden
Del Rey et al. — arXiv:2605.28986Studies T-count & MPS bond as control variables for learning simulabilityAtlas uses the same two as decision variables, cross-validated against Stim, in a verdict
Shao et al. — arXiv:2606.00474Pure theory: when a polynomial TN bond suffices under noise (no tool)Atlas is the running implementation; the Convergence Map aims to be measured against this bound
Zhang & Zhang — arXiv:2409.13809 (PRX Quantum 6, 010337)Magic-depth-one circuits are poly-simulable; a sharp P→GapP-complete jump at depth-2Grounds magic-depth as a decision axis — Atlas treats shallow-magic as the tractable regime and measures the T-structure
Camillo et al. — PRX Quantum 7, 020356Optimal stabilizer-extent (magic) decompositions for multiqubit unitariesAnchors the magic-cost estimator (stabilizer extent) behind the fold(magic) route
Reardon-Smith et al. — arXiv:2307.12702Free-fermion (matchgate) + k non-free gates: graded classical sim, O(4.5^k) in controlled-ZA complementary resource axis (matchgate, not Clifford+T). Atlas now routes it live (free-fermion theorem route, since 2026-07): nearest-neighbour matchgates → CPU in poly time (Valiant 2001; Terhal–DiVincenzo 2002). Internally validated (detection 12/12+16/16, simulation 32/32 exact vs statevector) but pending independent external review (Jordan–Wigner / Majorana-covariance) — every matchgate verdict carries that caveat. The graded k-non-free extension is still roadmap.
Dowling et al. — arXiv:2605.18943Noise-induced simulability transition (operator scrambling); finite noise ≠ automatically classicalGrounds the device-calibrated noise frontier: noise tightens the classical side only past a threshold
Tirrito, Turkeshi, …, Hamma et al. — arXiv:2304.01175 (Phys. Rev. A 109, L040401, 2024)Proves nonstabilizerness (magic) is quantified by the flatness of the entanglement spectrum — anti-flatness measures the magic that is coupled into entanglementAtlas ships their concept as the live effective-magic signal: anti-flatness of the exact central-cut Schmidt spectrum, extracted at zero extra cost from the MPS the router already builds — validated on 56 real circuits (8/8 Cliffords read 0.0000; qft_n18 with #T=459 reads 0.000 — its magic cancels; af↔bond|#T=+0.57 vs af↔#T|bond=+0.06). The concept is theirs; the routing-level exposure is ours. Display signal only — exact-spectrum regime, abstains when truncated, does not change the route

Every citation above links to its arXiv primary source — don't take our word for the characterisation, click through and check it yourself. That is the review. Full landscape: Evidence §6.

Open Atlas → See the measured evidence