Dynamical Decoupling Erased a Subthreshold Result on IBM Heavy-Hex
Add dynamical decoupling to a surface code running on IBM heavy-hex hardware and an apparent subthreshold scaling trend can vanish. That reversal is the sharpest caution in a Nature Communications paper published on 29 July 2026, and it raises the bar for what counts as a reported distance-scaling result.
Why weight-4 stabilizers do not fit a heavy-hex lattice
The surface code measures weight-4 stabilizers, so it assumes a data qubit can reach four neighbours. IBM's superconducting processors do not offer that. The heavy-hex lattice uses three distinct frequency assignments and places the control qubit only on edges connected to target qubits, which IBM says minimizes frequency collisions and the spectator errors generated by qubits that are idle during a gate (IBM Quantum).
That topology arrived with a code designed for it. Chamberland and colleagues introduced the heavy-hexagon code, a hybrid surface and Bacon-Shor subsystem code mapped onto the lattice, where flag qubits let a modified matching decoder correct to the full code distance under the limited connectivity (Physical Review X 10, 011022). Running a plain surface code on heavy-hex means paying in SWAP gates for connectivity the hardware never promised.
The fold-unfold embedding that holds a stabilizer round to depth 7
Vezvaee, Benito, Morford-Oberst, Bermudez and Lidar co-design the code embedding and the control pulses together (Nature Communications 17, 9201). A folding step applies CNOTs between data qubits to compress a weight-4 stabilizer into a weight-2 operator, an ancilla measures that weight-2 parity, and an unfolding step restores the original stabilizer.
Next-nearest-neighbour CNOTs, which heavy-hex does not couple directly, route through bridge qubits using SWAP gates. Those bridge qubits are never measured. The embedding holds a stabilizer measurement round to a depth-7 circuit, and because only half the stabilizers measure in parallel, a full syndrome extraction cycle takes two rounds.
The codes tested occupied 37 qubits at distance (3,3) and 65 qubits at both (3,5) and (5,3), on ibm_aachen with ibm_marrakesh as secondary validation. Both are 156-qubit Heron devices, which leaves the hardware 6 qubits short of a fully isotropic step from distance 3 to distance 5.
The subthreshold trend that dynamical decoupling reversed
Without DD, the paper reports that logical error rates can decrease as code distance grows, which reads as subthreshold behaviour. Apply DD and that trend may reverse. Once each code was optimised separately for DD inclusion, the subthreshold scaling result on ibm_aachen disappeared.
Idle time is the mechanism. Qubits waiting through a syndrome extraction cycle accumulate coherent ZZ crosstalk and non-Markovian dephasing, which the authors name as the main hardware error sources. DD suppresses both, and suppressing them removes the channel that was making the larger code look better than it measured.
Gap-aware means the pulse count is tuned to each idle window rather than fixed once. The authors searched the UR sequence family for pulse counts from 6 to 18 in steps of two, and used XY4 and RGA8a on shorter gaps, adding up to roughly 484 single-qubit DD pulses per cycle. Applying one sequence uniformly across every gap degraded performance, so the sequence search belongs inside the experiment rather than after it.
What anisotropic distance scaling bought and what it cost
Growing the X distance and the Z distance independently is cheaper than growing both, and the paper tests whether the trade pays. A (3,5) code protects the Z basis against phase flips, a (5,3) code protects the X basis against bit flips, and both reached an aligned-basis suppression factor between 1.23 and 1.46.
The basis-averaged factor sat well below 1. Protection bought this way is directional: a logical state in the orthogonal basis degrades without a matching benefit, so the code helps only if you already know which error type your algorithm is exposed to. Comparing the best sublattice narrows the margin to between 1.01 and 1.10.
For scale, Google Quantum AI reported a suppression factor of 2.14 ± 0.02 per two units of distance on Willow, whose square grid matches the code, culminating in a 101-qubit distance-7 memory at 0.143% error per cycle and a logical lifetime 2.4 times its best physical qubit (Nature 638, 920). Matching the lattice to the code is worth roughly a factor of two in Lambda.
What would push a heavy-hex surface code below threshold
The authors put a number on the gap. A calibrated noise reduction of around 30%, together with a slightly larger heavy-hex device, would deliver genuine subthreshold scaling by their projection. Both conditions are incremental, which is the encouraging half of a paper whose headline result is negative.
Until that lands, treat any heavy-hex distance-scaling claim as incomplete unless it states how DD was optimised for each code separately. The two error channels the paper blames, coherent ZZ crosstalk and non-Markovian dephasing across idle gaps, are precisely the ones a noise model fitted to a single calibration snapshot will miss.
SuperconducTED attacks that staleness with fuzzy inference over IBM calibration time series, and the earlier post on error detection with probabilistic error cancellation covers the mitigation side of the same hardware. Read the SuperconducTED calibration-drift approach before trusting a noise model across a recalibration.