For a noisy quantum channel, the most basic question you can ask is how many classical bits per use it can carry. For most channels we still cannot answer it, because the formula for the classical capacity involves an optimization over entangled inputs spread across arbitrarily many channel uses. A new preprint, “Classical Capacity and Entanglement Cost of the Amplitude Damping Channel” by Ziao Tang, Chengkai Zhu, Ge Bai and Xin Wang, claims to close that gap for the qubit amplitude-damping channel, the textbook model of energy relaxation, and for a whole class of qubit channels around it.
This article was drafted with Claude and fact-checked against the cited primary sources. The paper is a v1 preprint and has not yet been peer reviewed; below I separate what the paper proves from my own interpretation, which is labelled as such.
Why this paper is worth your attention
The Holevo–Schumacher–Westmoreland theorem expresses the classical capacity C of a channel N as a limit: C(N) = lim (1/n) χ(N⊗n), where χ is the Holevo information. Using product inputs gets you χ(N) per use; the question is whether entangling codewords across uses can ever do better. Hastings showed in 2009 that for some channels it can, so additivity cannot be assumed. Proving it for a specific channel requires a structural reason.
Amplitude damping is the obvious test case. Its quantum capacity and entanglement-assisted capacity have long had single-letter formulas, and its single-use Holevo information was worked out by Bennett, Shor, Smolin and Thapliyal (2002) and by Giovannetti and Fazio (2005). As the authors note, additivity results for entanglement-breaking channels and for unital qubit channels do not cover the nontrivial regime 0 < p < 1. Whether that known single-use rate is the true capacity has remained open. This paper says it is.
What the paper claims
The authors state three main results:
A class of additive qubit channels. For every qubit-to-qubit channel that admits a pure output (some pure input is mapped to a pure output state), the classical capacity equals the single-use Holevo information, C(N) = χ(N), and the parallel entanglement cost of simulating the channel equals the entanglement of formation of its normalized Choi state. Holevo information, channel entanglement of formation and parallel entanglement cost are all additive with any finite-dimensional partner channel. This class contains every qubit channel with at most two Kraus operators.
The classical capacity of amplitude damping. For damping probability p, C(Ap) equals the known one-parameter maximization over the input excitation probability q, and it is achieved by an equiprobable pair of pure states √(1−q*)|0⟩ ± √q*|1⟩ used as product codewords with collective decoding.
The entanglement cost of amplitude damping. Under parallel simulation by local operations and classical communication (LOCC), the cost is h2((1+√p)/2) ebits per channel use, where h2 is the binary entropy.
The authors are explicit that the single-use formula and optimal ensemble are not new. What is new is the matching many-use converse that turns them into capacities.
The key idea
Both results reduce to one inequality about entanglement of formation, EF. Fix a distinguished vector in each of two local systems A and B. If a state ρAB has no support on the sector where both systems are orthogonal to their distinguished vectors, the paper proves (Theorem 4.1) that ρ is strongly superadditive: for any global state Ω whose AB marginal is ρ, however it is correlated with auxiliary systems A′B′, EF(Ω) ≥ EF(ρ) + EF(σA′B′). In particular EF is additive on ρ and its entanglement cost equals EF(ρ).
Amplitude damping satisfies this constraint naturally. In its Stinespring dilation, an excitation either stays in the output or leaks into the environment, never both, so the joint output–environment state has no weight on |1⟩B|1⟩E. That single missing sector is what the whole argument runs on.
How it works
The proof of Theorem 4.1 has two ingredients:
A triangular block entropy inequality (Theorem 4.2). For a pure extension, the missing sector makes the coefficient matrix block-triangular. The authors show its entanglement entropy is at least a “scalar” term, fixed by the weights of the allowed sectors, plus weighted entropies of the individual blocks.
A decomposition that preserves two expectation values (Lemma 4.3). Any state can be written as a mixture of pure states within its support that each have the same expectation values of two chosen Hermitian operators. Choosing the two sector projectors bounds EF(ρ) by the scalar term, while local monotonicity and convexity bound EF(σ) by the block terms.
Notably, the argument never needs to evaluate either mixed-state convex roof; it only needs upper bounds on the two marginals.
For communication, the Matsumoto–Shimono–Winter correspondence converts this into strong superadditivity of the minimum output entropy at a fixed input, and hence strong additivity of Holevo information (Corollary 5.1). For simulation, the same theorem is applied to the Choi state: a qubit channel admits a pure output exactly when its Choi state has a product vector in its kernel, which makes the known lower and upper bounds on parallel entanglement cost coincide.
For amplitude damping specifically, the paper also derives an exact finite-block identity (Proposition 5.6) splitting the gap between nC(Ap) and the Holevo information of any n-use input into three nonnegative terms. It follows that for p < 1 the unique optimal average input at every block length is the n-fold product of the optimal single-use input.
Why it matters
At p = 1/2 the paper’s numerical evaluation gives a classical capacity of about 0.47173 bits per use, an entanglement cost of about 0.60088 ebits per use, a quantum capacity of zero, and an entanglement-assisted classical capacity of exactly 1. I reproduced the first two numbers independently from the paper’s formulas. In the high-damping limit p → 1, the authors derive that the entanglement-assisted capacity exceeds the unassisted one by a factor approaching 4, a ratio first obtained by Bennett et al. that the new theorem now attaches to the true capacity.
Beyond amplitude damping, the pure-output class gives a clean, checkable criterion for which qubit channels have additive classical capacity. The paper includes an explicit Kraus-rank-three example to show that the class is strictly larger than the Kraus-rank-two case.
Technical perspective (interpretation)
This section is my reading rather than the paper’s claims. What I find most useful here is that the argument is structural: a support condition, not a channel-specific calculation, drives both the capacity and the simulation results. That makes it a template. The authors themselves suggest looking for “excluded sectors” in other channels. The obvious next target is generalized amplitude damping at finite temperature, which the paper notes does not in general satisfy the hypothesis. A robust version of Theorem 4.1, with a penalty proportional to the weight in the excluded sector, would be a natural way to get approximate additivity there.
The paper is also a notable instance of AI-assisted mathematics. The authors disclose that frontier large language models were used for the proofs of Theorems 4.1 and 4.2, that they verified the statements and proofs themselves, and that a Lean 4 formalization accompanies the paper. If that formalization covers the key steps, it substantially raises confidence in an argument that has not yet been through peer review.
Limitations and open questions
Not yet peer reviewed. This is a v1 preprint. The claimed Lean formalization is referenced as a GitHub repository, but when I checked on 27 September 2026 that link returned “page not found”, so I could not inspect it. I have not verified the proofs independently.
AI-assisted proofs. Per the paper’s disclosure, language models were used for the two central proofs, and ChatGPT was used to help with the writing; the authors state that they take full responsibility for all mathematical claims.
Scope of the operational claims. The results concern vanishing-error classical communication and parallel LOCC simulation. Strong-converse thresholds, reliability functions and simulation against adaptive tests are left open.
Zero temperature only. The support condition is sufficient, not necessary. Generalized amplitude damping and most higher-dimensional channels are not covered.
Different resource models. The authors caution that comparisons between capacities and simulation costs use different free resources. That the quantum capacity vanishes for p ≥ 1/2 while the entanglement cost stays positive does not by itself imply irreversibility.
Paper information
Title: Classical Capacity and Entanglement Cost of the Amplitude Damping Channel
Authors: Ziao Tang, Ge Bai, Xin Wang (The Hong Kong University of Science and Technology (Guangzhou)); Chengkai Zhu (QudeLeap Research, Shanghai)
arXiv: 2609.28592v1 [quant-ph], submitted 23 September 2026, listed 25 September 2026
Length: 21 + 8 pages, 1 figure
Lean formalization (as cited in the paper): github.com/QuAIR/Classical-Capacity-Qubit-Amplitude-Damping (not publicly reachable as of 27 September 2026)
Primary sources
Z. Tang, C. Zhu, G. Bai, X. Wang, Classical Capacity and Entanglement Cost of the Amplitude Damping Channel, arXiv:2609.28592 (HTML full text)
C. H. Bennett, P. W. Shor, J. A. Smolin, A. V. Thapliyal, Entanglement-assisted capacity of a quantum channel and the reverse Shannon theorem, IEEE Trans. Inf. Theory 48, 2637 (2002)
V. Giovannetti, R. Fazio, Information-capacity description of spin-chain correlations, Phys. Rev. A 71, 032314 (2005)
K. Matsumoto, T. Shimono, A. Winter, Remarks on additivity of the Holevo channel capacity and of the entanglement of formation, Commun. Math. Phys. 246, 427 (2004)
M. B. Hastings, Superadditivity of communication capacity using entangled inputs, Nature Physics 5, 255 (2009)


