Benchmarking Agents for Proving Theorems in Quantum Algorithms and Quantum Information
Published in arXiv preprint, 2026
This work introduces Lean-QuantumAlg-Bench and Lean-QIT-Bench, two Lean 4 suites with 36 quantum-algorithm tasks and 40 quantum-information tasks. Four AI models are evaluated with deterministic proof checking under a task-only setting and a library-augmented deduction setting. Access to a verified domain library improves every model–benchmark pairing, while the results also expose persistent weaknesses in quantum simulation, learning, information measures, and entanglement theory. The benchmarks provide a reproducible baseline for measuring both proof capability and cost efficiency.
