
Nvidia benchmark finds average model accuracy falls 62.8% on 128,000-token tasks
AI agents · Friday, 2 October 2026
Why it matters
For an agent founder, the results point to pipeline and state-management choices—not just model selection—as major reliability levers: long context, unstable formats, missing IDs, and complex steps can sharply reduce exact task completion. The benchmark also provides concrete failure dimensions to test when evaluating long-running agent workflows.
What happened
Nvidia researchers evaluated seven open-weight models on the Long-Transduction benchmark and found average accuracy fell 62.8% when task context grew from 4,000 to 128,000 tokens. The benchmark used 1,440 documents per model, greedy sampling, and exact-match scoring; DeepSeek declined from 0.909 to 0.554, while Nemotron Super fell from 0.711 to 0.138. Varying input formats reduced performance by 36.5% on average, removing stable item identifiers caused relative drops of up to 64.3%, and greater local step complexity produced a 39.9% average decline. The researchers recommend decomposing long tasks into smaller subtasks and assigning each item a structured identifier.
Previously
30 Sep 2026 — Nvidia researchers released the paper “Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability,” introducing the Long-Transduction benchmark for evaluating long-task state tracking.(Nvidia via arXiv)
Players & places
- Nvidia