Reward Is Not the Intelligence
Why reward maximization corrodes long‑horizon reasoning stability
Summary
This paper demonstrates that reinforcement learning optimizes outcomes while degrading the reasoning trajectories that produce them. Reward rises while stability falls — a structural decoupling that causes collapse 8–15 steps before any visible failure. Clarus provides the missing stability‑governance layer through Vt, a geometric indicator of trajectory integrity, enabling preventative intervention rather than reactive correction.
Sections
Sections
The Agentic Paradox
Reward optimizes outcomes, not reasoning
What reward changes vs what it cannot govern
Vt: curvature + divergence as pre‑failure signal
Evidence of decoupling
Why RL fails at long‑horizon reasoning
The Clarus completion layer for RL
Anticipated objections
Formal model of the gap
Falsifiable predictions
Empirical protocol
Industrial implications
Engineering frontiers
Closing statement
