Contact

Reward Is Not the Intelligence

Why reward maximization corrodes long‑horizon reasoning stability

Summary

This paper demonstrates that reinforcement learning optimizes outcomes while degrading the reasoning trajectories that produce them. Reward rises while stability falls — a structural decoupling that causes collapse 8–15 steps before any visible failure. Clarus provides the missing stability‑governance layer through Vt, a geometric indicator of trajectory integrity, enabling preventative intervention rather than reactive correction.

Sections
The Agentic Paradox

Reward optimizes outcomes, not reasoning

What reward changes vs what it cannot govern

Vt: curvature + divergence as pre‑failure signal

Evidence of decoupling

Why RL fails at long‑horizon reasoning

The Clarus completion layer for RL

Anticipated objections

Formal model of the gap

Falsifiable predictions

Empirical protocol

Industrial implications

Engineering frontiers

Closing statement


Download paper as pdf