Clarus and AGI Safety: A Structural Layer for Frontier Systems
Behaviour shows performance. Structure shows stability. Only one keeps frontier systems safe.
Frontier AI systems can produce fluent, confident output even as their underlying reasoning substrate destabilises. Internal metrics — loss curves, activations, gradients, attention statistics — do not measure behavioural stability under load, leaving a structural blind spot where drift accumulates silently.
This paper introduces K, a behavioural invariant computed entirely from external outputs. K quantifies how stable a model’s behavioural geometry remains under perturbation, reframing, and long‑range compositional reasoning. Two companion metrics complete the framework: KEC, the load‑response curve that reveals brittleness and reserve capacity, and KSP, the propagation map that shows whether local instability remains contained or spreads system‑wide.
Together, K, KEC, and KSP form an architecture‑agnostic, falsifiable stability layer that detects coherence loss earlier, measures it more precisely, and makes it actionable before failures appear at the surface. This positions behavioural coherence as the structural prerequisite for alignment, oversight, interpretability, and all higher‑order safety mechanisms to function reliably at frontier scale.
Index
1. Abstract
Why behavioural stability must be measured externally, and why internal metrics fail to detect drift.
2. What Coherence Means in Clarus
Coherence as a geometric property of behaviour:
predictable perturbation response
recovery stability
temporal consistency
trajectory integrity
3. The Operational Form of K
K computed solely from external behaviour:
perturbation response
sequential divergence
recovery dynamics
symmetry under reframing
behavioural curvature
K = f(divergence, recovery, symmetry, curvature)
4. Why Coherence Is a Safety Primitive
Coherence as the structural substrate required for:
alignment
rule-following
long-range reasoning
compositional reliability
self-evaluation
stable oversight
5. KEC and KSP — The Diagnostic Layers
KEC: load‑response curve
reserve capacity
yield point
brittleness slope
strain
KSP: propagation map
localization
propagation pathways
systemic risk
fault clustering
6. A Complete Stability Architecture
K → Is the system stable?
KEC → Where does it break?
KSP → Will the break stay local?
7. Case Study: Long‑Context Narrative Drift
Constraint failure, reasoning drift, and end‑sequence collapse — invisible to internal metrics, visible to K, KEC, and KSP.
8. Interventions Enabled by K
Tactical interventions (low K, localized KSP)
Strategic interventions (KEC yield, rising KSP)
Systemic safeguards (high KSP, cascading instability)
9. Scaling Without K Is Scaling Blind
Why capability growth amplifies hidden instability and why behavioural invariants are required for safe scaling.
10. Why Clarus Is Foundational for Frontier Safety
Coherence visibility as the structural prerequisite for alignment, oversight, interpretability, and autonomous toolchains.
11. Conclusion — The Necessity of Invariant‑First Design
K, KEC, and KSP as the stability triad for reliable frontier systems.
Appendix A — Behavioural Geometry & Internal State
Why external behaviour necessarily reflects internal stability.
Appendix B — Mathematical Footprints of Instability
Perturbation sensitivity, recovery dynamics, curvature, temporal consistency.
Appendix C — External Stability Measures Across Engineering Domains
Control theory, robotics, circuits, ecology — stability inferred from behaviour.
Appendix D — Falsifiability & Empirical Validation
Why K is a scientific invariant, not a philosophical claim.
Appendix E — The Illusion of Transparency
Why internal access does not guarantee understanding.
Appendix F — Mathematical Inevitability of External Instability
Why internal instability must express externally.
Appendix G — Deliberate Scope of K
Why K measures stability, not mechanisms.
Download PDF
Download PDF
