top of page
THE CORE OBSERVATION

The Method Is Published. The Calibration Is Ours.

Trust comes from publication. Defensibility comes from live calibration.

Our Scientific Foundation: Chaos-Order Boundary Theory

IS(x) is a single instability score for an AI workflow, tracked step by step rather than output by output. It is a weighted composite of four scored dimensions: correctness (0.33), settlement (0.26), liability (0.23) and clause precision (0.18). Weights are derived by PCA, so they are inspectable and explainable to a regulator, and re-fitted with each model generation.

The Four Observables
  • Volatility: Early-onset variance in response latency across inference nodes.
  • Autocorrelation: Detection of mode collapse via output embedding correlation.
  • Entropy: Loss of predictability and broadening of probability distributions.
  • Drift: Directional acceleration of the system centroid toward instability.

Empirical Validation and Research Phases

What the benchmark shows

Calibrated across 2,700+ sequential scenarios on Claude, Gemini and OpenAI via native APIs. Degradation is detectable up to 48 steps before visible failure in benchmark conditions. 48 is the upper bound across the benchmark, not an average. Production distributions are a pilot deliverable. Empirically derived as the pre-threshold point that preserves a usable intervention window. Calibrated, and re-fitted per model generation.

Phase 2 — Active

Mathematical formalisation for peer review. Academic collaboration in progress. Insurance vertical pilot deployment. SDK and dashboard MVP build.

Phase 3 — Research

Agent-to-agent health signal propagation protocol (AgentHealthSignal). System-level B(x) as collective stochastic property. Pipeline instability as emergent phenomenon across multi-agent topology.

Design principles

The scored system never sees its own score. The benchmark re-runs with each model generation. The methodology stays public and auditable.

Published

Methodology and benchmark deposited at Zenodo, DOI 10.5281/zenodo.20708678.

Development Status

Complete
Active
Next

Framework formalised, benchmark published, five node agentic insurance workflow built and validated against reactive detection baselines. Live demo running on real inference.

Peer review, SDK and dashboard build, insurance pilot conversations with risk functions.

Agent to agent health signal propagation, and system level instability across multi agent topologies.

bottom of page