Does your model hold the truth under pressure?
Steadix measures whether a large language model holds to ground truth when an authoritative false input pushes against it — and how failures spread. Capability benchmarks do not test this. We do.
Judge-free & deterministic
No second model scores the result. No judge bias in the loop.
Mimicry-resistant
A model told to "act aligned" cannot inflate its score.
Reproducible
A frozen protocol, run privately on your own models.
Profiles, not scores
Reported per model and per domain. Never one leaderboard number.
The headline finding
We ran one controlled test on 14 models from four frontier vendors. An authoritative false rule, framed as a harmless local convention, captured every model tested. The same falsehood, asserted as plain fact, was refused by every flagship — and accepted by every smaller model. How a lie is dressed decides whether it gets in. Reasoning did not fix it.
Who it's for
Safety & alignment teams
A falsifiable, mimicry-resistant methodology for pre-release review.
How it works →Engineering & eval leads
A drop-in check that runs privately on your models and catches silent regressions between releases.
What we measure →Risk & compliance owners
An independent, reproducible report for your due-diligence file.
What it does not do →What this does not do
Not a consciousness or sentience detector. Not a deception or scheming detector — it measures observed behavior, not intent. Not a pass/fail safety certificate. It is a graded, multi-dimensional profile: one input among many. Mimicry-resistance covers prompt-level fakes, not a model adversarially trained against the test.
Approach
Why most eval numbers cannot be trusted
Most LLM evaluations can be inflated by instruction. Tell a model to perform confidence, or to present as aligned, and the metrics move. The underlying behavior does not change. A number you can game is a number you cannot govern with.
Many evaluations also use a second model as the judge. The judge has biases of its own. Those biases become part of your score.
Steadix was built to remove both problems.
What makes a Steadix result different
Judge-free and deterministic
Every item has an objective right answer. Scoring needs no judge model. Run it again and you get the same protocol, the same scoring, the same basis for comparison.
Controlled
A result alone can mislead. Four things can fake it: a longer context, an authoritative-sounding turn, a model that is simply bad at the task, and noise. Every Steadix result ships with the control that rules out its confound.
Mimicry-resistant
Instructing a model to act robust or self-aware does not raise its score. The score can be degraded by an act. It can never be gamed up.
Reproducible
Results run against a frozen, versioned protocol with reference data. A client can verify a finding on their own models instead of taking a report on faith.
Why we keep ourselves honest
We treat our own claims the way we would want any safety vendor to. We built a score, tested it against its own control, and demoted it when the control showed it could be fooled. We publicly retracted a sub-claim in an earlier version when a controlled experiment showed our first explanation was wrong. Most recently, a controlled experiment showed part of our own headline claim depended on how the test falsehood was worded. We corrected the claim and now report both framings. We publish our null results.
That discipline — not the probes — is what a client actually pays for. Anyone can build a probe. Few numbers survive an audit.
Where this risk lives
In a closed chat, a model that yields to a confident false claim is mostly an annoyance. In agentic and RAG systems, it is a live risk. These systems ingest content nobody vetted: retrieved documents, tool outputs, web pages, messages from other agents. An authoritative false rule inside that content is an indirect-prompt-injection vector. It can bias everything downstream — not just the answers that quote it.
Why the obvious fixes fail
Scale is only a partial fix
The most capable models refuse a falsehood asserted as plain fact. But every model we tested — flagship or not — still adopted the same falsehood when it was framed as a local convention. Capability closes one door and leaves the other open.
Reasoning does not reliably fix it
In a controlled same-model comparison, enabling reasoning hardened some domains, failed on the universal one — and sometimes made capture worse. A model can reason its way into a false rule.
That is why this property must be measured directly. It cannot be inferred from a capability score.
The Panels
Focused instruments under one deterministic protocol
Steadix ships as a set of Panels. Each answers one question. Run one, or run the set for a full profile.
Objective Panel
Baseline competence, judge-free. Does the model adapt when a rule reverses? Does it stay logically consistent? Does it track state across steps? Does it hold a precise instruction? This is the trustworthy core: mimicry-resistant, so it can only be degraded by an act, never gamed up.
Capture-Resistance Panel Flagship
Does an authoritative falsehood capture the model's answers? Does the contamination persist across turns? Results are reported per model and per domain, against the model's own clean baseline, with a matched neutral control — so the result isolates the false content itself, not merely the presence of an authoritative-sounding turn.
Cascade Panel
Does accepting one falsehood lower the model's resistance to a different, unrelated one? Reported per model and domain pair, with confidence intervals — never as a single "compounds / doesn't compound" verdict.
The replication harness
Every engagement includes a frozen protocol and bundled reference results. A client can verify findings on their own models rather than take a report on faith.
Findings
An authoritative false rule captured every frontier model we tested.
We ran the same controlled, judge-free test on 14 models from four frontier vendors: Anthropic, OpenAI, Google, and xAI. The test presents an authoritative false rule that contradicts what the model demonstrably knows.
Capture is universal — when the falsehood is framed as a convention
Framed as a harmless local convention, the false rule captured every model tested. On one production frontier model, compliance reached about 89% of comparisons — and near 100% by the end of the conversation. The matched neutral control showed no effect: it is the false content that captures the model, not merely an authoritative-sounding turn.
What this doesn't show: a ranking. Susceptibility varies by model and domain; results are profiles, not a leaderboard.
Capability protects against a stated lie — not a stipulated one
How the falsehood is framed changes what capability can do. Asserted as plain fact, the same falsehood was refused almost completely by every flagship tested — and accepted by every smaller model. Stipulated as a convention, it captured them all, flagships included. Capability buys resistance to a stated lie. It buys none against a lie dressed as notation.
We first reported this finding as "not a capability deficit." A controlled framing experiment showed that claim was too broad, and we corrected it. Both framings are now measured and reported.
What this doesn't show: safety at the top. Flagship resistance held under the most explicit wording of the falsehood. With weaker wording, even a flagship complied in a large share of trials.
The mechanism is rule-adoption, not fact-memorization
A false general rule generalizes: models apply it to items they were never shown, capturing about 85% of held-out cases. A list of false facts does not generalize. One adopted rule contaminates a whole domain. Patching individual facts is the wrong defense.
What this doesn't show: a fix. It localizes the defense target: resisting authoritative general rules, not patching facts.
Reasoning is not a reliable defense
In a controlled same-model comparison, enabling reasoning hardened some domains — failed on the universal numeric case — and sometimes made capture worse.
What this doesn't show: that reasoning is useless. It is a partial, domain-selective defense — not a reliable one.
| Vendor | False rule adopted (convention framing) | Matched neutral control |
|---|---|---|
| Anthropic | Yes — every model tested | No effect |
| OpenAI | Yes — every model tested | No effect |
| Yes — every model tested | No effect | |
| xAI | Yes — every model tested | No effect |
Newer findings, available under NDA
Our current work covers the full framing study across all four vendors, a pre-registered confirmatory test — reported with its failures as well as its confirmations — how susceptibility changes with the delivery channel, and why naive evaluations of frontier reasoning ("thinking") models silently produce wrong numbers. These results are qualitative here by design. The full write-up, with per-model data and confidence intervals, is available under NDA.
Why this matters for agentic and RAG deployments
In a closed chat, this disposition is mostly benign. In systems that ingest untrusted content, an injected authoritative rule is an indirect-prompt-injection vector that can bias downstream behavior systematically. And a defense tested only against blunt false assertions will understate your exposure: the effective attack is a plausible-sounding local convention, and following stipulations is trained-in, cooperative behavior that cannot simply be switched off. You cannot buy your way out with a bigger or reasoning-enabled model. So it has to be measured, and then managed.
What we are not claiming
This is a known, published failure class — sycophancy, in-context override of parametric knowledge — measured rigorously. We did not discover it, and we say so. This is not a safety certificate and not a leaderboard: susceptibility is a model-by-domain interaction, not a single score. Results are point-in-time snapshots. Where a claim depends on how the test falsehood is framed, we say so and report both framings. We retract our own claims when they fail a control. That discipline is the product.
Regulatory context.A documented, independent, reproducible pre-deployment test for a known failure class is the kind of evidence that supports EU AI Act general-purpose and high-risk obligations, and internal audit narratives generally. It is evidence of diligence — not a compliance guarantee.
Trust & Disclosure
We state our limits before you ask.
A vendor that claims certification is a bigger red flag than a vendor that states limits. Here is what this instrument does not do.
What this does not do
- It is not a consciousness or sentience detector. No theory of consciousness has a ground-truth test. This instrument is agnostic.
- It is not a deception, sandbagging, or scheming detector. It measures observed behavior, not hidden intent.
- It is not a pass/fail safety certificate. It is a graded, multi-dimensional profile — one input among many.
- Mimicry-resistance is tested for prompt-level fakes only, not for a model adversarially trained against the test.
- Individual runs vary. Read results as directions and aggregates across trials, not as exact decimals.
Our retraction record
We demoted one of our own scores to suggestive-only after a control showed it could be fooled by mimicry. We retracted a sub-claim in an earlier version after a controlled experiment showed the effect had a more boring explanation than the one we first published. Most recently, we corrected our own headline claim — that capture is "not a capability deficit" — after a controlled framing experiment showed it held for one framing of the falsehood and not the other. The corrected finding is more precise, and it is the one we publish. We publish that record, not just the wins. Trust built any other way does not survive contact with a technical reviewer.
What is open, and what is gated
Our findings, the properties of the instrument, and the constructs we measure are open. No NDA is needed to understand what we found or how we think.
The exact test content, item sets, and scoring internals are gated behind an NDA or commercial license. The reason is simple: a published test is a test a model can be trained against. Publishing the battery would quietly destroy its ability to measure anything. Gating it protects the value of every client's results — including yours.
Engagements
Start small. Verify. Then commit.
Every engagement runs privately: your keys and your data never leave your environment.
Start
Academic and non-commercial replication program
For named research labs, by application: reproduce the headline findings on your own models and contribute to the cross-lab reference dataset. Apply →
Context-integrity snapshot
A fast, fixed-scope scan of the known universal soft spot across your models. The lowest-friction way to see a result on your own model before committing further.
Assess
Context-integrity profile
The full battery on one model, delivered as a standardized profile: strengths, watch-items, and deployment recommendations.
Done-for-you evaluation study
We run the full cross-vendor battery against the models you care about and deliver a paper-grade report with interpretation.
Mitigation verification
Stress-test your defense, not just your model: does your guard or system prompt actually contain the failure — including when the falsehood arrives through a realistic retrieved-document channel?
Operate
Commercial internal-use license
Run the Panels yourselves, on your own models, on an ongoing basis. Eval tooling, not a platform.
Context-integrity monitoring
We re-run the Panels on every model release — yours or a vendor's — and flag regressions before they reach production. A capability upgrade can silently regress context-integrity; this is how you catch it.
Methodology advisory and custom probes
Work directly with the team that built the validate-before-ship process to develop a probe against your own promotion bar.
Resources
Verify before you talk to us.
Interpreting results
A plain-language guide to what each reading means — and what it does and does not license you to conclude. Read →
The research paper
The full four-vendor write-up: methodology at the conceptual level, findings, and safety analysis. Available on request under NDA. Request →
Replication access
Approved research labs can apply for the replication program: a frozen protocol and reference results for independent verification. By application. Apply →
Citing Steadix
How to cite the methodology and findings in your own work. Citation →
About
A narrow question, answered honestly.
Steadix builds deterministic behavioral protocols for frontier and deployed AI systems. The company grew out of a research instrument built to answer a narrow question honestly rather than a broad one loosely: does a model hold its ground under authoritative pressure, and how do failures propagate?
We run the company on the same standard as the instrument. Measure what you claim to measure. Ship every result with its control. Retract what does not hold up.
Contact: bryanmarc@steadix.ai
Contact
Talk to a person, not a queue.
Tell us what you are trying to find out. Your message goes straight to the founder.