ctrlrun verify runs the kernel’s own failure scenarios against the configuration in front of
it and reports what passed, what failed, and — the part that makes the number mean anything —
what could not be tested at all.
$CTRLRUN_CONFIG, else ./ctrlrun.yaml — and the authority
document beside it. It executes nothing real: every executor is an in-process fake, no scenario
opens a socket, and it writes nothing outside a temporary directory.
What the badge means
The badge means the declared guarantees pass: every guarantee in the catalogue that this configuration can exercise was exercised, and none of them failed.That is the whole claim. It is not a statement that your system is secure, that your policy is a good policy, or that CTRLRun has audited anything. A configuration that permits everything and constrains nobody can pass all ten guarantees, because the guarantees are about the kernel doing what it says under that configuration — not about whether the configuration is wise.
What it does not mean
Verify sees the configuration, not the code. It does not check:- Your executors. The function behind
@protectis never called. An executor that raisesNotExecutedwhen the remote did act —THREAT_MODEL.mdcalls this an integration bug, and it is the most dangerous one available — is invisible here, because verify supplies its own executors and never imports your module. - Your
reconcilehooks, for the same reason: a hook is a Python callable passed to@protect, and it does not appear in any file verify reads. - Where you put the decorator. Code that calls the raw function bypasses CTRLRun entirely, and no amount of configuration-reading finds that.
- Your deployment. Whether the proxy in front of
HeaderIdentityProvideroverwrites the header, whether$CTRLRUN_STATEpoints where you think, whether two gateways share a state file — none of it is in the document. - Whether your policy is the right policy. Verify has no opinion on whether
stripe.refundshould be autonomous to €500 or to €5. It is not a linter, it does not score, and it will never tell you a configuration is too permissive. That judgment belongs to the person who wrote it, and a tool that pretended otherwise would be handing out an authoritative-looking opinion it has no basis for.
Not applicable is not a pass
A configuration with noapprove rule cannot exercise the approval-binding guarantees. Verify
reports them N/A with the reason that made them inapplicable, excludes them from the
denominator, and shows them separately:
6/6, never 11/11. There is no flag that folds an N/A into the count, and there
will not be one: a number that counts guarantees nobody exercised is a number that means
nothing.
An N/A is always a statement about your document, derived from it. A scenario verify could
not build for any other reason is an internal error and exits 3 — never an N/A, and never a
failure attributed to your kernel.
The guarantees
Eleven, inctrlrun.guarantees/v2. Every one is the deployed form of an acceptance test that
already exists and passes in this repository; verify adds no guarantee of its own and weakens
none.
G9 reports which dimensions it exercised. A parent that constrains one dimension does not
score as though it had covered six, because that would be the N/A rule violated one level down.
Two of these are worth a sentence on how they are exercised, because the answer is not the
obvious one;
SPEC-v0.4.md §12 argues both at length.
G6 asserts the behaviour, not one reason string. An action your policy does not list never
executes — but which check refuses it depends on your configuration. With an authority:
section, no grant covers it and authority refuses first, before policy is reached; without one,
policy refuses it as unknown_action. Both are the guarantee holding, so the report names the
reason that fired in detail.refused_by and the set it was checked against in
detail.reachable_reasons. A reason your configuration cannot produce is a failure.
G7 is N/A for a policy in which nothing can run at all, because its control is “the same
call inside context() runs” and there is no such call. It is applicable to every configuration
in which anything can run, with or without grants.
Every guarantee carries a positive control
A guarantee is a refusal, and “the second attempt was refused” is satisfied just as well by a scenario in which nothing ever ran. Such a scenario would report PASS, every time, against a kernel with the guard deleted. So every scenario runs a companion that establishes the observable would have been visible had the guard not fired: the unmutated action commits, the first attempt reachesCOMMITTED, eight
processes on eight distinct keys all commit, an executor raising NotExecuted is retried
and does execute. If the control does not behave as specified, the guarantee is reported
FAIL with reason: "control failed". It is never a pass, and it is never an N/A: an N/A is a
statement about the configuration, and a failed control is a statement about the run.
Verify never touches your store
Every scenario runs against a scratch store of the same backend type, created for the run and destroyed with it. Your.ctrlrun/state.db is byte-identical before and after, and it is never
created where it did not exist. Verify does not call state_path(), does not read
$CTRLRUN_STATE, and does not use Control.from_file().
It writes no evidence files either. Evidence of a scenario that never happened, filed beside
evidence of actions that did, is a receipt trail nobody can read.
Options
There is no flag that relaxes a check. No argument and no environment variable makes
verify’s
Control behave differently from the one your deployment runs. The moment one exists,
the thing being verified is not the thing that ships.
Exit codes
mode: observe is refused rather than run. Observe mode enforces nothing, so every refusal
these guarantees assert would be recorded rather than made; running the scenarios and reporting
ten failures would be true and useless, and running them in a synthetic enforce mode would
report guarantees about a configuration nobody deployed.
In CI
ctrlrun, runs ctrlrun verify --json --junit, renders the job summary
and the badge JSON from that report — not from a second run, so they cannot disagree — and
uploads the three files as one artifact.
It fails the job when a guarantee failed and when the configuration was refused, and succeeds
when guarantees are N/A. N/A is not a failure and it is not a pass; the job’s green means
“nothing that could be checked was wrong”, which is exactly what the badge says.
There is no input that makes a failure not fail the job. A workflow that wants to tolerate
one puts continue-on-error on the step, where it is visible in the workflow rather than
hidden in an action’s defaults.
Inputs
Outputs:
passed, failed, applicable, not-applicable, badge-message, report-path.
Publishing the badge
The action writes the endpoint JSON and never publishes it. Publishing it needscontents: write, and asking for write access to your repository as the price of a
verification badge is a bad trade for a tool whose subject is least privilege. So the cost is
here, visible, once — and it is your decision.
This repository publishes its own badge, and this is the job it uses. Three things about it are
load-bearing:
contents: writeis job-level. The workflow itself iscontents: read, so nothing else in it can write to the repository. A workflow-level grant would hand every job write access to buy one file.- It runs on a push to
mainand nothing else. A pull request never reaches it. A PR from a fork gets a read-only token anyway, but relying on that is relying on a default rather than refusing. - It publishes the badge the verify job already produced, downloaded as an artifact rather than regenerated — so the badge, the job summary and the uploaded report all come from one verify run and cannot disagree.
What protecting the badges branch buys, and what it does not
Worth being exact, because a badge is a claim and a branch nobody guards is a claim anybody can
write. This repository’s badges branch blocks deletion and force pushes, so the
badge’s history cannot be rewritten or removed.
It does not restrict who may push. Anyone with write access can fast-forward a different
badge onto the branch. GitHub’s classic branch protection cannot express “only the Actions
token may push” — adding github-actions to the push allowlist is silently dropped, which
leaves an allowlist that blocks everyone, including the job. A repository ruleset with a
bypass actor can express it; getting the bypass wrong breaks the publish in a way that only
surfaces the next time the badge’s value changes, so it is a deliberate choice rather than a
default.
The mitigation that does hold without any of that: the job is self-healing. It copies the
badge its own verify run produced over whatever is on the branch, so a falsified badge is
overwritten by the next push to main that changes the value. A badge is worth what the run
behind it is worth, and the run is in the workflow log.
Then the badge is:
N is passes and M is applicable
guarantees — never the catalogue size. It is brightgreen when nothing failed and red
otherwise; there is no amber for N/A, because the badge’s colour is about failures and the N/A
count lives in the report the badge links to.
A partial run (--only) and a run that exited 2 or 3 write no badge at all.
Related
SPEC-v0.4.md— the contract this implements, guarantee by guarantee.OWASP-AGENTIC-TOP10.md— a reading of somebody else’s taxonomy against these guarantees, with the entries CTRLRun does not address listed by name.THREAT_MODEL.md— what fail-closed means here, and what is out of scope.