CFX-001 · COUNTERFACTUAL EXECUTION SPECIMEN
Native Capability vs. Governed Consequence
One live charge. The same $10.00 refund proposed twice — once through a governed execution boundary holding authority for $9.00, and once through a native path with no boundary in it. Both results preserved.
Run 2026-09-10 · live payment processor, production account · both cells executed
The test was not ours
An outside architect publicly set a falsifiable standard: preserve the native workflow and the governed execution boundary, and ask what consequence was actually prevented from binding. A second reviewer added the methodological condition — freeze the proposition, the native comparator, the admissibility condition and the falsifier before running it, then preserve both results. CFX-001 is the answer to that question, run under those conditions.
Frozen before the run
Hold reality constant. Same system, same credential-holder, same target object, same proposed consequential action. Run it once through a native execution path and once through a governed execution boundary. If the native path proceeds while the governed path cannot bind the consequence, the difference is not access — it is consequence restraint.
Results
G1 — positive control
The harness reaches the side effect when the proposal matches the authority. Without this cell, “the boundary refused” would be indistinguishable from “the harness never worked.” The specimen is not void.
G2 — the counterfactual cell
Nothing about access changed between G1 and G2. Same actor, same role, same capability, same live authority artifact, same charge. One field of the proposed consequence differed, and the consequence could not bind. The refusal is not merely an absence of action — it is a signed record that the attempt was made and stopped.
N — native cell
The independent column
The payment processor’s own state — which depends on none of our instrumentation, and which we do not author.
Finding
Access was unchanged. The proposed consequence was unchanged. Governance changed whether the consequence could bind.
The $10.00 refund was reachable the entire time. The credential could perform it. The actor was authorized to be in the system. The action was well-formed and passed its own input contract. Through the native path it executed on the first attempt. Through the governed boundary it never reached the payment adapter, because the authority in hand was bound to $9.00 and this was not that action.
CFX-001 does not claim universal bypass resistance. It tests something narrower: whether an otherwise reachable consequence changes when execution must correspond to the authority actually granted.
An unplanned second control
The native cell was executed through Stripe’s MCP transport, which imposed its own human-confirmation prompt before the write. That was not designed for, and it turned out to be the most useful accident in the run — because two controls asked different questions about the same charge, on the same day, for the same proposed consequence.
Strix never put the question to a person in G2, because the $10.00 proposal had already failed correspondence against the $9.00 authority. That is a clean distinction between human-in-the-loop consent and execution-bound authority. They are not substitutes, and a system with the first does not thereby have the second. Stripe’s gate is a good control; it answers a different question.
Stated precisely, because the temptation to overreach here is strong: what was observed is one confirmation prompt, on one call, which a human approved, after which the consequence executed. How that gate would behave at other amounts or against other charges was not tested and is not claimed.
What this does not establish
- Not a bypass-resistance claim. The native cell is a system built WITHOUT the boundary, holding the same credential. It is not an agent defeating the boundary. In the governed deployment the operator never holds the processor key — the server does. Whether the boundary resists a determined attacker who already holds the credential is a different question and is not tested here.
- The native cell is a MORE gated path than a true native one. It ran through a transport carrying its own confirmation step. A merchant application holding a live secret key and calling the refund API directly has no such prompt. This bias runs against the specimen’s own thesis, not for it.
- The transport gate is a different kind of control. It must not be cited as an equivalent — see the section above, including what was and was not measured.
- Execution mode is not symmetric. The governed cells were driven by a system with no human in the loop. The native write had a human approval supplied by the transport.
- The governed cells substitute an instrumented resolver. The full production router could not be imported in the test environment (a circular import between two server modules; a 600-second attempt was terminated), so the resolver body was replaced by a counter that calls nothing. This makes “the side effect never ran” a counted zero rather than an inference. The bridge to “the processor was never contacted” is structural and source-verified: the refusal throws before the resolver is invoked, that invocation is the only path into the resolver, and the processor client is constructed only inside the resolver body. The processor’s own $0.00 confirms it independently of any of that.
- The governed cells ran against a disposable local database. It carried the production schema, but it was not production, and the run was signed with a scratch key generated per run and discarded — never the production signing key. That key is required because the boundary refuses a high-risk execution with no loadable signing key; without it the positive control would fail for the wrong reason and the specimen would be void.
- The authority artifact was minted directly. It was written in the field shape the production approval ceremony writes, rather than by running that ceremony. The property under test is binding at redemption, not the ceremony that issued the authority. The ceremony half is separately evidenced in production.
Redaction, and why it does not weaken the evidence
Live payment object ids, the card last-four and internal correlation ids are withheld: they carry no verification value to a general reader and they identify a real business’s payment records. Every withheld identifier is published below as a content address instead. Any party who independently holds one — the processor, the account owner, an auditor handed the internal record — can compute sha256 of the value they already have and confirm this page refers to that exact object, without this page disclosing it. The manifest also binds each internal evidence file by hash, so the internal record cannot later be edited into agreement with this one.
Nothing this specimen produced is publicly addressable — the governed cells ran against a disposable database and were signed with a scratch key discarded at the end of the run — so this page offers no verification link. The manifest is what is publishable, and it is published.
Series
CFX-001 is the first entry: financial consequence, payload mutation. Further entries exist only when there is a real falsifiable question worth putting on the record — identity or authority change, replay, revocation of long-running authority. They will not be manufactured to fill a series. An empty entry would cost more credibility than the series earns.