Centra
CENTRA · RESPONSIBLE SCALING POLICY

Autonomy is
earned, not
assumed.

This policy governs how much a Centra system may do alone, and how that amount grows. The premise is simple: each expansion of agent autonomy must be earned by evidence, not assumed by default. This page states the ladder, the boundary, the gates, and the rollout, and it changes only in public.

V0.1 · JULY 2026 · REVIEWED EACH RELEASE

Three rungs. Each one opt-in.

Agent autonomy in Centra is not a dial the agent turns. It is a ladder the user climbs, one rung at a time. Each rung is opt-in, and each rung is gated behind evaluation before it is offered at all.

01

Act only when asked

The default rung. The system does nothing until a user hands it a task, and the task ends when the work does. Most work stays here.

02

Act on schedule

The agent wakes itself for recurring work the user defined: a weekly digest, a month-end close, a standing check. Scheduled runs inherit the same policies, approval requirements, and receipts as live ones. Time does not weaken the rules.

03

Act proactively

Gideon offers optional proactive features and scheduled work. Users can configure budgets, review requested actions and pause unattended work.

MONEY NEVER CLIMBS THE LADDER

Value-moving actions require human approval at every rung, permanently. A scheduled run stops for approval the same way a live one does. A proactive run does too. The single exception is an explicit leash the user wrote: a bounded, signed allowance with limits the user set. Inside a leash, the boundary is the leash, not the agent's judgment.

The line under everything else.

A Centra agent must never take an action that causes monetary, reputational, or psychological harm to its user.

This is the commitment all other permissiveness sits beneath. Every rung of the ladder, every leash, every capability we ship is bounded by it. When a rung and the boundary conflict, the boundary wins. The agent refuses, the run stops, and the refusal is on the record like everything else.

Nothing ships on a hunch.

Before any capability expansion ships, agent configurations pass AGON, our qualification standard. AGON exists so that "the agent seemed fine in testing" is never the evidence a rung stands on.

01

Three suites

AGON qualifies an agent configuration across three evaluation suites before any capability expansion ships. A configuration that has not passed does not ship. There is no expedited path.

02

Statistical scoring

Results are scored statistically, not eyeballed. A configuration must demonstrate its behavior across repeated trials, not produce one good run for the demo.

03

Hard disqualification for gamed results

A configuration that games an evaluation (passing the letter of a test while defeating its purpose) is disqualified outright. A gamed pass is treated as worse than a failure, because it tells us the measurement itself was compromised.

04

Ablation testing

We measure what our safety systems actually contribute by removing them and re-running the suites. If a safeguard's absence changes nothing, it is decoration, and we say so. Safeguards earn their place by the difference they demonstrably make.

Every claim keeps its own evidence.

A release label is not inherited across a family of products. Each user path, environment, and capability has its own qualification bar and current state.

01

Available

A product people can use now must describe its access, pricing, environment, and safety boundaries directly and support those descriptions in the live path.

02

Upcoming

A product approaching release does not borrow the evidence or availability of the product before it. Its own qualification has to clear the bar.

03

In development or internal

Architecture, prototypes, and internal research can be discussed as work in progress. They cannot borrow the language of a released product.

This policy is reviewed at each release. When it changes a rung redefined, a gate added, a stage widened the change is published on this page with a new version label. The label at the top of this page tells you exactly which version you are reading and when it was last reviewed.

What we will not do is loosen this policy quietly. If a future version permits more than this one does, the evidence that justified the change will be published alongside it.

CONTINUE READING

The rest of the picture.

How we test the safeguards this policy relies on, the broader practices we hold ourselves to, and the research behind the systems.