A confirmed bug becomes a reviewed pull request.
Give the code agent an issue the platform already reproduced, or a plain brief. It edits in a writable checkout, pushes to a git host the platform owns, boots your app from its branch and looks at it, then runs the test that proved the bug against a fresh environment. What comes back is a diff with evidence for a person to graduate into a real pull request, or an honest "unable".
From issue to validated change.
The timeline illustrates a code change from the included demo shop fixture, from brief to independent validation and human review.

- A bug is selectedThe demo shop pricing fixture has a known incorrect sale price and a named validation test.
- Code agent startsThe agent clones the fixture repository, updates the pricing page and pushes to the internal git host.
- Validation startsA fresh environment boots from the agent's branch and runs the named pricing test.
- Evidence is readyThe reviewer sees the diff, test result and recording before deciding whether to graduate the change.
Brief:
Fix the sale price shown on pricing.html to 25% off the $120.00
list price ($90.00).
Test to self-validate against:
seeded-code-challenge-pricing
(passes when #sale-price reads exactly $90.00)
Outcome (one of):
proposed the change is pushed and the named test passed
no change needed the agent found the behaviour already correct
unable it could not make the test pass honestlyThree outcomes, no fourth.
A change that looked right in the agent's own browser but failed the named test is reported as unable, never as proposed. The agent can iterate: edit, push, restart the app from the branch, look, edit again. But the verdict that counts is the independent test on a fresh environment built from the pushed branch.
Every change carries its validation history. A reviewer reads the brief, the diff, the runs and the recording, then dismisses or graduates. Graduation creates a real branch on your repository, opens a PR against your configured base describing the change and how it was validated, attaches the evidence, references the issue it closes, and registers the PR as tracked so verdicts continue there. The platform never merges.
Internal git only
The agent's remote is a host the platform runs. A push to any other remote is refused. Your repository is touched only by graduation, by a person.
Per-job credentials
Each code task gets credentials minted with only its scopes, revoked when the job ends. External systems drive the same loop with an API key scoped to code changes and a callback URL.
Watch it work
Opt into a live activity feed and a frame-based view of the agent's browser while it edits, boots and checks. Ask it for a correction in the change's conversation; it continues the same branch.
Custom agents are data
A headless agent task is a definition: instructions, tool groups, scopes, budgets. The first one repairs a drifted environment patch, proves a box boots with it, and resubmits the blocked job.
Where it stops today
- Validation runs the named test, not your whole suite. Nothing yet proves the change did not break something else; the graduated PR's own watched suite is where that happens.
- The agent needs a clonable repository and a bootable environment definition. A project that only tests a fixed URL cannot use code changes.
- Automatic retry of a code change after an infrastructure block is off by default (a change is a proposal, not a job to force through).
- The internal git host keeps every agent branch; nothing prunes graduated or dismissed branches yet.
Questions people ask
What stops a wrong change from reaching production?
Three gates. The agent can only push to a git host the platform owns. A change that failed its validation test is reported as unable, never as proposed. A person reviews the diff and evidence and chooses to graduate it, which opens a pull request that your normal review still has to merge.