← all writing

The work after “done”

Coding agents can hand us an implementation. What should count as a completed task?

Chirag Arora 9 min read

The awkward part of using a coding agent is often the moment it tells you it's done.

You asked it to fix a bug. It read the repository, changed a few files, added a test, and returned a tidy summary. A task that might have occupied your afternoon is now waiting for review.

Instead of asking a tool to complete a line of code, we can ask it to work through a problem: inspect the code, make a change, run commands, and revise the result.

It is no longer a fringe workflow. In Stack Overflow's 2025 developer survey, 84% of respondents said they used or planned to use AI tools in development. That includes ordinary assistants, not just autonomous agents. More specifically, GitHub reported in October 2025 that its coding agent was contributing to roughly 1.2 million pull requests a month.

But the last step of the workflow is still familiar. You open the diff and try to work out whether the task was actually completed.

The agent has saved you some implementation work. Now you need to know which parts of its summary you can rely on.

A convincing answer can still leave work unfinished

Consider a small pagination bug. A null cursor should return the first page, but the function returns an empty array instead.

An agent changes the function and adds a test. Its final message says:

Fixed the null-cursor bug.
Added regression coverage.
All tests pass.
No breaking changes.

Suppose the new assertion is:

assert.ok(Array.isArray(page(null)));

The test passes. Unfortunately, the old, broken implementation also returned an array. This test would not catch the bug returning tomorrow.

In this example, a test was added and the test passed. Neither fact establishes that it checks the behavior we care about.

You don't need to assume the agent is being dishonest. It may have run every command it says it ran. Its account of the work can be accurate while its conclusion is too broad.

Humans do this too. We confuse finishing the steps with achieving the result. An agent makes the distinction harder to notice because it can produce the code, the test, and the explanation in one smooth interaction.

Asking it “are you sure?” does not resolve this. What would help is something specific: show that the test catches the original bug.

Take “done” apart

The completion message above contains several claims with different kinds of evidence behind them.

“All tests pass” can be checked by running the relevant test command. “A test was added” can be checked against the task's starting state. “This regression test catches the bug” needs a before-and-after comparison. “No breaking changes” may need a much wider set of checks, including knowledge of callers that aren't in the repository.

Treating these as one yes-or-no answer loses information. A build can succeed while the regression claim is unsupported. A useful fix can coexist with an overly confident compatibility claim.

For each claim, we should be able to ask: what observation would support it, and what observation would contradict it?

That question is more useful than asking how confident the agent sounds. It is also more useful than having another model read the summary and say it looks reasonable. A second model can help with review, but agreement between two explanations is not the same as a passing execution or a reproduced failure.

I started building Redpen around this distinction. It is a local definition-of-done and verification layer for coding agents. The idea is simple:

“Done” is a claim. Evidence makes it true.

Its job is to keep the task, the agent's claims, and the evidence connected, so the reviewer doesn't have to reconstruct those connections from scratch.

Start before the implementation

To check what a task changed, you need to know what existed beforehand.

If a repository already had an uncommitted test when the task began, its presence at the end is not evidence that this task added it. Comparing only with the last commit would give credit for work that was already there.

Redpen records the task and a Git snapshot when you start. That snapshot includes the relevant staged, unstaged, and untracked state. Later checks compare with that starting point, even if changes have since been committed.

It doesn't establish who authored a change. It establishes a more modest fact: this changed after the task began.

The task also gets an editable definition of done. Redpen generates an initial checklist from the project's capabilities, but a generated checklist does not understand every requirement. The developer needs to make important requirements explicit. For a null-cursor fix, “tests pass” is weaker than “the null-cursor test fails before the fix and passes afterward.”

The basic workflow is:

redpen start "Fix pagination when cursor is null"

# work normally with Codex

redpen import codex
redpen check

The import brings in the completion message, not proof of its contents. Redpen parses recognizable claims and runs applicable verifiers independently. An agent's transcript saying that a command succeeded is not accepted as the command's result.

This separation also keeps the agent's prose from silently becoming the task specification. What we required and what the agent says it did are related, but they are not the same thing.

Make the test earn its place

The pagination example suggests a stronger kind of evidence than a changed test file.

Keep the new test fixed. Run it against the implementation from the task's baseline. Then run it against the current implementation.

For our example, the assertion should check the result:

assert.deepEqual(page(null), [1, 2]);

If this assertion fails on the baseline and passes on the current code, the test distinguishes the two implementations on the behavior we specified. If it passes on both, it hasn't demonstrated the regression.

Redpen v0.4 adds an opt-in verifier for this comparison:

redpen add "Null cursor regression" \
  --type regression-test \
  --path tests/null-cursor.test.mjs

redpen check

It creates baseline and current snapshots in temporary directories, puts the current test into the baseline snapshot, and executes both. A missing import doesn't count as catching the bug. The baseline must fail the same named assertion test that passes on the current snapshot.

The first implementation is narrow: a JavaScript node:test file with exactly one unskipped test, without dependency installation or build steps. Those temporary directories are not a security sandbox; tests still execute with your permissions.

Even when this check succeeds, it proves the tested case, not the entire pagination system. A poor assertion can still encode the wrong requirement. A human still has to decide whether the expected result is right.

“This test detects this difference” is a useful result. Calling it “the bug is definitely fixed” would throw away that precision.

Leave room for “unverified”

Redpen gives each condition or claim one of three results: PROVEN, FAILED, or UNVERIFIED.

The difference between the last two matters. A test command that runs and fails supplies negative evidence. A test command that cannot be discovered leaves a gap. Those shouldn't produce the same answer.

The same is true of compatibility. If the agent says “no breaking changes” and no verifier can establish it, that claim stays unverified. It may be true. Redpen does not know.

In the reproducible pagination demo, the result includes these lines, shortened here:

✓ Null cursor regression
  Fails with an assertion on the baseline; passes now.

✓ "All tests pass."
  npm run test · exit 0

? "No breaking changes."
  No deterministic evidence is available.

There is a separate decision about whether the task is done. By default, every required criterion must be proven, and a claim contradicted by evidence blocks DONE. An extra unsupported claim remains visible without enlarging the agreed task contract. With --strict-claims, every agent claim must be proven too.

A DONE verdict means the configured contract has been satisfied, not that every property of the software has been checked. If the contract is too weak, the verdict will be too weak. Making that contract inspectable is part of the point.

Evidence also has a lifetime. Tests that passed before another edit are evidence about the earlier code. Redpen checks whether repository or task inputs changed after verification; a saved receipt can become CHECK NEEDED. Otherwise, a report meant to reduce uncertainty could become another source of false confidence.

This doesn't replace CI or review

CI can already run excellent tests, including regression tests. A team can build task-aware checks into CI too. Redpen doesn't make those things newly possible.

It packages a particular workflow: record what this task requires, retain the agent's claims, collect evidence independently, and show which claims that evidence supports. The same underlying commands can be useful in CI and in a local task receipt.

Review still covers things the checks don't: whether the solution is appropriate, whether the tests reflect the requirement, and whether the change creates a problem elsewhere. A receipt should help someone find the remaining questions, not persuade them to stop asking.

The v0.4 release also includes a local Codex plugin beta to bring this workflow into the coding session. Automatic final-message capture and checks require a trusted hook and an explicitly bound session. Installing the plugin alone doesn't authorize execution.

Removing a manual step is useful only if the evidence remains independent.

A better handoff

I don't want the future of agentic coding to be a developer watching a terminal all day, checking every action as it happens. Delegation should give us time back.

But I also don't want the handoff to be a large diff accompanied by a reassuring paragraph. The person receiving it should be able to see what was required, what changed, which checks ran, what their results support, and what still needs judgment.

Redpen is an early attempt to make that handoff less expensive. It doesn't understand arbitrary software requirements or prove general correctness. Its next useful steps should make more real requirements checkable, not make the word DONE easier to print.

The measure of success is not how often it returns a green result. It is whether the person reviewing the work can make a better decision with less reconstruction.

Coding agents can take on more of the implementation. We still own the decision to ship. Better evidence is what lets those two things coexist.


Redpen v0.4 is open source and available as redpen-cli on npm. To try it, run npm install -g redpen-cli. The five-case demo uses real temporary Git repositories and simulated completion messages; its checks run independently.