AI coding, under real conditions.

Repository workflows, debugging, review and testing—evaluated on the work developers actually do.

What we investigate

Planned tests

Reviews

Task-led evaluations with environment, setup, intervention, failure cases and cost.

Framework

Comparisons

Like-for-like workflows, explicit criteria and a clear answer to who should avoid each tool.

In preparation

Real projects

Maintainable code, review notes and the reasoning behind every consequential decision.

Agent safety

Permissions, sandboxing, diffs, tests and human approval before consequential actions.

Read the publication standard