AI Coding Assistants: Measure the Complete Development Task

Updated September 8, 2026 · By Ryan Bold

An AI coding assistant can produce a patch quickly while still making the complete task slower. Understanding the request, reviewing the change, testing edge cases and repairing mistakes all count toward development time. The practical question is which tasks improve in your repository with your review standards.

What a productivity study actually tells us

METR’s early-2025 randomized study involved 16 experienced open-source developers completing 246 tasks in familiar projects. In that setting, allowing AI tools increased completion time by 19%. That result concerns a particular group, task selection and period; it is not a permanent verdict on all coding tools.

In a February 2026 update, METR explained that selection effects and measurement problems made its later experiment an unreliable estimate of the current productivity effect. Developers unwilling to work without AI and tasks omitted from the study could bias the results. Both reports matter when interpreting a headline about speed.

Choose a bounded first task

A small formatting bug, a well-specified parser change or an explanation of a local function is easier to evaluate than “modernize the application.” Give the assistant a concrete input, expected output and constraints. Keep credentials and private customer data out of example prompts unless the environment is explicitly approved for them.

For example, a date formatter requirement might state: accept an ISO timestamp, render it in the selected IANA time zone, handle invalid input visibly, and preserve the existing public interface. That is more reviewable than asking for a “smarter date system.”

Review behavior, not confidence

  1. Read the diff and confirm it solves the stated problem without unrelated changes.
  2. Check dependencies and APIs against the installed versions. A plausible function name may not exist.
  3. Test a normal case and the important boundary cases derived from the requirement.
  4. Inspect failure handling and sensitive data flows where relevant.
  5. Run the project’s required checks and record what remains untested.

For the date example, include a time-zone change that crosses midnight, an invalid timestamp and a daylight-saving boundary where applicable. A test that merely repeats the implementation’s calculation can agree with the same bug. Expected behavior should come from the requirement and trustworthy date/time rules.

A simple evaluation sheet

Record Why it matters
Task type and repository familiarity Results may differ between routine changes and unfamiliar systems.
Active work, waiting, review and repair time Generation time alone misses substantial work.
Accepted result and later regressions Fast delivery with defects can shift costs downstream.
Model, settings and tool permissions Another configuration may produce different results.

Compare similar tasks without forcing a universal percentage from a small sample. An assistant may help you explore unfamiliar code even when it does not shorten every task. Record that benefit separately from speed.

A time-zone task with checkable expected results

Here is a small requirement you can reuse: convert an ISO UTC timestamp into a selected IANA time zone, include the calendar date, and reject invalid input visibly. These are example acceptance cases, not measured results from a coding product:

Input and zone Expected local result
2026-12-31T23:30:00Z → Asia/Tokyo 2027-01-01 08:30
2026-12-31T23:30:00Z → UTC 2026-12-31 23:30
An invalid timestamp A visible validation error, not a plausible current time

Ask the assistant to explain where each expected result comes from before writing tests. Then inspect whether its tests assert a known answer or simply call the same formatter twice. The latter can agree with a bug.

For a hypothetical comparison, a manual task taking 30 minutes and an assisted task taking 12 minutes of implementation plus 23 minutes of review and repair yield totals of 30 and 35 minutes. The assisted implementation was faster, but the complete job took five minutes longer. Record both results honestly. A later task may differ; one example is not an estimate for your entire team.

Keep execution authority proportionate

A tool that suggests text differs from one that can install packages, edit files and deploy services. Start with reviewable changes and grant only the access required. Our agent workflow guide explains how to separate drafting from consequential actions. The developer remains responsible for understanding the change that reaches users.

For a small model you can evaluate on your own PC, follow our Ollama setup and local AI accuracy checks. The example separates a successful installation from answers that pass a defined test.