An AI coding agent can open a pull request in minutes. A human reviewer still needs the same hour they always needed, and the reviewer is now looking at code nobody on the team wrote line by line. The old review habits (skim the diff, trust the author, merge on green) quietly stop working when the author is a model that is confident, fast and occasionally wrong in ways that look perfectly reasonable.
This playbook is for teams that already let agents write code and now need a review process that keeps up. It is deliberately tool-agnostic: most of it is checklists, limits and roles you can adopt with nothing but your existing Git host and CI. We also show where a code graph, such as the one OpenVisio builds, fits in as one tool among several, specifically for the question reviewers find hardest: what else does this change touch?
Key takeaways
- Agent PRs fail differently from human PRs. The typical failure is not sloppy syntax, it is a plausible change that is subtly wrong about something outside the diff.
- Review the blast radius, not just the diff. Ask who depends on every changed export, and check that the PR accounts for them.
- Demand test evidence, not test existence. A new test that never failed before the fix proves little.
- Set a hard diff size limit and enforce it in CI. Small PRs are the single cheapest quality control you have.
- Keep a human in the loop at defined gates: merging, schema changes, auth, dependencies, and anything irreversible.
- Give agents their own branch or worktree per task so one runaway session never contaminates another.
- Split review into roles (author-owner, verifier, risk reviewer) so nobody rubber-stamps alone.
- Measure the process: review time, revert rate, escaped defects and diff size, then tune the limits with data instead of opinion.
Why agent pull requests need their own review process
What changes when the author is an agent
A human author carries context you can interrogate. You can ask why they chose this approach, and they remember the incident from last spring that explains the odd retry logic. An agent has none of that unless it was in the prompt or the files it happened to read. It also does not get tired, so it will produce the fifteenth consistent-looking file change of the afternoon without the hesitation a human would feel.
Three properties matter for review:
- Confident uniformity. Agent code is usually well formatted, well named and internally consistent. That removes the surface cues reviewers use to decide where to look closely.
- Local correctness, global blindness. The agent often reads a handful of files, makes a change that is right for those files, and misses a caller three directories away.
- Volume. Agents make it cheap to open many PRs. Review capacity, not authoring capacity, becomes the bottleneck.
The failure modes you will actually see
These come up repeatedly in practice, and the rest of the playbook is organised around catching them:
| Failure mode | What it looks like | Where the check lives |
|---|---|---|
| Missed callers | Function signature changed, two callers not updated | Blast radius check |
| Tests that prove nothing | Assertions mirror the implementation, or the test never failed | Test evidence |
| Scope creep | A bug fix also reformats and renames a module | Diff size limit |
| Invented APIs | Calls a method or flag that does not exist in your version | CI plus type checking |
| Silent behaviour change | Default value, ordering or error handling shifted | Risk review |
| Dependency drift | A new package added for one helper function | Approval gate |
| Stale base | Branch built on an old main, conflicts resolved by guess | Branch hygiene |
None of these is exotic. All of them are cheaper to catch with a checklist than with heroics.
The blast radius check
What blast radius means
The blast radius of a change is the set of code that can be affected by it: every importer of a changed module, every caller of a changed function, every consumer of a changed type or config key. A diff shows you what the author touched. The blast radius shows you what the author should have thought about.
For agent PRs this is the highest-value check, because agents are strongest at editing the file in front of them and weakest at knowing what depends on it.
A manual procedure that works anywhere
Before approving, for each changed exported symbol (function, class, type, constant, route, config key):
- Search for every reference. Use your editor's find-references, or at minimum a text search for the name.
- List the files that reference it that are not in the diff.
- For each, decide: is it unaffected, or should the PR have changed it?
- Check for indirect consumers: re-exports, barrel files, generated clients, documentation snippets, feature flags and environment variables.
Text search is imperfect. It misses dynamic imports, string-keyed lookups and re-exports under a different name. It also produces noise on common names. But it is a real improvement over reading only the diff, and it requires nothing but a terminal.
Where a code graph helps
A code graph answers the same question structurally instead of textually. OpenVisio parses a repository with tree-sitter into files, symbols and edges. Import edges are resolved file to file, and call edges are heuristic, symbol to symbol. Its get_dependents tool does directed impact analysis: who imports a target, or with direction=dependencies, what the target imports. get_neighborhood returns the local import subgraph around a file to a chosen depth, and get_hotspots ranks load-bearing files by import centrality, with git churn added when local history is present.
For a reviewer, that translates into a concrete step: take the changed files, ask for their dependents, and compare the list against the PR's file list. Anything important in the first list and absent from the second deserves a comment.
Two honest caveats. First, call edges are heuristic, so treat them as leads rather than proof. Second, a graph tells you what is connected, not whether the change is correct. It narrows where to look; it does not replace looking. If you use OpenVisio, the graph is built locally with no LLM and no network access, so running it as a review aid does not send your source anywhere.
Reading the result
Not every dependent needs a change. What you are checking is that the author, human or agent, can explain each one. A useful review comment is short:
get_dependents for src/billing/invoice.ts lists 11 files.
Only 3 are in this PR. Please confirm src/reports/export.ts and
src/api/webhooks.ts still behave correctly with the new rounding,
or add them to the PR with tests.
Better still, require the agent to produce that analysis itself in the PR description, then verify a sample.
A review checklist for agent PRs
The one-page checklist
Put this in your pull request template so it is visible to both the agent and the reviewer. Agents follow templates well, which makes this one of the easiest wins.
## Summary
What changed and why, in two or three sentences.
## Scope
- [ ] One concern only (no drive-by refactors or formatting)
- [ ] Diff is under the team limit (see CI check)
## Blast radius
- [ ] Changed exports listed below with their non-PR dependents
- [ ] Each dependent is either unaffected (reason given) or updated here
## Evidence
- [ ] New or changed tests listed
- [ ] Test failed before the change, passes after (output attached)
- [ ] Full suite, type check and lint pass in CI
## Risk gates (human approval required if ticked)
- [ ] Touches auth, permissions or secrets
- [ ] Database schema or migration
- [ ] New or upgraded dependency
- [ ] Public API or config contract change
- [ ] Deletes data or files
## Agent provenance
- Agent / model:
- Task or ticket:
- Branch or worktree:
The reviewer's pass order
Order matters because attention is finite. A pass order that works:
- Read the description first, then the tests. Tests state intended behaviour. If the tests do not match the description, stop and ask.
- Check scope. Does the file list match the stated task? Unexpected files are the cheapest signal of trouble.
- Run the blast radius check on changed exports.
- Read the implementation with the tests in mind. Look for what is missing, not just what is wrong.
- Check the risk gates. If any box is ticked, the right approver must sign off.
What to look for in the code itself
Agent code tends to hide problems in a few specific places:
- Error handling that swallows. A broad catch that returns a default is a classic way to make a test pass.
- Hard-coded values copied from a test fixture into production code.
- Duplicated helpers. The agent did not find the existing utility, so it wrote another. Search for similar names before approving.
- Comments that restate the code while the non-obvious decision goes unexplained.
- Deleted tests or loosened assertions. Always scan the diff for removed test lines.
Test evidence, not test existence
Why green CI is not enough
A passing pipeline says the suite passed. It does not say the suite tested the change. An agent asked to fix a bug and add a test will often add a test that passes, and a test that passes before the fix proves nothing about the fix.
Ask for the red-green record
The simplest standard: for any behaviour change, the PR shows that the new test fails without the change and passes with it. You can ask for it in the description, or check it yourself in under a minute:
# 1. Check out the PR and keep only the new tests
git checkout agent/fix-invoice-rounding
git checkout main -- src/billing/invoice.ts
# 2. The new test should now FAIL
npm test -- invoice
# 3. Restore the fix; the test should PASS
git checkout agent/fix-invoice-rounding -- src/billing/invoice.ts
npm test -- invoice
If step 2 passes, the test does not cover the bug. Send it back.
Evidence by change type
| Change type | Minimum evidence |
|---|---|
| Bug fix | Regression test that fails before and passes after |
| New feature | Tests for the happy path, one boundary case and one failure case |
| Refactor | Existing tests unchanged and passing; no assertion edits |
| Dependency bump | Changelog reviewed, lockfile diff scanned, full suite plus smoke test |
| Migration | Run forwards and backwards on a copy of realistic data |
| UI change | Screenshot or recording of before and after, plus keyboard and empty-state check |
For refactors the rule is strict: if the agent had to edit existing assertions to get green, it did not refactor, it changed behaviour. Treat any modified assertion in a refactor PR as a finding.
Mutation and property checks, selectively
For high-risk logic such as money, permissions or parsing, a stronger test than hand-written examples may be worth the cost. A property test or a mutation testing run on just the changed module shows whether the tests would notice if the logic were subtly broken. Reserve this for code where a wrong answer is expensive. Running it everywhere is a good way to make your pipeline slow and your team resentful.
Diff size limits
Why limits work
Review quality falls as diffs grow. Beyond a few hundred changed lines, reviewers skim, and skimming is exactly the failure mode that agent code exploits. A limit does two things: it protects reviewers, and it forces the task to be decomposed before the agent starts, which also improves the agent's output.
There is no universally correct number, and we will not pretend to have a benchmark for one. A reasonable starting point is a soft limit around 300 changed lines and a hard limit around 600, excluding lockfiles and generated files. Adjust it from your own review-time and defect data (see the metrics section).
Enforce it in CI
A limit that depends on reviewer goodwill will be ignored on the busy week. Make it a check:
# .github/workflows/diff-size.yml
name: diff-size
on: pull_request
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- name: Count changed lines
run: |
BASE=$(git merge-base origin/${{ github.base_ref }} HEAD)
LINES=$(git diff --numstat $BASE HEAD \
-- . ':(exclude)package-lock.json' ':(exclude)*.snap' \
| awk '{ s += $1 + $2 } END { print s + 0 }')
echo "Changed lines: $LINES"
if [ "$LINES" -gt 600 ]; then
echo "Diff exceeds 600 lines. Split the PR."
exit 1
fi
Allow an explicit override label for genuine exceptions, such as a mechanical rename, and require a named human to apply it. Exceptions should be visible and rare.
Make the agent split the work
The most effective instruction is upstream of the PR. Tell the agent in its task brief to propose a plan of independently mergeable steps first, then implement one step per PR. Stacked small PRs also give you natural rollback points: if step three is wrong, steps one and two still stand.
Human-in-the-loop approvals
Decide the gates in advance
Human approval is not a single checkbox at the end. It is a set of gates, defined before the agent starts, with named owners. The worst version is the implicit one, where whoever happens to be online clicks approve.
A workable gate model has three tiers:
| Tier | Examples | Who must approve |
|---|---|---|
| Routine | Tests, docs, small internal fixes, lint-level changes | One reviewer; CI green |
| Elevated | Public API changes, new endpoints, shared utilities with many dependents | Code owner of the area |
| Restricted | Auth, secrets, payments, schema migrations, new dependencies, infrastructure | Code owner plus a second named approver |
Use CODEOWNERS to make it mechanical
Your Git host can enforce the tiers. A CODEOWNERS file routes sensitive paths to the right people so the gate does not depend on anyone remembering:
# .github/CODEOWNERS
* @acme/engineering
/src/auth/ @acme/security @acme/platform
/src/billing/ @acme/payments
/db/migrations/ @acme/data @acme/platform
/package.json @acme/platform
/package-lock.json @acme/platform
/.github/workflows/ @acme/platform
Pair it with branch protection that requires code owner review and disallows the agent's own identity from approving. An agent should never be able to approve or merge its own pull request.
Separate permission from capability
Agents should run with the least privilege that gets the job done: write access to a feature branch, no ability to push to the default branch, no access to production credentials, and no authority to change branch protection or CI configuration. Treat changes to workflows and protection rules as restricted-tier by default, because an agent that can edit its own gates can silently remove them.
Branch and worktree hygiene
One task, one branch, one worktree
When several agent sessions share one working directory, they trample each other: uncommitted edits from task A end up in task B's commit, and a rebase in one session rewrites files under another. The fix is isolation. Give every task its own branch, and ideally its own Git worktree, which is a separate checkout of the same repository:
# Create an isolated checkout for one agent task
git fetch origin
git worktree add ../wt-fix-invoice-rounding \
-b agent/fix-invoice-rounding origin/main
# ... agent works in ../wt-fix-invoice-rounding ...
# Clean up after merge
git worktree remove ../wt-fix-invoice-rounding
git branch -d agent/fix-invoice-rounding
Worktrees share the object store, so they are cheap, and each has an independent index and working tree, so sessions cannot contaminate each other.
Naming and lifecycle conventions
- Prefix agent branches (for example
agent/) so reviewers and automation can recognise them at a glance. - Include the ticket or task id in the branch name so provenance is traceable.
- Always branch from a freshly fetched default branch. A stale base is a common cause of confusing conflicts that an agent will resolve by guessing.
- Delete merged branches and prune worktrees on a schedule. Orphaned worktrees accumulate quickly when agents are cheap to start.
- Never let an agent force-push to a shared branch. Force-push is acceptable only on its own task branch, and only before review begins.
Keep review context stable
Once a human starts reviewing, treat the branch as frozen except for requested changes. Agents that keep pushing "improvements" while a review is open invalidate the reviewer's mental model. Prefer follow-up commits with clear messages over history rewrites mid-review, and squash on merge instead.
Review roles for agent-heavy teams
Three roles, not one heroic reviewer
When one person is expected to read every agent PR in full, they either burn out or stop reading. Splitting the responsibilities makes each job smaller and more reliable:
| Role | Responsibility | Typical owner |
|---|---|---|
| Task owner | Wrote the brief, defines done, accountable for the PR existing at all | The engineer who assigned the work |
| Verifier | Confirms the evidence: red-green tests, CI, manual check of the behaviour | A peer, or a second automated pass |
| Risk reviewer | Checks blast radius and the restricted-tier gates | Code owner of the touched area |
For routine PRs one person can hold two roles. For restricted-tier changes they must be different people. The task owner should never be the only one to approve; they are the most anchored to what the agent was supposed to do rather than what it did.
Using a second agent as a reviewer
An agent reviewing another agent's PR can be useful as a first filter. It is good at consistency checks, missing tests, unused imports and obvious bugs. Use it, but with realistic expectations: it shares blind spots with the author model, it may be satisfied by the same plausible-looking reasoning, and it cannot be accountable. Treat its output as a prepared list of questions for the human reviewer, not as an approval. A human still owns the merge.
Rotate reviewers
Reviewer fatigue is real, and agent PR volume makes it worse. Rotate the verifier role so no one person becomes the permanent filter, and track the load (see metrics) so you notice when one reviewer is absorbing most of the queue.
A worked example: a rounding fix that touches more than it says
Consider a realistic scenario. An agent is asked: "Invoice totals are off by a cent for some orders. Fix the rounding." It opens a PR the same afternoon.
What the PR contains
src/billing/invoice.ts: replaces a floating-point sum with integer cents and rounds once at the end.src/billing/invoice.test.ts: adds a test for the failing order total.- 38 changed lines, CI green, description is tidy.
At a glance this is a good PR. It is small, in scope and tested.
Applying the playbook
Step 1, scope. Two files, both in billing, matching the task. Pass.
Step 2, evidence. The reviewer runs the red-green procedure. With the old invoice.ts restored, the new test fails; with the fix, it passes. Pass.
Step 3, blast radius. The reviewer lists changed exports. calculateTotal now returns integer cents instead of a decimal number. Searching references, then checking dependents with a code graph, shows that calculateTotal is used by:
src/billing/invoice.ts(in the PR)src/api/invoices.ts(not in the PR)src/reports/export.ts(not in the PR)src/webhooks/payment-succeeded.ts(not in the PR)
Three consumers outside the diff, and the return unit just changed from dollars to cents. Opening them, the reviewer finds that the export formats the value directly into a CSV, so customers would see totals multiplied by 100. The existing tests for that module mock calculateTotal, which is why CI stayed green.
The review comment
Blocking: calculateTotal now returns cents, but it has three callers
outside this PR (api/invoices.ts, reports/export.ts,
webhooks/payment-succeeded.ts). reports/export.ts formats the raw value
into CSV. Either keep the public return unit unchanged and convert
internally, or update all callers with tests that do not mock
calculateTotal. Please prefer the first option to keep the diff small.
The outcome
The agent keeps the public contract (dollars) and does the integer arithmetic internally, which shrinks the diff further. The reviewer approves, and the PR description now includes the dependents list as a permanent record. Nothing about this required a special tool, but the dependents check was the step that found the problem, and a graph made it a two-minute job instead of a grep session through string-keyed lookups.
Where OpenVisio and other tools fit
Use the right tool for each question
No single tool covers review. Here is how the common options map to the playbook:
| Question | Useful tools |
|---|---|
| Does it build, type check and pass tests? | Your CI pipeline |
| What else does this change affect? | Find-references in your IDE, text search, or a code graph such as OpenVisio's get_dependents and get_neighborhood |
| Is this file risky to touch? | get_hotspots (import centrality plus git churn), or git log analysis |
| Are there security issues? | Static analysis and secret scanners |
| Is the diff too big? | A CI diff size check |
| Who must approve? | CODEOWNERS and branch protection |
| What does the codebase look like overall? | OpenVisio's City and Atlas views via openvisio view, or any architecture diagram you trust |
A lightweight way to try the graph in review
If you want to test the graph-based blast radius check, the open-source CLI is the shortest path:
npm install -g openvisio
cd your-project
openvisio # registers MCP configs and runs a first index
openvisio view # optional: open the 3D map locally
Then, in a review session with a coding agent that has the server registered, ask it to call get_dependents for each changed file and report the importers that are missing from the PR. Verify a few of them yourself before trusting the pattern. The index is deterministic and local, so the same commit gives the same answer for every reviewer.
It also pays off earlier in the process. An agent that calls resolve_context before editing sees the neighborhood of the files it is about to change, which makes missed callers less likely in the first place. See /blog/reduce-ai-agent-token-costs for how that same structured context cuts discovery cost, and /compare for how a code graph relates to other approaches.
Metrics: how to know the process works
Measure outcomes, not activity
Counting PRs merged rewards volume, which agents provide for free. Choose measures that reflect quality and reviewer health:
| Metric | How to read it |
|---|---|
| Median time to first review | Rising means the queue is outgrowing reviewers |
| Median review time per PR | Falling alongside a flat defect rate suggests PRs are getting easier to review |
| Median and 90th percentile diff size | A creeping 90th percentile means the limit is being worked around |
| Revert rate of agent PRs | The clearest escaped-defect signal |
| Defects traced to agent PRs, per month | Compare against human PR baselines in your own repo |
| Review rounds per PR | High counts point to weak briefs or weak templates |
| Reviewer load distribution | Detects the one person who is the silent bottleneck |
| Override label usage | Frequent size or gate overrides mean the rules do not fit reality |
Compare to your own baseline
Do not borrow numbers from other teams or articles, including this one. Codebases, risk tolerance and agent setups differ too much. Take a month of data on human PRs, a month on agent PRs, and compare the shapes. The goal is not agent PRs that look identical to human PRs. It is to know when they diverge and why.
Close the loop
Every escaped defect from an agent PR is a prompt to improve the playbook. Ask which check would have caught it: a missing dependent, a vacuous test, an oversized diff, a skipped gate. Then adjust the template, the limit or the CODEOWNERS file. Over time the process should absorb your team's actual failure history.
Anti-patterns and common mistakes
Process anti-patterns
- Rubber-stamp on green. Approving because CI passed. CI checks what you wrote tests for, not what the agent forgot.
- Self-approval by proxy. The engineer who prompted the agent approves the PR in thirty seconds because they "know what it does."
- The mega-PR. One agent session, one enormous diff, one tired reviewer. Split before authoring, not after.
- Gate by memory. Relying on people to remember that migrations need a second approver instead of encoding it in CODEOWNERS.
- Agent reviews agent, nobody reviews. Using a model's approval as the approval.
Technical anti-patterns
- Reading only the diff. The diff is where the author looked. The risk is where they did not.
- Trusting mocked tests. Tests that mock the very function that changed will pass through any contract break.
- Editing assertions to get green. Always treat modified expectations in a supposed refactor as a finding.
- Letting agents edit CI and protection rules. Doing so lets them quietly remove the checks that constrain them.
- Long-lived agent branches. The longer a branch lives, the more likely it is built on a stale base and resolved by guesswork.
Cultural anti-patterns
- Treating review as an agent tax. Reviewers who feel the process is pure overhead will shortcut it. Make the checklist short and the automation do the boring parts.
- Blaming the model. When an agent PR causes an incident, the question is which gate failed, not which model is worse.
- Never loosening anything. Rules that never change get ignored. Review your limits quarterly against the metrics above.
A rollout plan you can start this week
Adopting everything at once is a recipe for resistance. A sequence that builds trust:
- Week one: add the PR template and the agent branch prefix. No enforcement yet; just visibility.
- Week two: add the diff size check as a warning, then flip it to blocking once the team has seen what it flags.
- Week three: write the CODEOWNERS tiers and enable required code owner review on restricted paths.
- Week four: introduce the red-green evidence requirement for bug fixes and a dependents list for any changed export.
- Month two: start collecting the metrics, review them together, and tune limits.
- Ongoing: move agents to per-task worktrees and tighten their credentials to least privilege.
Resist the urge to automate the judgment. Automate the counting, the routing and the reminders, and keep humans on the decisions.
FAQ
Should AI agents be allowed to merge their own pull requests?
No, not for anything that matters. Allow agents to open PRs and push to their own branches, but require a human approval and a human-triggered or policy-gated merge. For very low-risk categories such as documentation typo fixes you might relax this deliberately, but write that exception down rather than letting it emerge.
How big should an agent pull request be?
Small enough that a reviewer can hold it in their head. A soft limit near 300 changed lines and a hard limit near 600 (excluding lockfiles and generated files) is a sensible starting point, but tune it from your own review-time and defect data. The more important rule is one concern per PR.
Is a code graph required for a blast radius check?
No. Find-references in an editor and careful text search cover most cases. A graph helps when the codebase is large, when dependencies are indirect, or when you want the check to be fast and repeatable. Tools like OpenVisio add resolved import edges, heuristic call edges and centrality ranking, which narrows where to look, but you should still read the dependents it surfaces.
Can another AI agent do the review?
It can do a useful first pass: consistency, missing tests, obvious bugs and summaries of the change. It should not be the approver. It shares failure modes with the author and cannot be accountable for the outcome. Use it to prepare questions for the human reviewer.
How do we stop agents from sprawling across unrelated files?
Constrain it at three points. Put an explicit scope in the task brief, enforce a diff size limit in CI, and add a scope checkbox to the PR template. When the file list does not match the stated task, send the PR back rather than reviewing the extra files.
What if reviewers cannot keep up with the volume?
Reduce volume at the source before adding reviewers. Require smaller PRs, throttle how many agent PRs a person can have open, rotate the verifier role and automate evidence checks. If the queue is still growing, you are generating more code than your organisation can safely absorb, and the correct response is to slow authoring, not skip review.
Do these rules apply to human PRs too?
Mostly, yes. Blast radius, test evidence, size limits and gates are good practice for any author. Agents just make the gaps visible faster, so teams often find the playbook improves human PRs as a side effect.
Conclusion
Reviewing agent pull requests is not a new discipline so much as the old one applied with less trust and more structure. Check what the change touches, demand evidence that it works, keep diffs small, put humans at the gates that matter, isolate each task, split the review roles and measure the result. A code graph is a good fit for one of those steps, the blast radius, and unnecessary for the others.
If you want to go further, the /guides section walks through wiring a graph into Claude Code, Cursor and Codex step by step, /blog/reduce-ai-agent-token-costs covers how structured context lowers what your agents spend on discovery, and /compare shows how OpenVisio relates to other code-understanding tools.
- code review
- AI agents
- pull requests
- blast radius
- team workflow
- impact analysis
Put your team and your agents in one workspace
Create tickets, assign them to people or coding agents, and follow the work from channel to pull request.
Get started free


