Most teams meet their first coding agent as a tool: you open a terminal, type a prompt, watch it work, and paste the result somewhere. That is fine for one person on one afternoon. It stops working the moment you want three agents and five humans to ship in the same week, because nothing about a terminal session tells a teammate what the agent is doing, who asked for it, what it is allowed to touch, or whether the result was ever checked.
Running an engineering team with AI agents is mostly not an AI problem. It is the old management problem of roles, tickets, permissions, handoffs and review, applied to workers that are very fast, very literal and have no memory of last Tuesday. This playbook covers the structure that has held up in practice: how to split agent roles, how to write tickets an agent can actually finish, how to scope permissions and branches, how handoffs and shared channels work, how to keep cost under control, how things fail, and a 30-day rollout you can start on Monday. Where it describes product behavior, it refers to OpenVisio Team, where agents are assigned tickets on a shared board and talk in the same channels as people. The rest is generic engineering advice that works with any setup.
Key takeaways
- Give each agent one narrow role with a written definition of done, the same way you would hire a specialist rather than a generalist who "does everything".
- A ticket an agent can finish states the problem, the expected behavior, the files or areas in scope, and how the result will be checked. If you cannot write the check, the ticket is not ready.
- Scope permissions to the role: most agents should be able to push a branch and open a pull request, and nothing more. Humans merge.
- Use one branch and one worktree per task, including retries and review fixes, so concurrent work never collides.
- Hand off through artifacts (ticket comments, PR descriptions, channel threads), never through "the agent will remember".
- Cost control is a design decision, not a billing surprise: bound the context each task needs, cap retries, and measure per ticket instead of per month.
- Expect a small set of repeatable failure modes (vague tickets, scope creep, stale branches, review fatigue) and put a guardrail on each one.
- Roll out over 30 days: one agent, one ticket type, human review on everything, then widen only when the numbers say so.
Why a team of agents needs management
A single agent in a single terminal has an implicit manager: you. You watch it, you stop it when it wanders, you remember what you asked for. Add a second agent and a second human and that implicit management disappears. Three problems show up quickly.
First, invisibility. Work happening in a private terminal cannot be reviewed, prioritized or helped. A teammate cannot tell whether the login fix is in progress, blocked, or abandoned.
Second, collision. Two agents editing the same checkout will overwrite each other, and an agent editing your working tree while you edit it produces diffs nobody can explain.
Third, unbounded autonomy. An agent given a vague goal and broad permissions will do something plausible and large. Plausible and large is the most expensive kind of wrong, because it looks finished.
The fix is the same one human teams settled on decades ago: make work a first-class object (a ticket), give workers defined responsibilities (roles), put a gate between work and production (review), and keep the conversation where everyone can read it (channels). Agents change the speed and the failure modes, not the need for the structure.
Defining agent roles
The temptation is to run one all-purpose agent and ask it for everything. It works for a week. Then its context fills with unrelated history, its permissions creep up to cover the broadest task it ever received, and nobody can say what it is responsible for.
Prefer several narrow roles. A role is a short written contract: what the agent takes on, what it must never do, what "done" means, and who reviews.
A starting set of roles
| Role | Takes on | Never does | Done means | Reviewed by |
|---|---|---|---|---|
| Bug fixer | Reproducible bugs with a failing test or clear repro | Refactors beyond the fix, dependency upgrades | Failing test now passes, no other tests regress | Code owner of the area |
| Test writer | Coverage gaps in a named module | Changing production code | New tests pass and fail when the behavior is broken | Any engineer |
| Docs and changelog | Updating docs for merged changes | Editing source or config | Docs match merged behavior, links resolve | Whoever shipped the change |
| Refactorer | Mechanical, behavior-preserving changes in a bounded folder | Behavior changes, public API changes | Tests identical before and after, diff is reviewable | Code owner |
| Triage and research | Reading issues, reproducing, summarizing, proposing tickets | Writing code on the main product | A written summary with evidence and a proposed ticket | The person who asked |
| Reviewer assistant | First-pass review of PRs against a checklist | Approving or merging | Comments posted on the PR with file and line anchors | The human reviewer |
You do not need all of these. Many teams start with the bug fixer and the test writer, because both have objective checks. The refactorer and reviewer assistant earn their place later.
Write the role down
Put the contract where the agent will see it. In OpenVisio, an agent is registered as a team member with a name and a role description, for example "Backend fixes and tests", and it holds a team role on the same Owner, Editor, Commenter, Viewer scale humans use. Outside OpenVisio, the equivalent is a short section in your repo's agent instructions file. Either way, a role description should fit in a paragraph:
## Role: bug-fixer (agent "ada")
Takes: tickets labelled "bug" that include a reproduction or failing test.
Scope: the service named in the ticket. Do not edit shared libraries.
Never: upgrade dependencies, change CI config, touch migrations.
Done: a failing test is added first, then the fix; full test suite passes.
Output: a pull request from an agent/* branch, linked to the ticket,
with a summary of cause, change and how it was verified.
If blocked: comment on the ticket with what you tried, then stop.
The last line matters most. An agent that knows how to stop is worth more than one that is slightly smarter and never does.
Writing tickets agents can finish
Ticket quality is the largest single lever on agent output. The same agent that produces a clean pull request from one ticket produces a sprawling mess from another, and the difference is almost always the ticket.
The anatomy of a finishable ticket
OpenVisio's own agent guide asks for three things when you assign repository work: the problem, the expected behavior, and how the result should be checked. That is a good minimum. Extend it with scope and constraints and you get a template:
Title: Login rejects emails with a plus sign
Problem
Users with addresses like ana+work@example.com cannot sign in. The form
shows "Invalid email" before submitting.
Expected behavior
Any address that passes RFC 5322 basic validation is accepted. The
validator in the login form and the one in the API must agree.
Scope
- Login form validation (web/src/auth/)
- API request validation (api/src/auth/)
Out of scope: signup flow, password rules, UI copy.
How to check
- Add a unit test for ana+work@example.com on both validators.
- Existing auth tests still pass.
- No changes outside the two folders above.
Notes
Related decision: we do not normalize case in the local part.
Every section answers a question the agent would otherwise guess at. "Out of scope" is the cheapest scope-creep guard you will ever write.
Make the check executable
"Works correctly" is not a check. A command is a check. Prefer, in order:
- A failing test the agent must make pass (best, because it is objective).
- A command whose output you can read, such as a script or a curl against a local server.
- A short list of manual steps with expected observations.
- Reviewer judgment (weakest, and the slowest for you).
If you cannot write any of the first three, the ticket is probably a research or design task. Label it that way and expect a written proposal, not code.
Give context by reference, not by paste
Link the ticket to the relevant code areas, prior discussion and decisions rather than pasting walls of text into it. The agent can follow references; it cannot un-read a 3,000-line paste. If your agents query a code graph over MCP, the ticket can name entry points and let the agent pull the dependency neighborhood itself, which keeps context small. The cost angle is covered in how to reduce AI agent token costs.
Scoping permissions
Permissions are where role definitions become enforceable. A role description is a request; a permission is a fact.
Principle: least privilege, per role
Start every agent at the lowest level that lets it do its job, and raise it only for a specific reason. For most agents that means:
- Read the linked codebase and the project's tickets and channels.
- Comment on tickets and post in channels.
- Push to a branch with an agent-specific prefix and open a pull request.
It does not mean merging to the main branch, deleting remote branches, changing repository settings, or reading production secrets. In OpenVisio, coding mode lets an agent work in a configured workspace, push an agent/* branch and open a pull request, and the reference guide notes that remote branch deletion requires explicit authorization. Treat that default as a floor you build on, not a ceiling to remove.
Branches and worktrees per task
If you take one operational habit from this playbook, take this one: one branch and one worktree per task.
Why a worktree and not just a branch
A branch is a name; a checkout is the directory your tools actually read and write. Two agents sharing a checkout will switch each other's branch mid-edit. An agent sharing a checkout with a human will leave uncommitted changes in the human's way. A git worktree gives each task its own directory on its own branch, backed by the same repository, so tasks never touch each other's files.
# One worktree per ticket, on an agent-prefixed branch
git worktree add ../work/OVS-57 -b agent/ada/OVS-57-plus-sign-email origin/main
# Work, test, push from inside ../work/OVS-57
cd ../work/OVS-57 && npm test && git push -u origin agent/ada/OVS-57-plus-sign-email
# After the PR is merged
git worktree remove ../work/OVS-57
git branch -d agent/ada/OVS-57-plus-sign-email
Put the ticket reference in the branch name. It makes the branch searchable, makes the PR self-describing, and makes cleanup scriptable.
Reuse the branch for retries and review fixes
When a reviewer asks for changes, the agent should push to the same branch and update the same pull request, not open a new one. OpenVisio's agent guide spells out the rule: one branch and worktree per task in each repository, including retries and review fixes, and check for an existing branch and PR before creating either. Without this rule you get three half-finished PRs for one ticket and a reviewer unsure which is current.
Handoffs
Every handoff is a place where context dies. Agents make this worse because they have no hallway memory, but they also make it better, because they write everything down if you ask.
Handoff by artifact
A handoff should be readable by a stranger. For each transition, name the artifact:
| Transition | Artifact | Must contain |
|---|---|---|
| Human to agent | The ticket | Problem, expected behavior, scope, check |
| Agent to human reviewer | The pull request description | Cause, change, how verified, risks, what was not done |
| Agent to agent | A ticket comment or channel mention | One concrete request, the relevant link, the expected output |
| Agent back to requester | The final reply in the source thread | Result, link to PR, anything blocked |
| Anyone to future you | The merged PR and the ticket history | Why, not just what |
Agent to agent: keep requests concrete
When agents can ask each other for help, unbounded chatter is the failure mode. OpenVisio's agent guide describes a deliberate budget: automated turns require an explicit mention and a concrete request such as "@Alex verify the API contract", a limited number of distinct peer requests are accepted per human turn, and thanks-only replies or open-ended back-and-forth stay silent. You can apply the same rules by convention anywhere: a peer request names one agent, asks one thing, and says what a good answer looks like.
Shared channels and visibility
Channels are how a team of agents stays legible. The goal is that a human can glance at the right place and know what is happening without interrogating anyone.
Where agents talk
In a shared workspace, agents are members of the same channels as people. In OpenVisio, you can mention an agent in a channel to ask a question, reply in a thread to continue a conversation, and assign a ticket for repository work. Replies stay in the thread they came from. Completion messages go to the agent's dedicated channel by default, and you can explicitly ask it to announce something in another named channel, which it resolves from the live channel list and confirms in the source thread.
That gives you a simple channel design:
- One channel per agent or role for completion notices, so humans can mute or follow selectively.
- Project channels where humans and agents discuss work in threads, with the discussion attached to the thing being discussed.
- A triage channel where new issues land and the triage agent summarizes them.
Watch the work, not just the result
Final messages tell you what finished. They do not tell you what is stuck. OpenVisio publishes agent-wide presence (thinking, working, typing) so people see a live indicator when an agent is active, and Agent Studio on the machine running the watcher lets you open a cycle and inspect its plan, commands and response. Two caveats from the docs are worth repeating in your own process: a completed cycle does not mean the ticket is done, and a progress message describes ongoing work, so wait for a final result or a stated blocker.
Cost control
Cost scales with the number of agent turns, the context in each turn, and the number of retries. Control all three on purpose.
Know where the money goes
For most coding agents the dominant cost is context: the files and history re-read on every turn. An agent that explores blindly pays for every file it opens; one that is pointed at the right slice pays far less. A code graph exposed over MCP is one way to do that, replacing broad file crawls with structured queries for symbols, dependents and neighborhoods. OpenVisio's README reports roughly 30 to 90 times fewer tokens on discovery-heavy tasks than letting an agent crawl, a figure the project attributes to its own measurements, so measure on your repo before you plan around it.
Practical controls
- Right-size the model to the role. A docs agent does not need your most capable model. In Agent Studio you can set reply and coding models per agent, and saved settings apply to the next request on updated watchers.
- Cap retries. A ticket the agent has failed twice usually has a ticket problem, not a persistence problem. Put "after two failed attempts, comment and stop" in the role.
- Bound the task. Small tickets cost a predictable amount. Open-ended ones do not.
- Meter per ticket. Monthly bills hide which role is expensive. Record cost or token use against the ticket reference so you can see that the refactorer averages four times the bug fixer.
- Watch rate limits. Provider usage limits and charges apply when the agent runs. A rate-limited agent should be visible as blocked, not silently idle.
- Keep prompts and instructions short. Instructions load every turn. A 4,000-word agent instruction file is a tax on every task.
A worked example
Here is one ticket moving through the structure above. The scenario is illustrative, not a customer story.
Monday, 09:10. A user reports that addresses with a plus sign are rejected at login. The triage agent reproduces it from the report, finds both the form validator and the API validator, and posts a summary in the triage channel with file references. A human (the code owner of auth) approves turning it into a ticket.
09:25. The ticket is written using the template: problem, expected behavior, scope limited to two folders, out of scope signup and UI copy, check being a unit test on both validators. It is assigned to the bug-fixer agent in a project with a linked codebase.
09:26. The agent acknowledges in the ticket thread with a three-step plan: add failing tests for both validators, align the regex, run the auth suite. The code owner glances at it and says nothing, which is approval.
09:27 to 09:50. The agent creates a worktree on a branch named for the ticket, writes the failing tests first, fixes the validator, runs the suite. One unrelated test fails because of a pre-existing flaky timeout. The role says to comment and stop on ambiguity, so the agent notes the flake in the PR description rather than "fixing" it.
09:52. The agent opens a pull request with cause, change, verification and a note about the flaky test, links it to the ticket, and posts a completion message to its channel.
10:30. The reviewer reads the description, reads the small diff, asks for one change (a missing case with uppercase characters). The agent pushes to the same branch and updates the same PR.
10:45. The reviewer approves and merges. A human merged; the agent never could. The docs agent picks up a follow-up ticket to update the validation section of the docs.
What made this smooth was not a clever agent. It was a bounded ticket, an executable check, a role that said when to stop, a single branch for retries, and a PR description that made review cheap.
Failure modes and guardrails
Agent teams fail in repeatable ways. Each has a guardrail that costs little to add.
| Failure mode | What it looks like | Guardrail |
|---|---|---|
| Vague ticket | Large diff solving a problem nobody asked about | Ticket template with scope and an executable check; reject tickets without one |
| Scope creep | "While I was there I also refactored..." | Out-of-scope list; reviewer rejects unrelated changes; path restrictions in review |
| Duplicate work | Two agents or two PRs for one ticket | Check for existing branch and PR first; one branch per ticket |
| Stale branch | Agent works from an old base and conflicts on merge | Create worktrees from the current main; rebase before opening the PR |
| Silent blocker | Agent idle for hours, nobody knows | Role says "comment and stop"; presence indicators; daily look at blocked tickets |
| Review fatigue | Reviewers rubber-stamp big agent PRs | Small tickets; PR template; cap open agent PRs per reviewer |
| Phantom completion | Ticket marked done, behavior not actually verified | Done means the check ran; a completed cycle is not a completed ticket |
| Credential leak | Key pasted into a ticket or commit | Show-once keys, secret scanning in CI, rotate on suspicion |
| Cost drift | Bill grows with no visible cause | Per-ticket ledger; retry cap; model per role |
| Agent chatter | Agents replying to each other endlessly | Mention-only, one concrete request, a per-turn cap on peer requests |
Common mistakes
- One agent, every role. Context bloats and permissions creep. Split by responsibility.
- Letting agents merge. The time saved is small and the downside is large. Keep merge human.
- Sharing a checkout. Use worktrees. Always.
- No offboarding. When an agent role is retired or a machine is decommissioned, revoke the credential and remove the local connection. Agent Studio can remove an agent from a computer, including its local credentials, logs and memory, after you confirm by typing its name.
- Treating the ticket as a prompt. A ticket is a durable record. Write it for the reviewer six months from now, not just for the agent today.
A 30-day rollout plan
The aim is to widen autonomy only as fast as evidence supports it. Adjust the pace to your team, but keep the order.
Days 1 to 5: foundation
- Pick one repository with decent test coverage and a clear code owner.
- Choose one role: the bug fixer is usually the right first one.
- Register one agent with a written role description and Editor-level access or the lowest equivalent.
- Connect the agent on a machine you control, start its watcher, and confirm you can see its activity.
- Write the ticket template and the PR template. Agree on the branch naming convention.
- Decide who reviews and how quickly. Put it in writing.
Days 6 to 12: supervised tickets
- Assign three to five small tickets, each with an executable check.
- Read every plan and every pull request line by line.
- Keep a ledger: attempts, review rounds, tokens, outcome.
- After each ticket, fix the process, not just the output: if the agent guessed wrong, what was missing from the ticket?
Days 13 to 20: add a second role
- Add the test writer or docs agent, whichever has the most objective check.
- Give it its own branch prefix and its own completion channel.
- Run both roles in parallel on different tickets to confirm worktrees and channels keep them apart.
- Introduce the retry cap and the "comment and stop" rule if you have not already.
Days 27 to 30: decide
- Compare against your baseline: cycle time for comparable tickets, review rounds, defects found after merge, cost per merged ticket.
- Decide per role: widen, hold or retire. Widening means more ticket types or a second repository, not looser permissions.
- Write down what you learned in the repo's agent instructions so the next agent starts ahead.
How OpenVisio fits
OpenVisio Team makes this discipline the default path: agents are registered as team members with roles, tickets are assigned on a shared board, replies stay in threads, coding mode pushes agent/* branches and opens pull requests rather than merging, and Agent Studio shows each cycle's plan, commands and response. The MCP code graph gives an agent structural answers instead of making it read everything, which keeps tokens and tickets smaller. You can adopt the playbook without OpenVisio, and OpenVisio does not remove the need for it.
FAQ
How many agents should a small team start with?
One. Run a single agent in a single role for the first week or two, review everything, and keep a ledger. A second agent is worth adding only when the first is merging tickets reliably and review is not already your bottleneck. Adding agents faster than you can review them converts a productivity project into a queue.
Should agents be allowed to merge their own pull requests?
For almost every team, no. Merging is the human accountability point, and it is cheap compared with the cost of an unreviewed change reaching your main branch. If you later automate merges for a narrow class of changes (for example, docs-only changes with passing CI), make that an explicit, audited exception rather than a default.
What is the best first ticket for an agent?
A small bug with a reproduction, or a coverage gap in one module, where the check is a test that fails before and passes after. Avoid ambiguous product work, large refactors and anything touching authentication, billing or migrations for your first run.
How do I know an agent is stuck?
Look for it in three places: a live activity indicator in the workspace, the agent's recorded cycles and their outcome, and the ticket thread. Build "comment and stop on blockers" into the role so a stuck agent says so. A silent agent with a stale cycle and no comment is the signal to investigate.
What should I measure?
Merged tickets per week, review rounds per ticket, defects found after merge, cycle time for comparable tickets, and cost per merged ticket. Avoid counting lines of code or pull requests opened; those reward volume over correctness.
Conclusion
An engineering team that works with agents is still an engineering team. It needs roles that say what each worker owns, tickets that say what done means, permissions that make the boundaries real, branches that keep work separate, channels that keep it visible, and a human review step that nobody skips. Get those right on one agent and one ticket type, measure honestly, and widen slowly.
To go further, read how to assign tickets to Claude Code and Codex, the shared workspace for humans and AI agents, and AI project management for software teams. For the cost side, see how to reduce AI agent token costs. Step-by-step setup lives in /guides, and if you are evaluating tools, /compare lays out how OpenVisio differs from the alternatives.
- AI agents
- engineering management
- agent roles
- ticket workflow
- guardrails
Put your team and your agents in one workspace
Create tickets, assign them to people or coding agents, and follow the work from channel to pull request.
Get started free


