Ask a coding agent a question about a repository it has never seen and watch what it does. It runs a search, opens a file, opens another, searches again, and slowly assembles a mental model one tool call at a time. Every step costs tokens, and every step can miss something. A code knowledge graph is the alternative: do the structural work once, ahead of time, and let the agent ask precise questions of the result.
This guide is the long version. It explains what a code knowledge graph is, what lives in it, how one is built from source (parsing, symbol resolution, call edges), how agents query it, how to keep it fresh, where it falls short, and how it compares with embeddings and plain grep. It also covers connecting a graph to an agent over MCP and, importantly, how to tell whether the graph you have is any good. Where we mention specifics of OpenVisio, they come from how the open-source tool actually works, and we flag the limits as we go.
Key takeaways
- A code knowledge graph stores files, symbols, and the relationships between them (imports, definitions, calls) so an agent can query structure instead of reading text.
- Graphs are built with a real parser (tree-sitter is the common choice), a symbol resolution pass, and heuristics for call edges. Import edges are usually precise; call edges are usually approximate.
- Graphs, embeddings, and grep answer different questions. Graphs are best at "who depends on this", embeddings at "where is the thing about X", grep at "where does this exact string appear". Good agents use all three.
- Freshness is a first-class concern. Content hashes and incremental re-parsing keep a graph current without rebuilding it on every edit.
- Quality is measurable. Check parse coverage, import resolution rate, spot-check call edges against known truths, and compare agent tool calls with and without the graph.
- Graphs help most on large, structured repositories. Small or greenfield repos see little benefit.
What a code knowledge graph is
A code knowledge graph is a data structure that represents a codebase as nodes and edges. Nodes are things you can name: a file, a function, a class, a type. Edges are relationships between them: this file imports that file, this function calls that function, this class implements that interface.
That is the whole idea. The value comes from what you can do once the structure is explicit.
Why text is a poor interface for agents
Source code is text, but its meaning is relational. The question "what breaks if I change this function signature" is not answered by any single file. It is answered by following edges outward from a node. A language model reading raw files has to rebuild those edges from scratch, in its context window, every session. That is slow, token-hungry, and error-prone, especially in repositories too big to fit in context at all.
A graph precomputes the answer to the structural part of that work. The agent still reasons, still reads code when it needs to, and still writes the edits. It just stops spending its budget on rediscovering what a parser already knows. We cover the cost side of this in /blog/reduce-ai-agent-token-costs and the broader context problem in /blog/ai-codebase-context-guide.
Knowledge graph versus call graph versus dependency graph
These terms overlap, so it helps to separate them.
| Term | Nodes | Edges | Typical use |
|---|---|---|---|
| Dependency graph | Files or packages | Imports, requires | Build order, impact analysis |
| Call graph | Functions or methods | Calls | Tracing execution, finding dead code |
| Code knowledge graph | Files, symbols, types, sometimes docs and commits | Imports, calls, definitions, inheritance, ownership | General-purpose querying by humans and agents |
A code knowledge graph is the superset. It usually contains a dependency graph and a call graph as subgraphs, plus metadata such as signatures, line ranges, export status, and git history.
Nodes and edges: the schema that matters
You can design a graph schema with dozens of node and edge types. In practice a small schema covers most agent questions, and a small schema is easier to keep accurate.
Node types
- File: path, language, lines of code, a content hash, last-modified time.
- Symbol: function, class, interface, type alias, constant. Each carries a name, a signature, a start and end line, and whether it is exported.
- Package or module (optional): a grouping node for monorepos, so queries can stay inside a boundary.
- Commit or author (optional): useful for churn and ownership signals.
Edge types
- Imports (file to file): resolved to a real target path, not just a string like
../utils. - Defines (file to symbol): which file declares which symbol.
- Calls (symbol to symbol): which function invokes which.
- Extends or implements (symbol to symbol): inheritance relationships.
- References (symbol to symbol): type usage, constant usage, where a call edge is too narrow.
Weights and derived metrics
Raw edges are only the start. Derived values make a graph usable at scale:
- Centrality. Running PageRank over the import graph surfaces the files everything else leans on. OpenVisio computes this deterministically, with fixed damping and iteration count, so the same repository always produces the same ranking.
- Churn. Commit counts over the last 30 or 90 days, read from local git history, flag files that change often.
- Hotspots. Centrality multiplied by churn points at code that is both load-bearing and unstable, which is where bugs tend to cluster.
A table of what each signal answers:
| Signal | Question it answers | Caveat |
|---|---|---|
| Centrality | What is load-bearing? | Favors hub files like index barrels; needs sanity checks |
| Churn | What changes often? | Needs git history; shallow clones lose it |
| Hotspot | What is risky to touch? | A heuristic, not a defect predictor |
| Fan-in | Who depends on this? | Only as good as import resolution |
How a code knowledge graph is built
Building a graph from a repository is a pipeline with four stages. Each can fail in its own way, which is why it is worth understanding them.
Stage 1: discovery
First, decide which files exist. This means walking the directory tree, honoring .gitignore, skipping vendored and generated directories such as node_modules, dist, and build output, and detecting language by extension or content. Getting this wrong has an outsized effect. Index a minified bundle by accident and your top-ranked file is a build artifact.
Stage 2: parsing
Parsing turns source text into a syntax tree. Regex is tempting here and always fails eventually, because code is nested and context-sensitive. The robust approach is a real incremental parser.
Tree-sitter is the dominant choice. It has grammars for a very large number of languages, produces a concrete syntax tree, tolerates syntax errors (important, since agents often work on half-edited files), and is fast. OpenVisio uses tree-sitter WASM grammars and covers 40+ languages, including TypeScript, JavaScript, Python, Go, Rust, Java, C, C++, C#, Ruby, PHP, Swift, and Kotlin. Config and doc formats such as JSON, YAML, TOML, and Markdown are tracked as file nodes but not symbol-parsed.
Symbol extraction works by running queries against the syntax tree. A simplified tree-sitter query for TypeScript function declarations looks like this:
(function_declaration
name: (identifier) @name
parameters: (formal_parameters) @params
return_type: (type_annotation)? @return) @definition
(class_declaration
name: (type_identifier) @name) @definition
For each match you record the name, the signature text, and the line range of the node. Those become symbol nodes with exact path:line anchors.
Stage 3: symbol and import resolution
This is where graphs earn or lose their credibility. A parser can tell you that a file contains import { parse } from './parser'. It cannot, by itself, tell you which file ./parser means.
Resolution maps an import specifier to a concrete file. For JavaScript and TypeScript that involves:
- Relative paths with implicit extensions (
./parsercould beparser.ts,parser.tsx,parser/index.ts). - Path aliases from
tsconfig.json(@/lib/parser). - Workspace packages in a monorepo.
- Package exports and
indexbarrel files that re-export symbols from elsewhere. - External packages, which you should record as external and not resolve into
node_modules.
Other languages have their own rules: Python's package and relative import semantics, Go's module paths, Rust's mod and use trees, Java's package-to-directory convention. Each needs its own resolver.
Resolution rate is the single most useful quality number for a graph. If 95 percent of internal imports resolve to a file, the import graph is trustworthy. If 60 percent do, most downstream answers are suspect.
Stage 4: call edges
Call edges are harder than import edges, and it is worth being plain about why. Whether foo() refers to the foo in this file, an imported foo, a method on an object, or a dynamic dispatch target depends on types and runtime values that a syntactic parser does not have.
There are roughly three levels of precision:
- Name-based heuristics. Match a call site to a symbol by name, preferring symbols in the same file, then in imported files. Fast, language-agnostic, and approximate. OpenVisio's call edges are heuristic in this sense.
- Scope-aware resolution. Track local variables, parameters, and imports to disambiguate. More accurate, more work per language.
- Type-aware resolution via a language server. Ask a compiler or language server where a call actually goes. Most accurate, but you now run and maintain a language server per language.
The right answer depends on use. For navigation and "show me what is probably related", heuristic call edges are good enough and cheap. For automated refactors that rewrite call sites, you want type-aware resolution. Treat call edges as evidence, not proof, and mark them that way in the output.
Stage 5: ranking and serialization
Finally, compute centrality and churn, assign stable ids, and serialize. Determinism matters here. If ids are assigned in sorted path order, the same repository yields the same graph every time, which makes caching, diffing, and regression testing possible.
Querying the graph
A graph is only useful if the agent can ask good questions cheaply. The query surface matters as much as the data.
The questions agents actually ask
Most useful agent queries fall into a few shapes:
| Question | Graph operation |
|---|---|
| Where is this thing defined? | Symbol lookup by name or pattern |
| What does this task touch? | Rank files and symbols by relevance to a description |
| What breaks if I change this? | Reverse traversal of import and call edges |
| What does this file rely on? | Forward traversal to a given depth |
| What is the shape of this repo? | Top-N by centrality, with public symbols |
| Where is the risk? | Centrality times churn |
Query languages versus purpose-built tools
You can expose a graph through a general query language such as Cypher or Gremlin. That is flexible, but it asks the model to write correct queries against a schema, which costs tokens and sometimes fails. The alternative is a small set of purpose-built tools, each answering one shape of question and returning a ranked, bounded result.
OpenVisio takes the second route with seven tools: resolve_context, get_repo_skeleton, find_symbol, get_neighborhood, get_dependents, get_hotspots, and get_languages. Every tool takes a budget_tokens parameter and returns ranked, elided output with exact path:line anchors. The tool surface is intentionally tiny because tool schemas load into the agent's context on every turn. A graph server with fifty tools can burn more tokens in schema than it saves in answers.
Budgets, ranking, and elision
Three design rules make graph output agent-friendly:
- Rank before you truncate. When output exceeds the budget, drop the least relevant material first, not whatever happens to come last.
- Elide bodies, keep signatures. A signature plus a line range usually tells an agent what it needs. It can read the body if it needs more.
- Always anchor. Every returned item should carry a
path:lineanchor so the agent can jump to real source. A graph result without an anchor is a dead end.
A worked example: "add rate limiting to the login endpoint"
Here is how a task flows through a graph-backed agent, compared with a crawl. The repository is a mid-sized TypeScript service. The numbers below illustrate the shape of the process, not a benchmark.
Without a graph
- The agent greps for
login. It gets forty matches across routes, tests, docs, and frontend code. - It opens the three that look most relevant, perhaps two thousand lines in total.
- It notices the handler calls
authenticate, greps for that, and opens two more files. - It looks for existing middleware, greps for
middleware, and opens more. - After a dozen tool calls it has a picture, mostly built from whole files.
With a graph
- The agent calls
resolve_contextwith the task text. It receives a task-ranked skeleton plus the neighborhoods of the most relevant files, each with signatures and anchors, in one response. - It sees that
routes/auth.tsdefines the login handler, thatmiddleware/index.tsexports existing middleware, and thatroutes/auth.tsis imported by the app entry point. - It calls
get_dependentsonmiddleware/index.tsto see who relies on it before changing it. - It reads only the two function bodies it needs to edit, using the anchors.
- It makes the change and runs tests.
The agent reads far less source, and the reads it does make are targeted. The graph did not write the code. It made the discovery phase short and safer. For the numbers behind this pattern, see /blog/reduce-ai-agent-token-costs.
A task-ranked view can also be previewed from the command line without any agent:
npx -y openvisio skeleton . --budget=1500 --task="add rate limiting to the login endpoint"
Reading that output yourself is the fastest way to judge whether the ranking matches your intuition about the repo.
Keeping the graph fresh
A stale graph is worse than no graph, because the agent trusts it. Freshness is a design requirement, not a nice-to-have.
Incremental indexing
Rebuilding everything on every edit does not scale. The standard approach is content addressing:
- Hash each file's content.
- Keep a cache from hash to parse result.
- On change, re-parse only files whose hash changed.
- Recompute edges that touch changed files, then refresh derived scores.
OpenVisio uses a content-addressed parse cache for exactly this, and its --watch mode re-indexes as files change. The cache also keeps file ids stable mid-session, so an id the agent saw a minute ago still points at the same file.
Where freshness still breaks
- Branch switches and rebases. Hundreds of files change at once. Make sure your watcher handles bursts, or re-index on a git hook.
- Generated files. If a build step regenerates code, either exclude it or index after the build.
- Uncommitted work. A graph built from git objects misses working-tree changes. Build from the filesystem, not from commits.
- Derived scores lag. Centrality is global. A single new import can shift rankings slightly, so recompute it after a batch of changes, not per file.
A good rule: the graph should be able to answer "when was this built" and "which files changed since". If it cannot, you cannot reason about staleness.
Graph versus embeddings versus grep
These are not rivals. They answer different questions, and the best agent setups combine them.
Comparison
| Dimension | Code knowledge graph | Embeddings (semantic search) | Grep / ripgrep |
|---|---|---|---|
| Strength | Relationships and structure | Fuzzy intent, natural language | Exact strings, literal patterns |
| Typical question | Who calls this? What depends on it? | Where is the code that handles retries? | Where does this error message appear? |
| Precision | High for imports, moderate for calls | Probabilistic, ranked by similarity | Exact |
| Setup cost | Parse and resolve once | Chunk, embed, store vectors | None |
| Infrastructure | Local index | Model plus vector store | None |
| Failure mode | Missed dynamic edges | Plausible but wrong chunks | Too many or too few matches |
| Freshness | Incremental re-parse | Re-embed changed chunks | Always current |
| Explainability | Edges are inspectable | Similarity scores are opaque | Matches are literal |
When each wins
Use the graph when the question is about structure: impact analysis, finding entry points, understanding module boundaries, locating the definition of a symbol, or deciding what to read first.
Use embeddings when you do not know the vocabulary. If the code calls it reconcile and you are thinking of "sync", a lexical search misses and a semantic one can find it. Embeddings are also good for comments, docs, and issue text.
Use grep when you know the string. Error messages, config keys, feature flags, and TODO markers are all best found literally. Grep is also the ground truth for checking claims: if the graph says nothing calls a function, a grep for its name is a fast sanity check.
Why not just embeddings?
Embedding-only retrieval has a structural weakness for code. It returns chunks that look like your query, but it does not know that a chunk is imported by forty other files or that it is a test fixture. Similarity is not importance. A graph supplies the importance signal that embeddings lack. OpenVisio's own design notes explicitly leave semantic search out of the first tool surface and treat it as something that could be added later as a separate ranked source, rather than a replacement for the graph.
Using a graph through MCP
The Model Context Protocol gives agents a standard way to call external tools. A graph exposed as an MCP server works with any client that speaks the protocol.
Why MCP fits
MCP gives you a uniform tool-call shape, stdio transport for local servers, and wide client support, including Claude Code, Cursor, and Codex. The server runs on your machine, reads your files, and answers over stdio. OpenVisio's MCP server is read-only and makes no network calls; it contains no LLM.
Connecting it
The one-command route:
npm install -g openvisio
cd your-project
openvisio
That writes project-scoped MCP configuration and runs a first index. Manual equivalents point at openvisio mcp . --watch:
claude mcp add openvisio -- openvisio mcp . --watch
{
"mcpServers": {
"openvisio": { "command": "openvisio", "args": ["mcp", ".", "--watch"] }
}
}
[mcp_servers.openvisio]
command = "openvisio"
args = ["mcp", ".", "--watch"]
One practical gotcha: if Node is installed through nvm and your client is a desktop app, the app may spawn servers with a minimal PATH, and a bare openvisio command will not be found. Use absolute paths to node and to the CLI entry file in the config.
Steering the agent
Registering a server does not guarantee the agent uses it. Agents default to the tools they know best, which are search and file reads. Two things help. First, put a short instruction in your agent's project rules, for example: call resolve_context first on any task, drill in with find_symbol and get_dependents, and read source only when a returned slice is insufficient. Second, if you want to enforce the habit, a pre-tool hook can block raw reads and searches until the graph has been consulted. OpenVisio ships an optional hook for Claude Code that does this. Enforcement is blunt, so try the instruction first and add the hook only if the agent keeps skipping the graph.
Watching the agent
When a graph is also rendered visually, you can watch the agent work. OpenVisio's spotlight mode streams each tool call to an open viewer on a local port, so the files the agent is inspecting light up in the map. It is a useful debugging aid: if the agent keeps looking at the wrong area, you can see it.
Limits of code knowledge graphs
A guide that only lists strengths is selling something. These are the real limits.
Dynamic behavior
Reflection, dependency injection containers, dynamic imports built from strings, decorators that register handlers, and metaprogramming create relationships that no syntactic parser can see. If your codebase leans heavily on these, expect missing edges. Framework conventions (file-based routing, for instance) can be added as special-case rules, but each is extra work.
Heuristic call edges
As covered above, call edges from a syntactic pass are approximate. Common names (get, run, handle) produce false positives. Treat them as leads to verify.
Cross-language and cross-service boundaries
A TypeScript frontend calling a Python service over HTTP has no import edge between them. A graph of each repository will not see the link unless you add a layer that matches routes to clients, or a schema such as OpenAPI that both sides share.
Small and greenfield repositories
If the whole repo fits comfortably in context, a graph adds little. Gains concentrate in large, structured repositories where discovery dominates the cost of a task.
The graph does not understand intent
A graph knows that function A calls function B. It does not know that the business rule in A is wrong. It supplies structure, not judgment. Tests, review, and the model's reasoning still carry that load.
How to evaluate the quality of a graph
You do not have to take any graph on faith, whether you build one or adopt one. Quality can be tested, and the tests are not exotic.
Structural checks
- Parse coverage. What fraction of source files parsed without error? Track files that fell back to "counted but not parsed".
- Import resolution rate. Of all internal import statements, how many resolved to a file? Look at the unresolved list; patterns show missing alias or monorepo support.
- Determinism. Index twice. Do you get identical ids and scores? If not, caching and diffing will be unreliable.
- Noise. Do generated files, vendored code, or lockfiles appear in the top-ranked list?
Spot checks
Pick ten symbols you know well and check them by hand:
- Does
find_symbolreturn the right definition and the right line range? - Does
get_dependentson a core module match what grep for its import path finds? - Are the top-ranked files the ones a senior engineer would name as load-bearing?
A quick cross-check for dependents:
# What does the graph say imports this file?
# What does a literal search say?
rg -l "from ['\"].*lib/parser['\"]" src
Differences between the two lists are either graph gaps or aliases the regex does not cover. Both are informative.
End-to-end checks
Structural quality is necessary but not sufficient. The real test is whether agents do better work.
- Choose five to ten representative tasks from your backlog.
- Run each without the graph and record tool calls, tokens consumed, and whether the result was correct.
- Run the same tasks with the graph.
- Compare. Look at correctness first, cost second. A cheaper wrong answer is not a win.
Run each condition more than once if you can, because agent behavior varies between runs. Use your own repositories and tasks; numbers from someone else's codebase tell you little about yours.
OpenVisio's MCP server prints a one-line receipt on shutdown comparing tokens its tools returned against the estimated cost of reading the touched files whole. It is an estimate, not a measurement of a controlled experiment, so treat it as a directional signal and confirm with the A/B approach above.
Common mistakes
- Indexing everything. Vendored code, build output, and lockfiles pollute rankings. Respect ignore files and add explicit excludes.
- Trusting call edges as facts. They are heuristic. Use them to find candidates, then verify.
- Returning too much. A graph tool that dumps unbounded output recreates the problem it solved. Enforce budgets.
- Exposing too many tools. Every tool schema is paid for on every turn. Prefer a handful of well-named tools.
- Ignoring freshness. A graph built at session start and never updated will mislead after the first big edit. Use a watcher or incremental re-index.
- Skipping anchors. Results without file and line anchors force the agent back to search.
- Treating the graph as a replacement for reading. The goal is to make reading targeted, not to forbid it.
- Measuring only cost. Savings that come with worse answers are not savings. Evaluate correctness alongside tokens.
- Assuming one retrieval method is enough. Keep grep for literals and consider embeddings for fuzzy intent.
FAQ
What is a code knowledge graph in simple terms?
It is a map of your codebase stored as nodes (files, functions, classes) and edges (imports, calls, definitions). Instead of reading text, a tool or an agent can ask structural questions like "who imports this file" and get a direct answer.
Is a code knowledge graph the same as a vector database?
No. A vector database stores embeddings and returns items that are semantically similar to a query. A knowledge graph stores explicit relationships and returns items connected by those relationships. They complement each other: one finds things that look related, the other finds things that are actually wired together.
Do I need an LLM to build a code knowledge graph?
No. Parsing, import resolution, and ranking can all be done deterministically. OpenVisio's indexer uses tree-sitter and PageRank with no LLM and no network access. An LLM can sit on top, for example to explain the graph in prose, but it is not required to build it.
How accurate are the call edges?
It depends on the method. Name-based heuristics are approximate and can produce false positives for common names. Type-aware resolution through a language server is more accurate but heavier. Import edges, once resolved, are generally far more reliable than call edges.
How often should the graph be rebuilt?
Ideally never from scratch during a session. Use content hashing and incremental re-parsing so only changed files are processed, and run a watcher so the graph tracks edits. Do a full rebuild after large structural changes such as a big rebase if you suspect drift.
Does a graph help with small projects?
Not much. If the repository fits in the model's context with room to spare, a graph adds setup for little gain. The benefit grows with repository size and structure.
Will the graph send my code to a third party?
That depends on the tool. OpenVisio's graph indexer and MCP server run locally, are read-only, and make no network calls. Whatever tool you choose, check what leaves your machine, and whether the model provider you use with your agent sees the slices the graph returns.
Which agents can use a graph through MCP?
Any MCP client. Claude Code, Cursor, and Codex are the common ones, and each has a documented way to register a stdio server, as shown above.
How do I know the graph is actually helping?
Run the same tasks with and without it and compare correctness, tool calls, and tokens. Also check structural metrics such as import resolution rate. If you cannot show an improvement on your own tasks, do not keep the extra moving part.
Conclusion
A code knowledge graph is not magic. It is a parser, a resolver, some ranking, and a query surface, assembled so an agent can spend its context on thinking instead of rediscovery. Its strengths are real, for structure, impact analysis, and cheap orientation in big repositories, and so are its limits, which are dynamic behavior, approximate call edges, and cross-service links. Pair it with grep and, where it helps, embeddings, measure it on your own tasks, and keep it fresh.
If you want to try one, the /guides section walks through connecting OpenVisio to Claude Code, Cursor, and Codex. For the cost side of the argument, read /blog/reduce-ai-agent-token-costs, and for how it stacks up against other approaches, see /compare.
- code knowledge graph
- AI coding agents
- MCP
- tree-sitter
- codebase context
- static analysis
Put your team and your agents in one workspace
Create tickets, assign them to people or coding agents, and follow the work from channel to pull request.
Get started free


