Five weeks, one box: we taught our AI coding agent the whole company codebase

September 14, 2026

Every engineering team has a version of this problem. The answer to “where is this handled?” lives in a repository you do not have open, and the person who knows is in a meeting. Our stack spans a registry engine, EPP edge services, REST APIs, a React front end and a set of XML policy files across many repositories, in more than a dozen languages and config formats. The knowledge was all there. There was just no way to ask for it.

We could not solve that by pasting our code into someone else’s chat model. Source code is the company’s most sensitive asset: it contains policy logic, certificate handling, and the occasional credential that should never have been committed in the first place. Any tool that forwards our repositories to a third-party API was off the table.

So we built the retrieval layer ourselves, and five weeks later it is part of the daily routine. Here is what we did, and what the traffic says about it.

What we built at DNS

One NVIDIA DGX Spark runs the whole platform. There is no cloud vector database, no managed search service, no external index.

A Rust indexer walks the committed git tree of every repository we mount. It parses each changed file with Tree-sitter and chunks it around real symbols functions, classes, traits, XSD elements, policy templates rather than fixed windows of text. Chunking by symbol is the difference between finding a definition and finding twenty call sites.

Chunks land in two stores. Qdrant holds the vectors for semantic search. Neo4j holds a graph of files and symbols, which is what lets a text match expand into structure: what this class extends, what this file defines, who calls this function.

Retrieval happens in stages. Vector search and keyword search both run against that index and their rankings are fused, the results are expanded through the symbol graph, and a cross-encoder reranker scores the final candidates. The embedding and reranking models run on the box, so candidate ranking never involves a network call.

In front of all of it sits a small proxy. It issues per-user tokens, routes generation to the model that was asked for, counts the tokens each person spends, and scans every request before it can leave our network.

Retrieval is a tool, not a stuffed prompt

The easiest way to inject company context into an agent is to glue search results into the system prompt on every turn. It is also the fastest way to waste a context window and to teach the model to ignore its tools.

Instead, retrieval is three tools the agent calls when it needs them: search the code, walk the symbol graph around a name it already has, or fetch a complete file. Search returns labelled snippets with file paths and line ranges; the agent then reads the file locally before using any of it. When a policy XML is needed in full, it is stored whole at index time and fetched directly no git checkout, no reassembling fragments.

What never leaves the building

This was the constraint that shaped everything else.

Raw source files, full repository contents, git history and database contents are never sent anywhere. What can leave is a bounded set of retrieved snippets a few thousand tokens of sanitised context and only when the request is going to an external model. Whole-file fetches stay on the box.

Requests are scanned for credential patterns on the way out. A match blocks an external request outright; the same match on a request bound for the model running on our own hardware is logged but allowed, because refusing there breaks ordinary work agent tool results legitimately contain file contents, and patterns that look like API keys match plenty of ordinary code. The gateway is the only component that knows which route a request is taking, which is why the policy lives there.

Repositories can also opt out of indexing entirely. A single committed file at a repo root excludes paths from vectors, graph and whole-file storage, and evicts anything that was already indexed. Secrets never enter the index in the first place, rather than being filtered out of results.

Five weeks of traffic

Because we own the proxy, none of this is a guess.

Last seven days: 8 active accounts seven people and a CI job pushed 1.36 billion input tokens through the platform, and 93% of that input was served from prompt cache, which is what lets a single box keep up. They generated 6.5 million tokens of answers.

Last thirty days: 9 active accounts and 5.6 billion input tokens. The monthly cache share is 26%, not 93%, because that window includes an early, heavily uncached stretch.

Every account’s lifetime usage fits inside those 30 days. The platform is younger than the numbers it is serving.

What did all of that cost? A Spark and some time. The dollar column in our dashboard is a modelled figure at rates we configured ourselves, not an invoice by that card the month reads $677, and our actual spend was $0. What went in was the box it runs on and the weeks of engineering to build it.

The part that surprised us

We expected a rollout. There wasn’t one.

The retrieval plugin ships inside our agent and enables itself on login, so there was nothing to install, no server to add, no configuration to copy between machines. The same binary and the same token work in the terminal and inside the IDE.

Within days the questions changed. Not “can it do this” but “ask the other repo too” people stopped treating repository boundaries as walls, because the tool does not care which repository an answer lives in.

It also improves itself

Every search is recorded as a compact trace: the query and the citations it returned, never the code itself. When a result is actually used a file fetched, a graph walked that gets attached to the same trace.

That gives us something we would otherwise be guessing at: real labelled pairs of queries and the results that turned out to be useful, generated as a by-product of ordinary work rather than a benchmark we invented. Ranking changes get evaluated against how people actually search.

The lesson

The obvious version of this project was a chatbot with our documentation in it. The version that worked was narrower and more opinionated: build the index ourselves, keep raw code on the premises, expose retrieval as tools instead of prompt text, and account for every token so nobody has to argue about cost from memory.

Five weeks, from first commit to part of the daily routine. The feedback loop on your own tools is measured in days, not quarters and that is the whole reason to build them.