Introducing rEngin AI

August 27, 2026

How we built an in-house coding agent that knows the rEngin codebase — without ever sending the source out of the building.

Running a shared registry is not a single application. rEngin is a multi-repo stack: the registry engine, EPP edges, REST APIs, React and Next.js consoles, policy XML, billing, RDAP, DNSSEC tooling, and the operational systems around them. It spans Python, Rust, TypeScript, Java, C, XML, SQL and more, across fifteen-plus repositories.

That scale creates a practical problem. Institutional knowledge lives in git history, not in one engineer’s head. A question as simple as “where is the poll message submitted?” used to mean opening several checkouts and hoping you were looking at the right branch. Pasting proprietary registry code into a public LLM was never an option.

We needed a coding agent, not a chat window. The result is rEngin AI — DNS Africa’s internal AI platform. It searches company code, runs commands, edits files, and accounts for every token. Retrieval is a first-class tool. Raw source never leaves the box.

Why we built it ourselves

Four constraints drove the design.

Knowledge is distributed. The answer to most engineering questions sits somewhere across the rEngin SRS core, the Registry Operator Console, the Registrar Portal, policy engines, and the supporting services. Search had to be one tool call, not fifteen clones.

Proprietary code cannot leave the building. A coding agent that reads the registry and forwards whole files to a third-party API is a data-disclosure risk. The platform had to give us access to strong frontier models for generation while guaranteeing that raw source stays on-premises.

Chat is not enough. We needed an agent that can search, retrieve structural context, run commands, and edit — the same class of workflow we use to ship rEngin itself.

Spend must be visible. Once both cloud and local models are available, “who used which model, and what did it cost?” is a product requirement, not an afterthought.

Those constraints map cleanly onto how we already think about rEngin: operators stay in control, data stays inside the trust boundary, and the platform is built for production rather than demos.

Design principles

Source stays local. Only retrieved, sanitised context snippets — budgeted to roughly 8,000 tokens / 24,000 characters — are sent to cloud APIs. Full-file fetches never leave the machine.

Secrets are stopped at the edge. Every generation request is scanned before it can leave. A match blocks cloud-bound traffic. Local generation is only logged (not blocked), because an agent’s tool results legitimately contain file contents and naive patterns produce false positives. Matched text is never written to logs — only the pattern.

Cloud first, local fallback — but not mid-conversation. Cloud frontier models are the default generation path. If they are unavailable or rate-limited, generation falls back to the local model. The switch is not made mid-session: changing model families mid-conversation would invalidate tool-call history.

The context window is enforced. The gateway budgets every request against the model’s real window. It estimates tokens from request size, self-calibrates from measured prompt_tokens, shrinks max_tokens before refusing, and returns a structured overflow response with real numbers instead of an opaque failure.

Accounting is part of the product. Every proxied call records the user token, upstream, model, input and output tokens, status, duration and bytes. Usage rolls up per token, per model and per day in an admin UI.

Retrieval is a tool, not prompt stuffing. The agent pulls context on demand. The proxy never silently injects RAG into a chat payload.

One public front door. Vector search, the graph store, embeddings, the reranker and the local model are not published on the host. Everything external goes through a single authenticated gateway.

How it works

The platform runs on a single NVIDIA DGX Spark workstation. One box hosts the local model, embeddings, vector and graph stores, the retrieval gateway, the generation proxy, the indexer, and private web search.

Generation proxy

The front door is a thin FastAPI proxy. It authenticates an issued token, chooses the upstream from the model id, and forwards the request path unchanged. User tokens never leave the proxy — upstream credentials are swapped at the edge.

Cloud generation currently routes to xAI (Grok) and Qwen Cloud. Local generation is an OpenAI-compatible SGLang endpoint serving Qwen3-Coder-Next-FP8 with a 262,144-token context — large enough for whole-file, tool-heavy agent sessions.

Retrieval

A docker-internal retrieval gateway is the only component that can correctly answer three policy questions at once: how large is this window, where may this payload go, and what may this request cost.

Search is a three-stage pipeline:

  1. Hybrid vector search (dense + sparse) in Qdrant
  2. Graph expansion in Neo4j — file → symbol → relationship
  3. Cross-encoder rerank, returning the top results

Vector search finds the text. The graph adds structure the registry actually needs: what a class extends, what a file defines, how a policy XML fragment relates to the engine that evaluates it. Reranking keeps the final context tight.

Indexing

A Rust indexer incrementally tracks committed git trees. It diffs HEAD against the last indexed commit, parses changed files with Tree-sitter, chunks around symbols, embeds on GPU, and updates Qdrant and Neo4j.

That matters for rEngin specifically. Registry policy is not just code — it is XML, XSD and XSLT. Tree-sitter gives symbol-aware chunking across more than fifteen languages, including those policy formats. Selected policy files can be stored whole so the agent can fetch them intact. Secrets and generated dumps are excluded from the index.

Re-embedding a changed file is cheap enough that a full repo refresh is measured in minutes, not hours. An optional watcher can overlay uncommitted work during active development.

The agent

rEngin AI is the agent harness on top of that stack: tools, sub-agents, sandboxing, and IDE integration. Generation is pointed exclusively at our proxy. There is no third-party login surface and no path for source to bypass the egress controls.

The same front door also exposes token-gated private web search and read-oriented operations tools (cluster state and logs). Production remains read-only.

What this means for rEngin

rEngin already exists to give registry operators control without giving away the stack. rEngin AI applies the same idea to how we build it.

A developer can ask where an EPP poll message is submitted, how a reserved-name rule is enforced, or how the Console talks to the registry engine — and get an answer grounded in the actual repositories. The agent can then propose an edit, run a command, or pull neighbouring symbols from the graph.

When cloud models are the right tool, they see only budgeted snippets. When they are not available, the local 256k-context coder continues the session. In both cases the spend is attributed, the secrets stay in-house, and the index tracks shipped code rather than a stale dump.

That is the point. We did not want a chatbot that pretends to know the registry. We wanted an agent that can work inside it.

In short

We wanted a coding agent that knows the rEngin codebase without ever leaking it. So we built one: a single DGX Spark runs the platform — a Rust incremental indexer that turns every repo into vector chunks and a symbol graph, a three-stage retrieval gateway, a local 256k-context coding model, and a transparent generation proxy that routes to frontier models with per-user accounting and egress secret scanning.

The result is rEngin AI. It searches company code with one tool call, falls back locally when the cloud is down, and keeps raw source inside the building.

rEngin AI is DNS Africa’s internal engineering platform. It is not a public product endpoint. For rEngin SRS, Console and Portal, visit renginsrs.com.