IntelliDocs
An enterprise knowledge platform where a LangGraph supervisor routes queries to specialized agents, decomposes multi-part questions into explicit plans, and answers from hybrid retrieval with measured groundedness.
RoleArchitecture & implementation
- Orchestration
- LangGraph · Supervisor routing · Plan-and-Execute
- Retrieval
- Dense vectors · BM25 · CrossEncoder reranking · pgvector · ChromaDB
- Backend
- Async FastAPI · JWT auth · Redis · Streaming responses
- Operations
- AgentOps · LLM-as-judge evaluation · Guardrails
- Frontend
- Vue.js admin dashboard · Embeddable chat widget
Problem
Enterprise knowledge is not one corpus. It is many departmental domains with different vocabularies, different access rules, and different notions of what counts as an authoritative answer. A single retrieval index over all of it answers narrow lookups adequately and fails on anything that spans domains.
The second problem is that complex questions are not single retrievals. A query with three parts needs three retrieval strategies and a synthesis step — but a one-shot RAG pipeline flattens it into one embedding lookup and returns something fluent that answers a third of the question.
Approach
I designed a multi-agent orchestration layer in LangGraph where a supervisor node routes queries to specialized retrieval and synthesis agents, coordinating explicit handoffs across departmental knowledge domains. The supervisor holds routing logic; the specialists hold domain knowledge. Neither has to know the other's internals.
On top of that sits a Plan-and-Execute workflow: complex multi-part queries are decomposed into an explicit step plan, each step executes against the appropriate retrieval strategy, and the results synthesize into a grounded answer. Making the plan explicit is what makes the system debuggable — you can read what it decided to do before looking at what it produced.
Memory spans both tiers. Long-term vector memory persists user context across sessions while short-term conversational state handles the current exchange, so agents carry prior interactions forward instead of restarting cold on every query.
Engineering challenges
Retrieval quality was the load-bearing problem. Dense vector search alone degrades badly on exact identifiers, policy codes, and rare proper nouns — precisely the terms enterprise queries turn on. I engineered hybrid retrieval combining dense vector search with BM25 lexical matching, then CrossEncoder reranking over the merged candidate set, running across pgvector and ChromaDB.
Multi-tenancy made this harder. Isolation had to be enforced at the retrieval boundary rather than filtered after the fact — a post-filter that runs after ranking can silently return a result count that reveals what exists in another tenant's corpus, even when the content itself never surfaces.
The third challenge was knowing whether any of it worked. I established an LLM-as-judge evaluation harness scoring response accuracy, groundedness, and task success, which converted agent quality from subjective review into a measured, trackable signal — and instrumented AgentOps observability across agent runs, token consumption, latency distribution, and failure modes, surfaced in a real-time monitoring dashboard.
Architecture
- Input
- Orchestration
- Processing
- Retrieval
- Storage
- Output
Key decisions
Supervisor-coordinated agents over a single retrieval chain
Departmental domains have different vocabularies and access rules. One chain either over-fits to one domain or is generic enough to serve none of them well. A supervisor routing to specialists keeps domain logic where the domain knowledge is.
Explicit step plans rather than an open-ended agent loop
A loop that runs until the model decides it is finished has unbounded cost and no inspectable structure. Decomposing into a visible plan means you can read the system's intent before reading its output — which is the difference between debugging and guessing.
Hybrid retrieval with reranking, not dense vectors alone
Embeddings blur exact identifiers and rare terms, which is what enterprise queries are frequently about. BM25 covers that blind spot; reranking resolves the merged candidate set. Each method compensates for the other's failure mode.
Tenant isolation enforced at the retrieval boundary
Filtering after ranking leaks information through result counts and latency even when content never surfaces. Enforcing at the boundary makes cross-tenant exposure structurally impossible rather than dependent on a downstream filter being correct.
Outcome
Agent quality became measurable rather than subjective: the LLM-as-judge harness scores accuracy, groundedness, and task success, and AgentOps surfaces run traces, token consumption, latency distribution, and failure modes in a real-time dashboard.
Shipped for enterprise rollout as an embeddable chatbot widget plus a Vue.js admin dashboard, so deployment into an existing product is a plug-in rather than an integration project.
- 40%
- Cost and latency reduction