Internal AI Platform
MKCL's self-hosted LLM and agent infrastructure — moving inference in-house to eliminate per-token spend, meet data-sovereignty mandates, and give engineering teams reusable agent scaffolding.
RoleArchitect
- Serving
- Ollama · vLLM · Self-hosted inference
- Models
- Qwen · Gemma · Model gateway / routing
- Interface
- OpenWebUI · FastAPI · Unified API contract
- Platform
- Guardrails · LLM-as-judge evaluation · Agent + RAG framework
Problem
Running institutional workloads against hosted model APIs created three compounding problems at once: per-token spend that scaled with adoption rather than value, sensitive institutional data leaving organizational control, and a hard dependency on a vendor's model availability, versioning, and deprecation schedule.
For an organization handling sensitive institutional data under compliance mandates, the sovereignty constraint was not negotiable. That made self-hosting a requirement rather than an optimization — and turned the question into how to make in-house inference genuinely usable by teams who had been building against a hosted API.
Approach
I architected the platform around a shared model gateway exposing a unified API across backing models. Product teams integrate against one contract; underneath, models can be swapped, versioned, or routed without touching consumer code. That indirection is what makes model choice an operational decision rather than a migration.
Serving runs on Ollama and vLLM across the Qwen and Gemma model families, with OpenWebUI as the operator-facing interface for internal teams. The two serving runtimes cover different needs — throughput-oriented batch work and interactive single-request use — behind the same gateway.
Engineering challenges
The harder problem was not serving models but preventing every team from rebuilding the same infrastructure. Each new agentic service otherwise reimplements orchestration scaffolding, retrieval pipelines, and memory components — and each implementation carries its own subtly different bugs.
I packaged those as a reusable agent and RAG framework, so a team stands up a new agentic service by configuring components rather than writing them. The same reasoning drove making guardrails and evaluation shared platform components: prompt injection detection, PII masking, and LLM-as-judge scoring are defaults teams inherit, not per-project work each team remembers to do.
Architecture
- Input
- Orchestration
- Processing
- Inference
- Output
Key decisions
A model gateway rather than direct client integration
Without the indirection, every model change becomes a coordinated migration across every consuming service. One API contract turns model choice, versioning, and routing into an operational decision made in one place.
Two serving runtimes behind one interface
Ollama and vLLM optimize for different access patterns. Standardizing on one would force half the workloads onto the wrong runtime; exposing both directly would leak that choice to every consumer. The gateway absorbs it.
Guardrails and evaluation as platform components, not project code
Safety and quality measurement implemented per-project means implemented inconsistently, and skipped under deadline. Making them defaults that teams inherit changes the failure mode from 'someone forgot' to 'someone actively opted out'.
Outcome
LLM workloads run on self-hosted infrastructure, removing per-token API spend and keeping sensitive institutional data inside organizational control — the compliance requirement that motivated the platform.
Engineering teams stand up new agentic services against reusable orchestration, retrieval, and memory components rather than rebuilding core infrastructure per project, with guardrails and evaluation inherited by default.