All work

01AI Platform2026

Internal AI Platform

MKCL's self-hosted LLM and agent infrastructure — moving inference in-house to eliminate per-token spend, meet data-sovereignty mandates, and give engineering teams reusable agent scaffolding.

RoleArchitect

Serving
Ollama · vLLM · Self-hosted inference
Models
Qwen · Gemma · Model gateway / routing
Interface
OpenWebUI · FastAPI · Unified API contract
Platform
Guardrails · LLM-as-judge evaluation · Agent + RAG framework

Problem

Running institutional workloads against hosted model APIs created three compounding problems at once: per-token spend that scaled with adoption rather than value, sensitive institutional data leaving organizational control, and a hard dependency on a vendor's model availability, versioning, and deprecation schedule.

For an organization handling sensitive institutional data under compliance mandates, the sovereignty constraint was not negotiable. That made self-hosting a requirement rather than an optimization — and turned the question into how to make in-house inference genuinely usable by teams who had been building against a hosted API.

Approach

I architected the platform around a shared model gateway exposing a unified API across backing models. Product teams integrate against one contract; underneath, models can be swapped, versioned, or routed without touching consumer code. That indirection is what makes model choice an operational decision rather than a migration.

Serving runs on Ollama and vLLM across the Qwen and Gemma model families, with OpenWebUI as the operator-facing interface for internal teams. The two serving runtimes cover different needs — throughput-oriented batch work and interactive single-request use — behind the same gateway.

Engineering challenges

The harder problem was not serving models but preventing every team from rebuilding the same infrastructure. Each new agentic service otherwise reimplements orchestration scaffolding, retrieval pipelines, and memory components — and each implementation carries its own subtly different bugs.

I packaged those as a reusable agent and RAG framework, so a team stands up a new agentic service by configuring components rather than writing them. The same reasoning drove making guardrails and evaluation shared platform components: prompt injection detection, PII masking, and LLM-as-judge scoring are defaults teams inherit, not per-project work each team remembers to do.

Architecture

Product teamsInternal servicesModel gatewayUnified API contractGuardrailsInjection · PIIAgent frameworkOrchestration · RAGOllamaInteractive servingvLLMThroughput servingModelsQwen · GemmaEvaluationLLM-as-judge
  • Input
  • Orchestration
  • Processing
  • Inference
  • Output
A shared gateway fronts multiple self-hosted serving runtimes, so product teams integrate against one contract while models are swapped or routed underneath. Guardrails and evaluation are platform components rather than per-project code.

Key decisions

  1. A model gateway rather than direct client integration

    Without the indirection, every model change becomes a coordinated migration across every consuming service. One API contract turns model choice, versioning, and routing into an operational decision made in one place.

  2. Two serving runtimes behind one interface

    Ollama and vLLM optimize for different access patterns. Standardizing on one would force half the workloads onto the wrong runtime; exposing both directly would leak that choice to every consumer. The gateway absorbs it.

  3. Guardrails and evaluation as platform components, not project code

    Safety and quality measurement implemented per-project means implemented inconsistently, and skipped under deadline. Making them defaults that teams inherit changes the failure mode from 'someone forgot' to 'someone actively opted out'.

Outcome

LLM workloads run on self-hosted infrastructure, removing per-token API spend and keeping sensitive institutional data inside organizational control — the compliance requirement that motivated the platform.

Engineering teams stand up new agentic services against reusable orchestration, retrieval, and memory components rather than rebuilding core infrastructure per project, with guardrails and evaluation inherited by default.