Agentic AI EngineerMKCL · Pune

I buildagent systemsthat surviveproduction.

Three years turning agent systems from prototypes into production infrastructure — orchestration, retrieval, evaluation, and the self-hosted platform underneath them.

Hover any stage to inspect the pipeline

agent tracereq_8f2ka91c
running

Tap any stage to inspect it

latency
0ms
docs
—
confidence
—
throughput
—
Supervisorroute · tenant scope

A LangGraph supervisor routes to the right domain agent.

emitsrouted query

outputstreaming
  • awaiting request…
recent
  • req_7c04ba1f340ms6 docs0.94cached
  • req_29e8dd5a1240ms8 docs0.88
  • req_d15f0c731185ms7 docs0.91

An interactive trace of a retrieval-augmented generation pipeline. Each stage can be focused to read what it does and what it produces.

Systems built end to end

Four production systems built at MKCL — agent orchestration, enterprise retrieval, and the self-hosted platform underneath them.

How the problems changed

Three years at MKCL, and the shift in what a unit of work meant across them — from features, to services, to the platform other teams build on.

  1. A feature

    Features

    2022 — 2023

  2. A service

    Services

    2024

  3. A platform

    Architecture

    2023 — Present

  1. 2023 — Present (current role)

    MKCL

    Senior Project Associate II

    Maharashtra Knowledge Corporation · Pune

    • LangGraph
    • LangChain
    • Python
    • FastAPI
    • pgvector
    • Redis
    • Docker
    • MCP

    Making agent systems safe, measurable, and reusable across an organization.

    Delivered five production AI systems across edtech, agritech, and public-sector domains, owning agentic architecture, multi-model orchestration, evaluation, and secure API design. Currently building MKCL's internal self-hosted AI platform, delivering data-sovereign inference and reusable agent infrastructure to engineering teams.

    • Agentic architecture across five production systems

      Architected ReAct and Plan-and-Execute agent workflows in LangGraph with supervisor-coordinated routing, explicit inter-agent handoffs, tool calling via MCP servers, and layered fallback logic. Designed agent memory across both tiers — short-term conversational state and long-term cross-session vector memory — so agents carry context forward instead of restarting cold.

    • Production readiness as an owned concern

      Owned the path from working prototype to operable system end-to-end: LLM-as-judge evaluation scoring accuracy and task success, AgentOps observability across runs and failure modes, and security guardrails against hallucination, prompt injection, and PII leakage. The point was converting agent quality from something reviewed subjectively into something measured.

    • Self-hosted platform infrastructure

      Architected the internal AI platform to move LLM workloads onto self-hosted inference — removing per-token API spend, meeting data-sovereignty mandates for sensitive institutional data, and bringing model choice and versioning under organizational control. Packaged agent and RAG scaffolding so teams build services rather than infrastructure.

    • Technical leadership

      Led a team of four engineers delivering production GenAI and agentic systems, owning technical design decisions, code reviews, sprint planning, and production deployments. Mentored engineers across LangGraph orchestration, RAG architecture, prompt engineering, agent evaluation, and AI security practices, and translated stakeholder use cases into agent workflows with defined handoffs, fallback paths, and safety boundaries.

  2. 2024

    MKCL

    Campus Live — University Management System

    Platform engineering

    • Go
    • Vue.js
    • MySQL
    • REST APIs
    • RBAC

    Building auditable systems for academic data across many institutions.

    Backend and full-stack work on a university management system deployed across institutions in Maharashtra — the pre-AI half of the same tenure, and where the service-design habits came from.

    • Core module development

      Developed five core modules — admissions, exams, fees, reporting, and notifications — in Go and Vue.js against MySQL, deployed across 20+ institutions in Maharashtra.

    • Access control for academic records

      Implemented role-based access control ensuring secure, auditable handling of academic data across institutions with different administrative structures — the constraint that shaped how the modules were bounded.

  3. 2022 — 2023

    CDAC

    Advanced Computing

    AIT YCP

    • Java
    • J2EE
    • React JS
    • JavaScript

    Learning to write correct code against someone else's design.

    Advanced Computing programme covering Java, J2EE, React JS, and programming fundamentals — the formal grounding before the MKCL tenure.

    • Foundations

      Java, J2EE, React JS, and programming fundamentals, alongside a web development certification from Internshala.

Capability map

What each set of tools is for, and how I use it. Marked items are what I reach for first.

Agentic Systems

Multi-agent architectures where routing, handoffs, and fallback paths are designed up front — because an agent loop without them is a demo, not a system.

  • LangGraph
  • LangChain
  • Multi-agent orchestration
  • ReAct
  • Plan-and-Execute
  • Supervisor & router agents
  • Agent handoffs
  • Tool calling
  • MCP servers
  • Short & long-term memory

Retrieval & Grounding

Hybrid retrieval with reranking, because dense vectors alone blur exactly the identifiers and rare terms real queries turn on.

  • Hybrid search
  • pgvector
  • CrossEncoder reranking
  • BM25
  • ChromaDB
  • Chunking strategies
  • Source traceability
  • Multilingual RAG

Evaluation, Safety & Operations

Agent quality treated as a measured signal rather than a subjective review — and safety as a platform default rather than a per-project effort.

  • LLM-as-judge evaluation
  • AgentOps
  • Prompt injection detection
  • PII masking
  • Groundedness scoring
  • Response validation
  • Hallucination prevention
  • Cost & latency optimization

Platform & Backend

Async services and self-hosted inference, chosen for operational cost and data control rather than novelty.

  • Python
  • FastAPI
  • Ollama / vLLM
  • Go
  • Microservices
  • Redis
  • Docker
  • JWT / RBAC
  • Vue.js
  • Model gateway & routing

Anyone can make a model produce an answer.
The engineering is in what happens when it’s wrong.

How I think about building

Positions I hold about engineering — including the ones with a real cost attached.

  1. A demo proves the idea. Production proves the engineering.

    The distance between a working prototype and a system that survives contact with real inputs is where most of the actual work lives — error handling, degradation under load, the malformed input nobody anticipated. AI makes this gap wider, not narrower, because a model that fails gracefully in a demo fails confidently in production.

  2. If you cannot measure it, you are not improving it — you are changing it.

    Retrieval quality, agent reliability, and answer correctness all feel improvable by intuition and are not. I build the evaluation harness before the thing it evaluates, because otherwise every subsequent change is a plausible-sounding guess with no way to distinguish progress from regression.

  3. Complexity should be paid for, not accumulated.

    Every abstraction, service boundary, and dependency has a carrying cost that is paid on every future change. The question is never whether a pattern is good practice but whether this system, at its current size, is getting more back than it pays. Most systems are simpler than their architecture suggests.

  4. Language models are components, not architectures.

    A model is a probabilistic function with a latency profile, a failure mode, and a cost per call. Treating it as a component subject to the same engineering discipline as any other — bounded, observable, replaceable — produces systems that keep working when the model does something unexpected. Treating it as the architecture produces systems nobody can debug.

  5. Interfaces should expose the decision, not the data model.

    The structure that makes a system easy to build is rarely the structure that makes it easy to operate. Users arrive with a question, not a schema; an interface earns its complexity by answering that question faster than the raw data would.

  6. Performance is a design constraint, not a later optimization.

    Speed is not something applied to a finished product — it is a consequence of decisions made early about payload size, rendering strategy, and how much work happens before a user sees anything. Retrofitting it means undoing those decisions rather than tuning around them.

Things I've learned the hard way

Observations from building these systems — the kind that only show up after something has been in production for a while.

  • Retrieval

    Chunking is a semantic operation disguised as a string operation

    Splitting on a token count severs claims from the conditions that qualify them, so a passage can be retrieved with its meaning inverted — and it still looks like a clean, relevant result. The failure is invisible at retrieval time and only surfaces in the answer.

  • Evaluation

    A vibes-based eval is worse than no eval

    Reading twenty outputs and judging them good is not measurement, but it produces the same confidence as measurement. It licenses changes that feel like progress with no way to detect regression. A small fixed question set beats a large informal one every time.

  • Agents

    Retry safety is a property of the tool, not the caller

    Every agent loop eventually retries a step. If the tool signature does not distinguish read-only from mutating, that distinction lives only in the author's memory — and gets lost the first time someone adds a step under deadline.

  • Interfaces

    Live-updating tables are unreadable

    Rows that re-sort while you are reading them make an interface technically fresh and practically useless. Buffering updates until an interaction boundary, or pinning the row under inspection, costs a little staleness and buys back the ability to actually read the thing.

  • Context

    More context is not more grounding

    Filling a context window with everything plausibly related lowers answer quality: the model has to locate the relevant span inside the noise, and it does that imperfectly. Precision in retrieval matters more than recall past a fairly low threshold.

  • Systems

    The interesting failure is the one that succeeds

    A crash announces itself. A model that returns a confident, well-formatted, wrong answer does not — and nothing downstream is built to catch it. Most of the engineering in an AI system goes into making silent failures loud.

Background

I build systems that make decisions — and I care most about what happens when those decisions are wrong.

I work at MKCL in Pune, where I've spent three years taking agent systems from prototype to production across edtech, agritech, and public-sector work — and more recently building the self-hosted platform the rest of it runs on.

The part I find genuinely interesting is the gap between the two. A model that produces a good answer in a notebook and a system that produces a defensible answer under load are different engineering problems — the second one needs evaluation, observability, fallback paths, and guardrails that hold when a model does something you did not anticipate. Most of my work lives in that gap.

I lead a team of four and spend a fair amount of time mentoring engineers on agent orchestration, retrieval architecture, and AI security — which has made me better at explaining why a design is right, not just that it is.

Based in
Pune, India
Focus
Multi-agent systems · RAG platforms
Currently
Self-hosted AI platform at MKCL
Open to
Agentic AI & platform engineering roles
Education
Bachelor of EngineeringShram Sadhana Bombay Trust, Jalgaon · 2017 — 2022

Available for work

Have a problem that
doesn’t have an obvious solution?

The interesting work usually starts there. If you’re building something with AI at its core — or trying to make something existing actually reliable — I’d like to hear about it.