Profile
AI Engineer with 9+ years of software engineering experience, including 3+ years shipping production LLM applications. Delivered 5 GenAI products end-to-end on a multi-tenant SaaS platform at Intellora AI — a conversational WhatsApp agent, RAG-powered retrieval, AI lead qualification, a multi-step travel-planning agent, and an AI-powered CRM — driving a +41% lift in lead-scoring precision, −38% LLM costs, and 100% sales-team adoption in 4 weeks. Deep expertise in RAG & vector search, agent orchestration, LLM evaluation & observability, and real-time streaming AI interfaces. Previously scaled systems to 50K+ concurrent users at Licious and built a telehealth platform 0→1 as founding engineer at Allo Health.
Tech Stack
AI & ML Foundations
Machine Learning
Deep Learning
Natural Language Processing (NLP)
Transformers
Generative AI
LLMs & Foundation Models
AIOpenAI GPT-4o
AClaude
GGemini
L3Llama 3
MMistral
HFHugging Face
RAG & Vector Databases
PPinecone
pgpgvector
WWeaviate
QQdrant
CChroma
coCohere Rerank
Retrieval-Augmented Generation
Hybrid BM25 + Dense
Semantic Chunking
Grounded Citations
Agents & Orchestration
LGLangGraph
LCLangChain
LiLlamaIndex
VVercel AI SDK
PyPydantic
Function Calling
MCP · Tool Use
ReAct · Multi-Agent
Evals, MLOps & Observability
LSLangSmith
LfLangfuse
HHelicone
Golden Datasets
CI Regression Gates
Production Monitoring
Guardrails & PII
Languages & Backend
PyPython
FaFastAPI
TSTypeScript
NNode.js
NeNestJS
KKafka
Microservices
SSE · WebSockets
OAuth2/JWT · Multi-Tenant
Data, Cloud & Frontend
PgPostgreSQL
RRedis
awsAWS Bedrock
GGCP
AzAzure OpenAI
DDocker
K8Kubernetes
PtPyTorch
ReReact 18
NNext.js
Testing & AI Tooling
CCClaude Code
CuCursor
CyCypress
PwPlaywright
JJest
Experience
AI Engineer · Intellora AI
Jan 2024 — Present
Bengaluru · Multi-tenant GenAI SaaS platform
- Shipped 5 production LLM products end-to-end on a multi-tenant SaaS platform (NestJS, Next.js, PostgreSQL + pgvector, Kafka, OAuth2/JWT); defined the 12-month AI roadmap across 4 product squads.
- WhatsApp AI Agent: conversational agent on WhatsApp Business API (GPT-4o + LangGraph tool orchestration) with context persistence, per-tenant token metering, guardrails, and human handoff — +35% qualified leads, −28% low-intent traffic.
- AI Lead Qualification: LLM scoring agent combining structured CRM features with conversation embeddings (text-embedding-3-large + pgvector), evaluated on a 5K-row golden dataset — +41% sales-accepted-lead precision over the rule-based baseline.
- AI Chat Summarisation: Claude-based hierarchical map-reduce pipeline (topics, sentiment, action items) over multi-day conversations — manager review time cut from ~12 min to <90 sec per thread.
- AI Travel Itinerary Agent: multi-step planning agent on LangGraph state machines calling flight/hotel/POI APIs as tools, with Pydantic validation and SSE streaming — <6s p95 via prompt caching and parallel tool execution.
- AI-Powered CRM: auto-drafted follow-ups, deal summaries, next-best-action, and NL-to-SQL with read-only guardrails — 100% sales-rep adoption within 4 weeks.
- RAG Pipelines: retrieval over tenant knowledge bases (semantic chunking, hybrid BM25 + dense, Cohere reranking, citations); owned the prompt–data–eval lifecycle with LangSmith/Langfuse tracing and CI regression tests gating every prompt change.
- Cut LLM spend −38% at constant eval quality via Haiku/Sonnet/Opus model routing, prompt caching, structured-output retries, and per-tenant token budgets with cost/latency observability.
- Built token-streaming chat UIs in Next.js (Server Components, SSE/WebSockets); drove team-wide adoption of Claude Code and Cursor — feature delivery time −30%.
Founding Full-Stack Engineer · Allo Health
Jun 2022 — Dec 2023
Bengaluru · Consumer telehealth, 0 → 1
- Built the entire consumer telehealth platform (Next.js, TypeScript, Node.js, PostgreSQL) — booking, payments, diagnosis flows — 0 to production in 4 months; grew engineering from 1 → 6.
- Established testing & delivery culture: Jest units, 120+ Cypress and 150+ Playwright E2E journeys in CI; led checkout A/B experiments — +22% conversion.