AI Engineer focused on building reliable intelligent systems across LLM applications, RAG, evaluation, production ML, cloud platforms, and developer tooling. I work at the intersection of AI systems and software engineering, turning experimental ideas into measurable, maintainable, and production-ready software. Outside the terminal, chess, binge-watching, and travelling are my favourite ways to disappear — always curious about new ideas, unfamiliar places, and things I haven't figured out yet.

More about me

WHAT I DO

01

LLM Applications & RAG Systems

Designing retrieval-augmented generation pipelines and intelligent language-model applications that handle real-world input, context, and failure modes.

02

Evaluation Pipelines & ML Infrastructure

Building structured evaluation frameworks that measure model behaviour, retrieval quality, and system reliability across repeatable test scenarios.

03

Production AI Workflows & Cloud

Taking AI prototypes to production — event-driven pipelines, failure handling, monitoring, and cloud deployment on AWS and similar platforms.

04

Developer Tools & Platforms

Creating tools, dashboards, and platforms that help other engineers build, evaluate, and ship AI systems with confidence.

EXPERIENCE

Lamatic.ai logoLamatic.ai | Applied AI Intern

Jan 2026 – May 2026
  • Engineered AI/ML product workflows and evaluation pipelines for LLM applications, structuring repeatable evaluation paths to assess model behavior and reliability.
  • Built event-driven integrations between AI workflows and application systems, automating multi-step execution paths and reducing manual intervention across workflows.
  • Developed production-oriented components for AI applications with an emphasis on reliable execution, failure handling, and maintainable integration boundaries.
  • Contributed fixes and feature improvements to LangChain and Metaflow, extending practical experience from production AI systems into open-source infrastructure.

Hasprana logoHasprana Health Care Soln Pvt Ltd | AI Engineer

Jul 2026 – Present
  • Engineer AI systems spanning RAG retrieval, model evaluation, and production ML components, taking features from implementation through production-ready integration.
  • Design evaluation workflows for LLM applications that assess retrieval quality, model behavior, and system reliability across repeatable evaluation scenarios.
  • Build and refine AI service components with attention to reliability, maintainability, failure handling, and practical production constraints.
  • Audit and review AI-service code for correctness, maintainability, and production readiness, identifying implementation issues before they propagate into downstream workflows.

SKILLS

  • Python
  • Java
  • JavaScript
  • TypeScript
  • Rust
  • C
  • C++
  • Python
  • Java
  • JavaScript
  • TypeScript
  • Rust
  • C
  • C++
  • PyTorch
  • TensorFlow
  • scikit-learn
  • NumPy
  • Pandas
  • LangChain
  • LangGraph
  • PyTorch
  • TensorFlow
  • scikit-learn
  • NumPy
  • Pandas
  • LangChain
  • LangGraph
  • Transformers
  • Hugging Face
  • MLflow
  • OpenCV
  • FastAPI
  • Next.js
  • Streamlit
  • Transformers
  • Hugging Face
  • MLflow
  • OpenCV
  • FastAPI
  • Next.js
  • Streamlit
  • PostgreSQL
  • SQL
  • Redis
  • Temporal
  • Docker
  • AWS
  • Git
  • GitHub
  • GitHub Actions
  • Linux
  • REST APIs
  • pytest
  • Ollama
  • Jupyter
  • OpenCode
  • PostgreSQL
  • SQL
  • Redis
  • Temporal
  • Docker
  • AWS
  • Git
  • GitHub
  • GitHub Actions
  • Linux
  • REST APIs
  • pytest
  • Ollama
  • Jupyter
  • OpenCode

PROJECTS

01

Kairos

IN-PROGRESS

Open-source platform for transparent RAG development, evaluation, experimentation, and explainable AI.

  • TypeScript
  • RAG
  • Explainable AI
02

RedOps

IN-PROGRESS

Production-grade LLM evaluation and red-teaming platform covering hallucination, jailbreak resistance, latency, cost, and token usage across providers.

  • Python
  • LLM Evaluation
  • Docker
More Projects

RESEARCH & PUBLICATIONS

RESEARCH / 2026

Benchmarking Small Language Models for Domain-Specific Question Answering

A Comparative Study of Phi-3-mini, Mistral-7B, and Gemma-2 on SQuAD v2

SLMQUESTION ANSWERINGSQuAD v2PHI-3-MINIMISTRAL-7BGEMMA-2

Benchmarks small language models for domain-specific question answering. Compares Phi-3-mini, Mistral-7B, and Gemma-2 on SQuAD v2 using Exact Match and F1 evaluation, while examining inference time and the generative–extractive mismatch in instruction-tuned models.

01

Compared Phi-3-mini-4k-instruct, Mistral-7B-Instruct-v0.2, and Gemma-2-2b-it under identical zero-shot conditions on SQuAD v2.

02

Evaluated 200 validation examples using Exact Match, F1 score, and average per-query inference time with a fixed seed for reproducibility.

03

The experiments expose a trade-off between QA accuracy and inference latency across the three SLMs, with no single model dominating every measured dimension.

04

The study identifies a generative–extractive mismatch: instruction-tuned models can produce semantically correct full-sentence answers that receive poor EM/F1 scores when evaluated against short extractive spans.

View Research Paper

OPEN SOURCE

CONTRIBUTIONS3 PRs
01 Ollamaollama
OLMo3 Terminal-Chunk Tool Call ParsingPR #18677 ↗

Fixed OLMo3 streaming tool-call parsing when a complete tool call arrives in the terminal chunk, ensuring final `done=true` responses are parsed instead of being emitted as raw assistant content.

GoLLM InferenceStreamingTool Calling
DeepSeek3 Partial Tool-Call Opening TagsPR #18682 ↗

Fixed DeepSeek3 streaming parser handling for tool-call opening tags split across chunks, preserving partial delimiters until they can be reconstructed into a valid tool call.

GoStreamingTool CallingParser Reliability
02 LangChainlangchain-ai
Post-Generation VerificationPR #36480 ↗

Contributed a post-generation verification component that validates LLM outputs through a Runnable-based fact-checking workflow.

PythonLLM ValidationRunnable
03 MetaflowNetflix
Artifact SerializationPR #3061 ↗

Improved error handling around artifact serialization failures, making failure paths more explicit and robust.

PythonML PipelinesError Handling

GITHUB ACTIVITY

LessMore

कर्मण्येवाधिकारस्ते मा फलेषु कदाचन।
मा कर्मफलहेतुर्भूर्मा ते सङ्गोऽस्त्वकर्मणि ॥ २-४७

Your duty is to uphold dharma
Never to claim the reward
Let not the promise of victory guide you
The battlefield summons, be relentless in action

- Bhagavad Gita - Chapter 2, Verse 47