Data Scientist · Generative AI & LLM Engineer

Systems that go to productionand stay there.

I build LLM systems that go to production and stay there — agent architectures, the evaluation stacks that prove they work, and the guardrails that keep them honest.

Michelin / Pune, India / vama0259@gmail.com
Live figures — basis logged
Analysts servedUS and Europe, daily50+
Accuracy500-question benchmark95%
Analyst-hours savedstakeholder-reported27/wk
Cost / querytracked per query in MLflow$0.05–0.20
01

Results, plotted

n = 500-question benchmark unless noted

Accuracy

Basis — 500-question benchmark
Accuracy: 95%. Basis: 500-question benchmark.100095%

Consistency

Basis — run-to-run, same question
Consistency: 85%. Basis: run-to-run, same question.100085%

Component ID

Basis — GPT-4-mini, before and after grounding
Component ID: 30% before, 59–68% after. Basis: GPT-4-mini, before and after grounding.100030%before59–68%after
03

Publications

MLDS 2025

Autonomous web agent for enterprise workflows

First author

Peer-reviewed
MLDS 2024

VAE and GAN models for 3D reconstruction of mechanical parts

Peer-reviewed
04

About

I'm Varun Malhotra, a Data Scientist · Generative AI & LLM Engineer at Michelin in Pune, India. My work sits across three layers of the same problem: agent architecture — deciding what an LLM system is allowed to do and when it should stop and ask; evaluation — building the benchmarks that prove a system works before a user finds out it doesn't; and guardrails — the read-only boundaries and human-in-the-loop gates that keep an autonomous system honest in production.

On Talk to Data I own the agent architecture, LLMOps, security and frontend workstreams — I architect, build, and engineer the system end to end.

Education
B.Tech, Computer Science (Big Data Analytics)
SRM Institute of Science and Technology · 9.25 / 10 GPA · Sep 2020 — Jun 2024

Let's talk shop.

Michelin · Pune, India