Pratik Bhavsar's Blog Posts
Pratik Bhavsar is an AI Engineer at Splunk focusing on agent evaluation, reliability, and observability. He has spent a decade building across the full AI stack, from training transformers, building semantic search, and shipping ML systems to designing agentic architectures.
He joined Cisco through the acquisition of Galileo, where he led open-source evaluations and developer relations. He built the Agent Leaderboard, an open benchmark measuring AI agent performance on real-world tasks, and the Hallucination Index, a systematic study of factual reliability across foundation models.
He is the author of five technical books covering eval engineering, agentic systems, RAG, multi-agent architectures, and LLM-as-a-Judge methodologies. Prior to Galileo, Pratik was a founding NLP Scientist at Enterpret and Senior Data Scientist at Morningstar. Pratik holds an M.Tech from IIT Bombay and loves to share his thoughts on Substack.

What Is SME-in-the-Loop in AI Evals?

What Is Context Engineering?

How To Build Deep Research Agents With GPT-5.6 and Evals

How To Tune and Scale LLM Judges: The Complete Guide

Choosing an Embedding Model for RAG: Why Your Custom Evals Beat the Public Leaderboard

Optimizing RAG Retrieval: How to Select the Right Reranking Model

Advanced Chunking Techniques: What The Benchmarks Actually Show

LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)

LLM Judges vs. SLM Judges: When To Use Which

SLM-as-Judge: How to Build and Deploy an SLM Judge

How to Architect an Enterprise Retrieval-Augmented Generation (RAG) System
Can AI Agents Catch Their Own Mistakes? Evaluating for Self-Correction

Token Meter: A Live Cost Meter for Your Coding Agents

Mastering RAG: 8 Failure Modes to Evaluate Before Going to Production

How To Reduce Hallucinations in RAG Applications

How Evals Become Guardrails

Agent Incident Response: How Output Guardrails Help Remediate Agent Failures
