Pratik Bhavsar's Blog Posts

Pratik Bhavsar

Pratik Bhavsar is an AI Engineer at Splunk focusing on agent evaluation, reliability, and observability. He has spent a decade building across the full AI stack, from training transformers, building semantic search, and shipping ML systems to designing agentic architectures.

He joined Cisco through the acquisition of Galileo, where he led open-source evaluations and developer relations. He built the Agent Leaderboard, an open benchmark measuring AI agent performance on real-world tasks, and the Hallucination Index, a systematic study of factual reliability across foundation models.

He is the author of five technical books covering eval engineering, agentic systems, RAG, multi-agent architectures, and LLM-as-a-Judge methodologies. Prior to Galileo, Pratik was a founding NLP Scientist at Enterpret and Senior Data Scientist at Morningstar. Pratik holds an M.Tech from IIT Bombay and loves to share his thoughts on Substack.

What Is SME-in-the-Loop in AI Evals?
Learn
6 minutes

What Is SME-in-the-Loop in AI Evals?

AI evals aren’t a set-and-forget activity. A big value-add comes from bringing SMEs into the eval loop, and this article shows you how to do exactly that.
What Is Context Engineering?
Learn
7 Minute Read

What Is Context Engineering?

Understand context engineering as a systematic discipline for curating agent inputs to improve reliability, reduce costs, and optimize performance.
How To Build Deep Research Agents With GPT-5.6 and Evals
Learn
6 Minute Read

How To Build Deep Research Agents With GPT-5.6 and Evals

Build a deep research agent with plan-and-execute in LangGraph, GPT-5.6, and Tavily. Plus span-level evals that catch planning and retrieval failures.
How To Tune and Scale LLM Judges: The Complete Guide
Learn
6 Minute Read

How To Tune and Scale LLM Judges: The Complete Guide

A working LLM judge is not a finished judge. Learn how panels of judges, bias, and configuration choices can help you tune and scale LLM judges.
Choosing an Embedding Model for RAG: Why Your Custom Evals Beat the Public Leaderboard
Artificial Intelligence
10 Minute Read

Choosing an Embedding Model for RAG: Why Your Custom Evals Beat the Public Leaderboard

Public embedding leaderboards help shortlist models, but only corpus-specific evaluations confirm fit. Learn how to measure performance.
Optimizing RAG Retrieval: How to Select the Right Reranking Model
Artificial Intelligence
11 Minute Read

Optimizing RAG Retrieval: How to Select the Right Reranking Model

Rerankers optimize retrieval order for production RAG systems; learn how to select reranker architectures for your latency budget, and domain requirements.
Advanced Chunking Techniques: What The Benchmarks Actually Show
Artificial Intelligence
7 Minute Read

Advanced Chunking Techniques: What The Benchmarks Actually Show

Optimize retrieval performance by matching chunking strategies to your document structure, latency budgets, and cost constraints
LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)
Learn
7 Minute Read

LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)

Understand the tradeoffs between LLMs and humans for generative AI evaluation
LLM Judges vs. SLM Judges: When To Use Which
Learn
8 Minutes Read

LLM Judges vs. SLM Judges: When To Use Which

LLM judge vs SLM judge: an SLM costs 10-30x less and runs in 15-150ms, making 100% coverage affordable. See where each wins and when to make the switch.
SLM-as-Judge: How to Build and Deploy an SLM Judge
Learn
7 Minute Read

SLM-as-Judge: How to Build and Deploy an SLM Judge

How to build an SLM judge: See how to size the model, fine-tune with LoRA, validate it, and serve it in production.
How to Architect an Enterprise Retrieval-Augmented Generation (RAG) System
Artificial Intelligence
7 Minute Read

How to Architect an Enterprise Retrieval-Augmented Generation (RAG) System

Learn how to implement an end-to-end enterprise RAG architecture, step by step through the pipeline, by understanding common failures.
Can AI Agents Catch Their Own Mistakes? Evaluating for Self-Correction
Artificial Intelligence
12 minutes

Can AI Agents Catch Their Own Mistakes? Evaluating for Self-Correction

Evaluate the effectiveness of AI agent self-reflection to determine when internal revision passes actually improve accuracy versus when they introduce new errors.
Token Meter: A Live Cost Meter for Your Coding Agents
Artificial Intelligence
8 Minute Read

Token Meter: A Live Cost Meter for Your Coding Agents

See what your coding agents cost before the bill arrives. Token Meter tracks Claude Code, Codex, and Cursor usage live — free, open source, and 100% local.
Mastering RAG: 8 Failure Modes to Evaluate Before Going to Production
Artificial Intelligence
6 Minute Read

Mastering RAG: 8 Failure Modes to Evaluate Before Going to Production

Review the eight critical RAG failure modes and learn how to build a robust evaluation
How To Reduce Hallucinations in RAG Applications
Artificial Intelligence
15 Minute Read

How To Reduce Hallucinations in RAG Applications

Learn to diagnose the six primary hallucination patterns and implement a robust pipeline of retrieval, reranking, and guardrail enforcement to ensure reliability.
How Evals Become Guardrails
Learn
7 Minute Read

How Evals Become Guardrails

Learn the eval-to-guardrail lifecycle to transform offline evaluation criteria into runtime policies that intercept and block agent failures in production.
Agent Incident Response: How Output Guardrails Help Remediate Agent Failures
Observability
8 Minute Read

Agent Incident Response: How Output Guardrails Help Remediate Agent Failures

Use real-time output guardrails and remediation strategies to surgically contain AI agent failures before they reach your users.
Operationalizing Your Vector Database for Production RAG
Artificial Intelligence
5 Minute Read

Operationalizing Your Vector Database for Production RAG

Learn how to maintain production RAG performance by monitoring filter selectivity, tuning hybrid search parameters, and tracking retrieval drift.