Jackson Wells' Blog Posts
Jackson Wells is a writer and marketer at Splunk focused on how teams evaluate and monitor AI agents in production. He came to Splunk through the Galileo acquisition, and before that spent four years writing for developer audiences at LinearB. His work covers evals, guardrails, and what it takes to trust a AI systems in production.
Display Mode
Paginated
Filter
Author
Author URL
Limit
18

OpenAI Swarm Multi-Agent Orchestration: The Complete Engineering Guide
Learn how to build production-ready multi-agent systems using the OpenAI Agents SDK, focusing on clear role boundaries and effective handoff orchestration.

Top Agent Frameworks To Use in 2026: A Comparison
Compare top agent frameworks like LangGraph and CrewAI based on operational resilience, durable execution, and production-grade observability.

Measuring Production Agent Performance: A Deep Dive Into How Elite Teams Measure Agent Performance
Learn how to measure autonomous agent performance with this guide. Elite teams layer observability in to isolate workflow failures and improve production reliability.

F1 Score in AI Evaluation: How to Balance Precision and Recall
F1 score provides a more reliable performance metric than accuracy for imbalanced AI tasks, such as toxicity detection and safety classification.

10 Common Hallucinations: When Bad AI Impacts Trust and Revenue
Hallucinations in autonomous agents create business liability. Learn how teams can verify model outputs against certified source data, in a systematic way.

7 LLM Metrics to Enhance AI Reliability
Measure LLM reliability and performance using seven key metrics to isolate system health from generation quality.

AI Agent Systems: A Field Guide To Types and Levels
This field guide categorizes AI agents across seven levels of autonomy, providing a roadmap for mapping system architectures against common failure modes and production-ready requirements.

AI Accuracy Explained and How to Improve It
Measure and improve AI accuracy by distinguishing between classification metrics, generative faithfulness, and agentic trajectory completion.

Best Practices for AI Model Validation in Machine Learning
Learn how to validate AI models across the full lifecycle. From development evals to runtime guardrails, close the gap between benchmarks and production behavior.

The AI Production Readiness Checklist: Infrastructure Requirements for Enterprise AI
Enterprise AI requires core components—including data pipelines, security, and observability—to successfully move agentic applications from demo to production.

Benchmarks for Multi-Agent AI Systems
Evaluate multi-agent AI systems using benchmarks that prioritize coordination, reliability, and policy adherence over simple accuracy scores to ensure production readiness.

How to Evaluate AI Systems
Explore a detailed step-by-step process on effectively evaluating AI systems to boost their potential.

LLM Benchmarks: Top Categories for Evaluating AI Beyond Conventional Metrics
Evaluating LLMs requires moving beyond general metrics to domain-specific, agentic, and adversarial benchmarks that reflect real-world performance.

How On-Premises AI Observability Works, and Why Regulated Enterprises Need It
How on-premises AI observability works — what agent telemetry contains, why evaluation is the hard part to keep behind the firewall, and how the deployment models compare.

How To Test AI Agents Effectively
Unlock the key to AI agent testing with our guide. Discover metrics, best practices, and innovative techniques to evaluate your AI agents.

How Custom AI Evaluation Metrics Close the Confidence Gap
Learn how custom AI evaluation metrics outperform standard benchmarks to ensure GenAI reliability, compliance, and business alignment for confident deployment.

What Is BERTScore and How Does It Work for NLP Evaluation?
Discover BERTScore’s transformative role in AI, offering nuanced and context-aware evaluation for NLP tasks, surpassing traditional metrics.

Practical AI in Business: How Companies Capture Real Value in 2026
Practical AI delivers business value by connecting agents to measurable outcomes like revenue and cost savings, supported by observability and incident response infrastructure.