SnagTrace | Deterministic AI Agent Fault Detection Library
SnagTrace is an open-source Python library that helps developers identify common failures in AI agent execution traces using deterministic analysis. Instead of relying on another language model to judge whether an agent failed, it detects structural issues such as execution loops, invalid tool arguments, and cost budget overruns, then reports the exact step where the problem occurred. The library is designed to integrate with modern agent frameworks such as LangChain and LangGraph while remaining lightweight, reproducible, and easy to extend. Overview Debugging AI agents can be difficult because a single execution may involve hundreds of tool calls, multiple agents, and long execution traces. While existing observability platforms provide tracing and replay, developers often still need to manually inspect traces to locate the source of a failure. SnagTrace focuses on failure types that can be detected reliably without an LLM. By analyzing execution traces with deterministic rules, it produces consistent results and helps reduce the time required to investigate common agent failures. Features Key Capabilities Loop Detection Detects repeated tool invocations and recursive execution patterns using configurable sliding-window analysis. Tool Argument Validation Identifies invalid tool inputs, schema mismatches, truncated arguments, and parsing failures before they propagate through an agent workflow. Cost Budget Monitoring Tracks cumulative execution cost and reports the precise step where a configured budget threshold is exceeded. Technologies Design SnagTrace follows a deterministic approach to agent debugging. Each detector analyzes execution traces independently without calling an external language model. This makes diagnoses reproducible, inexpensive to run, and suitable for automation in development and CI workflows. The detector system is modular, allowing additional rule-based detectors to be added without changing the core architecture. Performance The library is designed for efficient analysis of large execution traces while keeping detectors stateless and easy to test. For long-running agents, traces can be analyzed periodically or after execution completes to avoid unnecessary overhead. My Role This project was designed and developed independently. My work included: Outcome SnagTrace provides a lightweight way to detect common AI agent failures without introducing additional LLM-based analysis. It is intended for developers who want deterministic, reproducible diagnostics that can be integrated into local development workflows or automated pipelines.