Market Opportunity
Unreliable agent behavior - testable harness for LLM tool calls targets a $6.0B = 100,000 engineering teams building agent-enabled features x $60k ACV. Assumes broad developer adoption across startups and enterprises needing observability and test harnesses for LLM agents. total addressable market with medium saturation and a year-over-year growth rate of 40-60% - rapid LLM adoption and agentization of workflows accelerates demand for observability and testing.
Key trends driving demand: Function and tool-calling APIs -- platforms like OpenAI and others added direct tool calling, increasing complexity and brittleness in agent behavior; Agentization of workflows -- more products use multi-step LLM agents, raising frequency of regressions and need for harnesses; Shift from manual to automated QA -- developers expect CI and automated tests, creating demand for testable LLM contracts; Observability demand for LLMs -- teams want tracing, attribution, and diffs of prompt and tool usage to debug agent flows.
Key competitors include LangChain Labs - LangSmith, PromptLayer, Guardrails (open-source and commercial variants), Homegrown tests and manual debugging.