AI agent observability and security
Our analysis
Sentry leads at 26%, ahead of LangSmith at 17%. Established monitoring products may have an advantage as AI observability converges with conventional application monitoring.
Question
Have you used any of the following tools for AI agent observability, prompts, or evals in the past year? Do you want to use any of the following in the next 6 months?
- scale
- Optional
- v2026.1
Data
| Respondents | Percent | |
|---|---|---|
| Sentry | 592 | 26.0% |
| LangSmith | 393 | 17.3% |
| MLflow | 347 | 15.2% |
| Langfuse | 345 | 15.2% |
| Datadog LLM | 271 | 11.9% |
| Weights & Biases | 244 | 10.7% |
| Arize AI | 133 | 5.8% |
| Braintrust | 132 | 5.8% |
| DeepEval | 131 | 5.8% |
| Respan | 127 | 5.6% |
| Promptfoo | 121 | 5.3% |
| Agenta | 114 | 5.0% |
| Ragas | 111 | 4.9% |
| Galileo AI | 110 | 4.8% |
| Portkey | 109 | 4.8% |
| LangWatch | 101 | 4.4% |
| Helicone | 95 | 4.2% |
| Humanloop | 93 | 4.1% |
| PromptLayer | 92 | 4.0% |
| Patronus AI | 91 | 4.0% |
| HoneyHive | 90 | 4.0% |
| Confident AI | 86 | 3.8% |
| Traceloop | 86 | 3.8% |
| Chamber | 85 | 3.7% |
| Opik | 84 | 3.7% |
| Future AGI | 80 | 3.5% |
| Sentrial | 79 | 3.5% |
| Athina AI | 78 | 3.4% |
| Lunary | 77 | 3.4% |
| Moda | 76 | 3.3% |
| Ashr | 70 | 3.1% |
| Parea AI | 69 | 3.0% |
Use this data
Response data is released under the ODbL 1.0, which asks that you attribute it.
Structured for code, with the question, the year and the source url alongside the numbers.
A formatted table, ready to paste into a document, an issue or a prompt.
Spreadsheet
.csv is a plain text format which most spreadsheet software can open.