---
title: "AI agent observability and security — AI data 2026"
url: https://survey.stackoverflow.co/2026/ai/data/ai-obs
year: 2026
chapter: "ai"
question: "AIObs"
section: "AI agent tools"
licence: "ODbL 1.0 (https://opendatacommons.org/licenses/odbl/1-0/)"
---

# AI agent observability and security

Sentry leads at 26%, ahead of LangSmith at 17%. Established monitoring products may have an advantage as AI observability converges with conventional application monitoring.

Asked as: Have you used any of the following tools for AI agent observability, prompts, or evals in the past year? Do you want to use any of the following in the next 6 months? (`AIObs`, scale, optional, v2026.1)

## Options offered

1. Respan
2. LangSmith
3. Weights & Biases
4. MLflow
5. Langfuse
6. Arize AI
7. Helicone
8. Traceloop
9. Datadog LLM
10. HoneyHive
11. Braintrust
12. Promptfoo
13. Patronus AI
14. Humanloop
15. Portkey
16. Sentry
17. DeepEval
18. Ragas
19. Galileo AI
20. LangWatch
21. PromptLayer
22. Confident AI
23. Opik
24. Agenta
25. Future AGI
26. Lunary
27. Parea AI
28. Sentrial
29. Athina AI
30. Moda
31. Ashr
32. Chamber

## Results

### Used

n = 2,277

|  | Respondents | Percent |
| --- | ---: | ---: |
| Sentry | 592 | 26.0% |
| LangSmith | 393 | 17.3% |
| MLflow | 347 | 15.2% |
| Langfuse | 345 | 15.2% |
| Datadog LLM | 271 | 11.9% |
| Weights &amp; Biases | 244 | 10.7% |
| Arize AI | 133 | 5.8% |
| Braintrust | 132 | 5.8% |
| DeepEval | 131 | 5.8% |
| Respan | 127 | 5.6% |
| Promptfoo | 121 | 5.3% |
| Agenta | 114 | 5.0% |
| Ragas | 111 | 4.9% |
| Galileo AI | 110 | 4.8% |
| Portkey | 109 | 4.8% |
| LangWatch | 101 | 4.4% |
| Helicone | 95 | 4.2% |
| Humanloop | 93 | 4.1% |
| PromptLayer | 92 | 4.0% |
| Patronus AI | 91 | 4.0% |
| HoneyHive | 90 | 4.0% |
| Confident AI | 86 | 3.8% |
| Traceloop | 86 | 3.8% |
| Chamber | 85 | 3.7% |
| Opik | 84 | 3.7% |
| Future AGI | 80 | 3.5% |
| Sentrial | 79 | 3.5% |
| Athina AI | 78 | 3.4% |
| Lunary | 77 | 3.4% |
| Moda | 76 | 3.3% |
| Ashr | 70 | 3.1% |
| Parea AI | 69 | 3.0% |

### Want to use

n = 2,277

|  | Respondents | Percent |
| --- | ---: | ---: |
| Sentry | 551 | 24.2% |
| LangSmith | 478 | 21.0% |
| MLflow | 442 | 19.4% |
| Langfuse | 445 | 19.5% |
| Datadog LLM | 445 | 19.5% |
| Weights &amp; Biases | 357 | 15.7% |
| Arize AI | 274 | 12.0% |
| Braintrust | 282 | 12.4% |
| DeepEval | 302 | 13.3% |
| Respan | 299 | 13.1% |
| Promptfoo | 277 | 12.2% |
| Agenta | 270 | 11.9% |
| Ragas | 277 | 12.2% |
| Galileo AI | 271 | 11.9% |
| Portkey | 264 | 11.6% |
| LangWatch | 285 | 12.5% |
| Helicone | 260 | 11.4% |
| Humanloop | 269 | 11.8% |
| PromptLayer | 254 | 11.2% |
| Patronus AI | 254 | 11.2% |
| HoneyHive | 265 | 11.6% |
| Confident AI | 261 | 11.5% |
| Traceloop | 256 | 11.2% |
| Chamber | 266 | 11.7% |
| Opik | 264 | 11.6% |
| Future AGI | 261 | 11.5% |
| Sentrial | 246 | 10.8% |
| Athina AI | 243 | 10.7% |
| Lunary | 251 | 11.0% |
| Moda | 245 | 10.8% |
| Ashr | 243 | 10.7% |
| Parea AI | 237 | 10.4% |

In context: https://survey.stackoverflow.co/2026/ai/data.md
