19 companies
Loading the map…
Tooling for seeing what a model-backed application actually did in production: tracing calls, logging prompts and responses, scoring output quality, and catching regressions between versions. Some entries are dedicated LLM tools. Others are established observability vendors that bolted model-specific features onto what they already had.
Check the tracing integration first. If it does not fit the framework you already run, nothing else about the product matters. Then compare whether evaluation happens offline, online or both, whether scoring is automated, human or model-graded, and how prompts and responses get stored, because this category handles the most sensitive data of any tooling layer.