26 companies
Loading the map…
Services that sit between an application and one or more model providers, offering a single API across many models plus some combination of routing, caching, failover, rate limiting and spend tracking. Several run inference themselves. That is why companies here turn up in the GPU cloud category too.
Compare the model catalogue, whether the service hosts models or brokers out to others, the latency it adds, and what happens when an upstream provider falls over. Billing shapes vary. Some resell provider pricing at a margin, others charge a flat gateway fee on top of your own provider keys.