20 companies
Loading the map…
Speech recognition, speech synthesis, voice cloning, and the orchestration layers that stitch them into a real-time conversation. Component APIs and end-to-end platforms both appear, including the ones built for agents that answer the phone.
For real-time use, latency to first audio beats any quality benchmark. After that: language and accent coverage, streaming in both directions, interruption handling, and for cloning, what consent and provenance controls exist. That last one is turning into a compliance question rather than a feature.