
On September 22–23, I was in Chicago for the IIT Real-Time Communications Conference at Illinois Tech, an annual IEEE conference where industry engineers and academic researchers present work on real-time communications. This year’s tracks covered programmable networks, WebRTC applications, emergency communications, and AI in real-time communications. I

A production voice agent has a latency budget of roughly 800 milliseconds before a conversation starts to feel slow. That budget is spent across eight sequential stages, typically split across three to five different vendors depending on how the stack is assembled. This post maps that voice

Self-hosting a voice AI pipeline gives you control that managed APIs can’t offer: you choose the models, you decide where they run relative to each other, and you keep audio data within your own infrastructure. Getting there means managing GPU workloads, model placement, and service networking yourself.

Now that everyone’s back from summer break here in Madrid, it feels like the right time to recap AI Tinkerers Madrid: AI in Real-Time Communication Systems, which WebRTC.ventures sponsored and hosted this past July alongside BTS, who provided their offices and a great spread of local snacks.