Once again, WebRTC Live broadcasted direct from the RTC.ON conference in Kraków, Poland where WebRTC.ventures CTO Alberto Gonzalez interviewed two speakers:
- Luke Curley, MoQ co-creator at moq.dev, ex Discord, ex Twitch. Luke’s talk, “On a Boat: Why Pull-Based Media Streaming Beats Push,” argues that pull-based protocols like MoQ handle bandwidth-constrained, high-camera-count streaming better than push-based options like WebRTC, SRT, and RTMP, covering multipath QUIC, AI-driven rendition switching, and lossy live plus pristine VOD from a single stream.
- Pratim Mallick, Staff Software Engineer at Stream. Pratim’s talk, “SVC on Android, Debugged,” walks through a debugging journey into why hardware VP9 encoders on Android fail to produce multiple spatial layers as expected, and how Android 16’s new webrtc.svc.l1tN schema addresses the gap.
And one guest:
- Charles Duyk, Principal Engineer on AWS Interactive Video Service (IVS)
WebRTC Live #117: Live from RTC.ON 2026
Episode highlights below the video.
Episode Highlights
Pratim Mallick, Staff Software Engineer at Stream
- Start debugging with one quality score per call. Stream uses a single call-quality metric to decide whether a complaint (black video, bad audio) comes from the publisher, the media server (SFU), or the receiver. From there, the team separates device problems from network problems using WebRTC’s built-in stats and broad telemetry, and tries to reproduce the issue only after that.
- SVC cuts the publisher’s cost of serving mixed networks. Simulcast makes a publisher encode and send several versions of one video, which costs bandwidth and CPU. SVC (scalable video coding) with the VP9 codec sends a single stream with quality layers built in, so a viewer on a large screen and a viewer on a weak mobile connection each receive a suitable resolution. Pratim sees the biggest payoff in one-to-many calls. A one-to-one call between known devices on good networks gains little.
- Android hardware quirks require custom fixes. Stream maintains its own fork of libwebrtc. Early Pixel phones distorted H.264 video at resolutions that weren’t multiples of 16, and some devices advertise hardware echo cancellation they don’t deliver, which can cause echo for every participant in a call. The workaround is a device blocklist that falls back to software processing.
- Production data replaces a full device lab. Android’s spread of OS versions and manufacturers makes exhaustive lab testing impractical. Stream keeps a modest lab and relies on monitoring live data, investigating spikes tied to specific devices.
Charles Duyk, Principal Engineer, AWS Interactive Video Service (IVS)
- Client variability is the hardest problem at scale. IVS offers low-latency HLS and WebRTC, and its WebRTC product targets audiences from hundreds of thousands to millions of concurrent participants. Charles says the toughest challenge is the long tail of client devices and the behaviors customers push them to perform, which regularly break assumptions built into the infrastructure.
- CDN cacheability is the strongest case for MOQ. MOQ is an emerging streaming protocol designed to work with CDNs the way HLS and DASH do. Charles calls this the least glamorous and most important reason it could succeed HLS. Delivering the same cacheability over WebRTC would turn a system into an RTP replay server.
- MoQ’s large spec cuts both ways. WebRTC is a bundle of existing specs (SRTP, DTLS, ICE, STUN), which kept it lightweight but led to fragmentation. MOQ is going the opposite direction with one large spec covering everything. Charles says that could speed adoption or collapse under its own weight, and he expects interoperability work similar to the WHIP and WHEP projects for WebRTC.
Luke Curley, MOQ co-creator (moq.dev)
- MOQ started as a fix for a latency problem at Twitch. Luke’s team had pushed HLS as low as it would go, then built a custom WebRTC stack and SFU. He found WebRTC a poor fit for high-definition broadcast, since Twitch viewers wanted roughly one-second latency and had no use for 100 milliseconds. MOQ grew out of newer browser technologies: WebTransport, QUIC and WebCodecs.
- MOQ works like a live version of HTTP. A viewer asks a CDN for a media track. If the CDN doesn’t have it, the CDN requests it from upstream, and layered caches provide the scale. Luke describes it as pub/sub HTTP, which lets generic CDNs handle live media without a custom SFU.
- Two simple flags give CDNs congestion rules. MOQ splits media into independent groups (each keyframe starts one), and a publisher can attach a priority from 0 to 255 plus an ordering preference to each track. During congestion, the CDN delivers by those rules. This avoids the head-of-line blocking (a late frame stalling everything behind it) that produces buffering with HLS and RTMP.
- Voice AI needs different delivery guarantees than human calls. WebRTC drops late data because humans can’t absorb audio faster than real time. AI agents can, and a lost prompt or response breaks the interaction, which is why most voice AI uses WebSockets. Luke says MOQ exposes the quality-versus-latency tradeoff so a product team can set a rule, such as waiting up to five seconds for a prompt. MOQ is available in the Pipecat voice AI framework.
- WebRTC keeps the conferencing crown. Luke expects WebRTC to remain the best option for conference calls for a long time, given decades of tuning. He wants WebRTC improvements such as combined ICE, DTLS and STUN/TURN handshakes, and playout delay extended beyond video.
Up Next! WebRTC Live #118
Building the Data Layer for Voice AI with vCon
With guest Jeff Pulver, VoIP industry pioneer and founder of the vCon Foundation.
Wednesday, October 7, 12:30 pm Eastern
