If you run WebRTC at scale, a large share of your media goes through TURN relays. Restrictive NATs, enterprise firewalls, and clients that hide their IP addresses all push sessions onto relayed candidates, and when TURN server performance drops under load, call quality drops with it. This post shows how eBPF TURN offload moves ChannelData relay into the Linux kernel so media packets never reach the userspace process. We demonstrate it with a TURN server written in Python and measure per-packet processing time before and after.

The approach builds on the 2023 paper “Supercharge WebRTC: Accelerate TURN Services with eBPF/XDP” by Tamás Lévai, Balázs Kreith, and Gábor Rétvári, presented at the eBPF ’23 workshop. The authors showed how to accelerate TURN with eBPF on top of pion/turn and reported 2-3x improvements in throughput and delay. Offload leaves the Python server exactly as slow as before. It speeds up relay by taking the server off the per-packet path.

I also presented this content at both ClueCon 2026 and the the IIT RTC Conference 2026. See the recording of my ClueCon presentation and read my recap of the RTC conference: WebRTC.ventures Visits IIT RTC Conference 2026: Scale, Latency, and Trust.

Why TURN server CPU usage climbs at scale

Here’s the problem: in a traditional userspace TURN server, every relayed packet crosses the kernel network stack twice:

Diagram of a userspace TURN server relaying a ChannelData packet without eBPF TURN offload. The packet travels from Client A through the kernel network stack (NIC, driver and sk_buff allocation, netfilter and routing, socket buffer), context-switches to the userspace TURN server, then context-switches back and passes through the kernel stack again before reaching Client B.
Without eBPF TURN offload, every relayed packet crosses the kernel network stack twice and triggers two context switches, which limits TURN server performance at scale.

Each traversal means a packet going through:

NIC → driver → sk_buff allocation → netfilter → socket buffer → context switch → userspace → and then back down the whole stack again.

At a few hundred concurrent streams, this is where things start to hurt:

  • CPU saturation from context switches
  • Rising tail latency under load
  • Jitter spikes when the scheduler gets busy.

How eBPF TURN offload works: TC vs XDP

eBPF lets you run small, verified programs inside the Linux kernel, attached to specific hooks, without writing a kernel module. For TURN offload, two attachment points matter.

eBPF/TC (Traffic Control)

The TC hook sits at the traffic control layer, after the kernel allocates an sk_buff for the packet. It has access to packet metadata and can use bpf_redirect() to push packets toward egress. Its big advantage: it works on any NIC, no special driver support needed. You still skip netfilter, routing decisions, and the userspace context switch.

eBPF/XDP (eXpress Data Path)

XDP hooks in even earlier, down at the NIC driver level, before the sk_buff is ever allocated. It works on the raw packet (xdp_md) with zero per-packet memory allocation. It can rewrite headers and send the packet straight back out with XDP_TX (same NIC) or XDP_REDIRECT (another NIC). It’s faster than TC, but it asks for native XDP support in the driver (ENA, mlx5, ixgbe, and friends), which in practice means bare-metal clusters where you control the NIC.

Both hooks can do the actual TURN relay work, because the thing they need to parse is delightfully simple.

The format that makes this easy: ChannelData

Once a WebRTC client binds a channel, it stops wrapping media in bulky STUN messages and switches to TURN’s ChannelData format. Per RFC 8656, it’s just a 4-byte header followed by the payload:

  • 2 bytes: channel number (in the range 0x4000–0x4FFF)
  • 2 bytes: payload length
  • N bytes: the actual media

Here’s the entire serializer from the Python server, and that’s not an abridged version:

def serialize_channel_data(channel: int, payload: bytes) -> bytes:
    """Serialize a ChannelData message.

    Produces a 4-byte header (channel number + payload length) followed by
    the payload. No padding is applied for UDP transport (RFC 8656 §11.5).
    """
    header = struct.pack("!HH", channel, len(payload))
    return header + payload

Telling ChannelData apart from a STUN control message is just as cheap. STUN messages start with a byte in 0x00–0x03; ChannelData starts in the 0x4000–0x4FFF range:

def classify_message(data: bytes) -> Literal["stun", "channel_data", "discard"]:
    first_byte = data[0]
    if first_byte <= 0x03:
        return "stun"
    if len(data) >= 2:
        first_two = struct.unpack("!H", data[:2])[0]
        if 0x4000 <= first_two <= 0x4FFF:
            return "channel_data"
    return "discard"

That’s the whole trick. A few bytes of header tell you everything you need to make a forwarding decision, and that decision is simple enough to make inside the kernel. The choice between TC and XDP is a deployment tradeoff, not a technical limit. BPF maps carry the shared state between userspace and the kernel program either way.

The demo: a deliberately slow TURN server

To prove the point, the demo uses a TURN server written in Python, an interpreted language nobody would pick for line-rate packet pushing. That’s on purpose. If we can make a slow server scale, the lesson transfers to any server.

The setup:

  1. A simple Python TURN server (full RFC 5766 control plane: Allocate, CreatePermission, ChannelBind, ChannelData).
  2. A web client that streams video to the server, which echoes it back, and displays per-packet processing time.
  3. Synthetic load to stress the relay.
  4. We flip on eBPF offload and watch processing time fall off a cliff.
  5. We compare against Coturn under the same load.

Step 1: the userspace path (the slow way)

When offload is off, every media packet rides the full Python path. The server parses the ChannelData header, looks up the channel binding, and forwards the payload through the per-allocation relay socket:

def datagram_received(self, data: bytes, addr: tuple[str, int]) -> None:
    msg_type = classify_message(data)
    if msg_type == "stun":
        self._handle_stun(data, addr)
    elif msg_type == "channel_data":
        self._handle_channel_data_msg(data, addr)
# Forward payload to peer via the relay socket
channel_number, payload = parse_channel_data(data)
# ... look up allocation, look up channel binding, check expiry ...
allocation.relay_transport.sendto(payload, dest_addr)

# Record processing time for the Python path
if t_start is not None:
    elapsed_us = (time.perf_counter_ns() - t_start) / 1_000.0
    self._metrics_state.python_us.append(elapsed_us)

Every byte of that media packet you take crosses into userspace and back. Functionally correct, but it’s paying the full round-trip tax on every single frame.

Step 2: Relay ChannelData with TC redirect and XDP forwarding

The offload itself is a small C program loaded into the kernel at the TC ingress hook. It parses Ethernet → IP → UDP, confirms it’s ChannelData on the TURN port, looks up the forwarding state in a BPF hash map, rewrites the L2/L3/L4 headers with helper functions, and redirects the packet to the egress interface. No socket, no context switch, no round trip through userspace.

First, the forwarding state shared between userspace and the kernel:

/* BPF map key: uniquely identifies a channel binding from a specific client. */
struct channel_key {
    __u16 channel_number;  /* TURN channel (0x4000–0x4FFF) */
    __u16 _pad;
    __be32 src_ip;         /* Client source IP (network byte order) */
    __be16 src_port;       /* Client source port (network byte order) */
    __u16 _pad2;
} __attribute__((packed));

/* BPF map value: everything needed to rewrite headers and redirect. */
struct fwd_entry {
    __be32 dst_ip;            /* Peer destination IP */
    __be32 src_ip;            /* Relay source IP */
    __be16 dst_port;          /* Peer destination UDP port */
    __be16 src_port;          /* Relay source UDP port */
    unsigned char dst_mac[6]; /* Peer destination MAC (next-hop) */
    unsigned char src_mac[6]; /* Source MAC (server's interface MAC) */
    __u32 ifindex;            /* Egress interface for bpf_redirect() */
} __attribute__((packed));

Now the heart of the TC program. Notice how little there is to it: bounds checks for the verifier, a couple of range checks to find ChannelData, one map lookup, and a decision:

int turn_offload(struct __sk_buff *skb) {
    void *data = (void *)(long)skb->data;
    void *data_end = (void *)(long)skb->data_end;

    /* ... parse Ethernet / IPv4 / UDP, bounds-checked ... */

    /* Only process packets destined to the TURN port */
    if (udp->dest != __constant_htons(TURN_PORT))
        return TC_ACT_OK;

    /* ChannelData lives in 0x4000–0x4FFF; STUN starts 0x00–0x03 */
    __u16 first_two_bytes = ntohs(*(__u16 *)udp_payload);
    if (first_two_bytes < CHANNEL_DATA_MIN || first_two_bytes > CHANNEL_DATA_MAX)
        return TC_ACT_OK;

    /* Look up forwarding state for (channel, client IP, client port) */
    struct channel_key key = {};
    key.channel_number = first_two_bytes;
    key.src_ip = ip->saddr;
    key.src_port = udp->source;

    struct fwd_entry *fwd = channel_map.lookup(&key);
    if (!fwd)
        return TC_ACT_OK;   /* not offloaded — let userspace handle it */

If there is no match, TC_ACT_OK passes the packet up to userspace like nothing happened. When there is a match, the program rewrites the packet in-place using TC helper functions that handle sk_buff bookkeeping and incrementally fix the checksum:

  /* Rewrite Ethernet MACs */
    bpf_skb_store_bytes(skb, offsetof(struct ethhdr, h_dest),
                        fwd->dst_mac, 6, 0);
    bpf_skb_store_bytes(skb, offsetof(struct ethhdr, h_source),
                        fwd->src_mac, 6, 0);

    /* Rewrite IP source address + incremental checksum fix */
    bpf_skb_store_bytes(skb, ip_src_offset,
                        &fwd->src_ip, sizeof(fwd->src_ip), 0);
    bpf_l3_csum_replace(skb, ip_csum_offset,
                        old_saddr, fwd->src_ip, sizeof(__be32));

    /* Rewrite IP destination address + incremental checksum fix */
    bpf_skb_store_bytes(skb, ip_dst_offset,
                        &fwd->dst_ip, sizeof(fwd->dst_ip), 0);
    bpf_l3_csum_replace(skb, ip_csum_offset,
                        old_daddr, fwd->dst_ip, sizeof(__be32));

    /* Rewrite UDP ports, zero the checksum (valid for IPv4, RFC 768) */
    bpf_skb_store_bytes(skb, udp_src_offset,
                        &fwd->src_port, sizeof(fwd->src_port), 0);
    bpf_skb_store_bytes(skb, udp_dst_offset,
                        &fwd->dst_port, sizeof(fwd->dst_port), 0);
    __u16 zero_csum = 0;
    bpf_skb_store_bytes(skb, udp_csum_offset,
                        &zero_csum, sizeof(zero_csum), 0);

    /* Redirect to the egress interface — done */
    return bpf_redirect(fwd->ifindex, 0);
}

The XDP version does the same thing at the driver level (before sk_buff allocation) and returns XDP_TX to bounce the packet back out the ingress NIC. Same idea, different door: TC’s bpf_redirect() steers packets to an egress interface, XDP’s XDP_TX sends them back the way they came.

Step 3: wire it up from userspace

Userspace still runs the full TURN control plane, however its only job is to populate the BPF map when a client binds a channel, so the kernel knows where to send things. When a ChannelBind succeeds, the server resolves the next-hop MAC and inserts a forwarding entry:

# --- BPF offload: insert forwarding entry ---
if self._offload_manager is not None:
    peer_mac = self._resolve_peer_mac(dst_client_ip)
    src_mac = self._get_interface_mac()
    self._offload_manager.add_channel_binding(
        channel_number=channel_number,
        client_ip=addr[0],
        client_port=addr[1],
        peer_ip=dst_client_ip,
        peer_port=dst_client_port,
        peer_mac=peer_mac,
        src_mac=src_mac,
        relay_port=self._port,   # TURN server port
        relay_ip=relay_ip,
    )

Under the hood, add_channel_binding just packs a ChannelKey and FwdEntry into the shared hash map. From that moment on, the kernel handles every ChannelData packet for that channel:

key = ChannelKey()
key.channel_number = channel_number
key.src_ip = struct.unpack("<I", socket.inet_aton(client_ip))[0]
key.src_port = socket.htons(client_port)

value = FwdEntry()
value.dst_ip = struct.unpack("<I", socket.inet_aton(peer_ip))[0]
value.src_ip = struct.unpack("<I", socket.inet_aton(relay_ip))[0]
value.dst_port = socket.htons(peer_port)
value.src_port = socket.htons(relay_port)
# ... copy MAC addresses, set ifindex ...
self._channel_map[key] = value

Picking TC or XDP is a single argument. The loader gives both modes the same interface, so the rest of your code doesn’t care which hook is doing the work:

class BPFLoader:
    """Unified BPF loader supporting TC and XDP attachment modes."""
    VALID_MODES = ("tc", "xdp")

    def start(self) -> None:
        if self._mode == "tc":
            self._start_tc()   # bpf_redirect() at the TC ingress hook
        else:
            self._start_xdp()  # XDP_TX in native driver mode

Step 4: measure it

The kernel program samples its own per-packet processing time with bpf_ktime_get_ns() and ships it to userspace over a ring buffer, so the before/after comparison is apples-to-apples. Note that this measuring logic does add some overhead on its own, and we include it for the purpose of the post, on a real implementation you want to keep this route as lean as possible.

__u64 start_ns = bpf_ktime_get_ns();
/* ... do the relay work ... */
struct sample_event event = {};
event.processing_ns = bpf_ktime_get_ns() - start_ns;
events.ringbuf_output(&event, sizeof(event), 0);
return bpf_redirect(fwd->ifindex, 0);

The web client shows the result directly. Before offload, processing time tracks the full Python round trip and climbs as load increases. After offload, the same packets are handled in the kernel in a tiny fraction of the time, and they stay flat as you pile on synthetic load. The Python server, the one we admitted was slow, stops being the bottleneck because it’s barely in the path anymore.

Results: TURN server performance before and after offload

It’s worth being precise about the win, because it’s easy to misread.

We did not write a faster TURN server. The Python server is exactly as slow as it always was. What changed is that we eliminated the userspace round trip for the hot path. Media packets that used to climb all the way up to userspace and back now get rewritten and sent on their way down in the kernel, near the metal.

A few things fall out of that:

  • It’s language-agnostic. The eBPF program doesn’t know or care whether your control plane is Python, Go, Rust, or C. The same approach drops in front of an optimized TURN server such as Coturn to maximize performance at higher loads.
  • You choose your hook based on deployment. TC for portability across any NIC and most cloud/Kubernetes setups; XDP for maximum performance on bare metal whose driver supports it.
  • State stays in sync. BPF maps are the single source of truth shared by userspace and kernel, so bindings, permissions, and forwarding entries don’t drift apart.

Every breath your media takes through a relay used to be a round trip. Now, for the packets that matter most, it isn’t.

Where this fits: scaling TURN servers on Kubernetes

This is the kind of deep systems work that pays off precisely when you’ve already done the easy optimizations and you’re still watching CPU climb and costs follow it. If you’re scaling WebRTC, running TURN behind STUNner on Kubernetes (this is available as a premium feature), or just tired of throwing more relay nodes at a problem that’s really an architecture problem, kernel offload changes the math.

At WebRTC.ventures, this is our home turf: designing and building real-time media applications that hold up under real load. If your TURN tier is struggling, we can help you take the round trip out of the hot path, on your stack, in your environment. Contact us today and let’s make it live!

Recent Blog Posts