Imagine you are standing in a data center roughly the size of three football fields. Around you, 32,768 NVIDIA H100 GPUs are humming, consuming enough power to run a medium-sized city. You are currently training a trillion-parameter Large Language Model (LLM). The bill for this single training run? North of $50 million.
Then, the "Tail Latency Monster" strikes.
Somewhere in the fabric, a single switch buffer overflows. A burst of packets from a "noisy neighbor" checkpointing their job hits the same egress port as your critical All-Reduce operation. The network triggers Priority Flow Control (PFC). A "PAUSE" frame ripples back through the leaf-spine architecture. For 200 microseconds, a whole pod of GPUs sits idle. At this scale, 200 microseconds of idle time across 32k GPUs isn't just a glitch—it’s a massive waste of capital.
For years, the industry answer was simple: Buy InfiniBand. It was the only way to get the lossless, low-latency, credit-based flow control required for massive GPU clusters. But as we move into the era of 100k+ GPU clusters and multi-tenant AI clouds, the "InfiniBand Tax" and its proprietary ecosystem are becoming bottlenecks.
Enter RoCE v2 (RDMA over Converged Ethernet).
The promise is intoxicating: InfiniBand-level performance on standard, commodity-ish Ethernet hardware. But the reality is a brutal engineering challenge. You can't just plug in a 400G switch and hope for the best. To make RoCE work for multi-tenant LLM orchestration, you have to solve the "Lossless Ethernet" paradox.
Let’s go down the rabbit hole.
The Physics of the Problem: Why Ethernet is "Broken" for AI
Ethernet was born in a world of "Best Effort." If a packet gets dropped, TCP will eventually notice and retransmit it. This is fine for Netflix; it is catastrophic for an LLM.
In distributed LLM training, we rely on Collective Communications (All-Reduce, All-to-All, Reduce-Scatter). If one GPU out of 4,096 is slow because it’s waiting for a retransmitted packet, the entire synchronous SGD (Stochastic Gradient Descent) step waits. Your effective TFLOPS drop off a cliff.
RDMA (Remote Direct Memory Access) solves this by bypassing the CPU kernel entirely. It allows one GPU to write directly into the memory of another GPU across the network. RoCE v2 wraps these RDMA InfiniBand packets inside UDP/IP/Ethernet headers.
But here is the catch: RDMA assumes a lossless fabric. If an RDMA packet is dropped in a standard Ethernet environment, the Go-Back-N retransmission mechanism kicks in, which is incredibly inefficient. To make RoCE viable, we have to turn "Best Effort" Ethernet into a "Lossless" fabric.
The Pillars of the RoCE Stack: PFC and ECN
To achieve losslessness, we rely on two primary mechanisms that are notoriously difficult to tune at hyperscale.
1. Priority Flow Control (PFC)
PFC is the "Stop Light." When a switch buffer fills up to a certain threshold, it sends a PFC PAUSE frame back to the sender.
- The Engineering Challenge: If misconfigured, PFC leads to Head-of-Line (HoL) Blocking. If Job A is congesting the network, it can trigger a PAUSE that stops Job B, even if Job B's path is clear. Worse, in a circular dependency, you get PFC Deadlocks, where the whole network just... stops.
2. Explicit Congestion Notification (ECN)
ECN is the "Slow Down" signal. Instead of waiting for a buffer to overflow and triggering a PAUSE, the switch marks the IP header of packets (using the ECN bits) when buffers start to get slightly full. When the receiver sees these bits, it tells the sender to throttle its injection rate.
- The Tuning Nightmare: This is governed by the DCQCN (Data Center Quantized Congestion Control) algorithm. It involves a dozen variables:
Target Rate,Alpha,Beta,G,W, and timers. If you’re too aggressive, you underutilize the 400G link. If you’re too slow, you trigger PFC anyway.
Multi-Tenancy: When "Noisy Neighbors" become "Deadly Neighbors"
In a hyperscale environment, you aren't just running one LLM job. You have:
- Job A: A 175B parameter training run (High throughput, latency-sensitive).
- Job B: A 7B parameter fine-tuning job (Burstier).
- Job C: A storage-heavy checkpointing process (Massive elephant flows).
In a naive Ethernet setup, Job C will destroy Job A’s performance. To solve this in a multi-tenant GPU orchestrator, we have to implement Traffic Class Isolation.
DSCP Mapping and Virtual Lanes
We use DSCP (Differentiated Services Code Point) values to map different types of traffic into different hardware queues (Priorities) on the NICs and switches.
# Example: Mapping NCCL traffic to DSCP 46 (Expedited Forwarding)
# and setting the NIC to use Priority 3 for RoCE
export NCCL_IB_GID_INDEX=3
export NCCL_IB_TC=184 # 184 >> 2 = 46 (DSCP)
By isolating RDMA traffic on a specific Priority (e.g., Priority 3) and giving it its own buffer space, we ensure that a massive TCP-based data transfer on Priority 0 won't trigger a PFC PAUSE on our training traffic.
The Architecture: Building the "Lossless" Spine-Leaf
For a 10,000+ GPU cluster, a standard Clos topology is used, but the oversubscription ratio must be 1:1. You cannot have any oversubscription in the fabric if you want to maintain RoCE performance at scale.
Rail-Optimized Design
This is the "secret sauce" used by the likes of Meta and CoreWeave. In a rail-optimized topology, we ensure that all "GPU 0s" across all servers are connected to the same Leaf switches.
- Why? Most LLM communication happens within a "rail." When you do an All-Reduce, GPU 0 talks to GPU 0 on other nodes. By keeping this traffic on the same set of switches, you minimize the number of hops and reduce the probability of cross-job interference.
Adaptive Routing: The Holy Grail
In standard Ethernet, we use ECMP (Equal-Cost Multi-Path). ECMP hashes a flow (based on IP/Port) to a specific path. If two "elephant flows" happen to hash to the same physical link, that link gets congested while others sit idle. This is the "Incast" problem.
Modern RoCE deployments (using Broadcom Tomahawk 5 or NVIDIA Spectrum-4 ASICs) use Adaptive Routing (AR). The switch hardware monitors the load on its egress ports. If one link is congested, it dynamically reroutes packets of the same flow to an underutilized link.
The Catch: Adaptive Routing can cause packets to arrive out of order. RDMA originally hated this. Modern NICs (like the ConnectX-7) now support Hardware Out-of-Order Placement, which reassembles the data in the NIC's hardware before handing it to the GPU memory. This is a massive leap forward for Ethernet.
Orchestration: Making Kubernetes RoCE-Aware
You can’t just use a standard K8s CNI and expect RoCE to work. You need a specialized stack.
1. The Secondary Network
We never run RoCE over the primary K8s interface (which handles logs, health checks, etc.). We use Multus CNI to attach multiple interfaces to the Pod.
- eth0: Management/Control plane.
- net1-net8: The high-speed RoCE interfaces (one per GPU).
2. Topology-Aware Scheduling
A standard K8s scheduler might place half of your GPUs on Rack A and the other half on Rack Z. The latency penalty of crossing the spine multiple times will kill your training performance.
We use a custom Topology-Aware Scheduler that understands the physical layout of the network. It attempts to bin-pack GPU jobs into "Network Islands"—groups of racks connected to the same cluster of aggregation switches.
3. NCCL Tuning: The Software Glue
The NVIDIA Collective Communications Library (NCCL) is the library that actually moves the data. Tuning NCCL for RoCE is an art form.
# A typical environment config for a RoCE-based LLM Job
env:
- name: NCCL_DEBUG
value: "INFO"
- name: NCCL_IB_HCA
value: "=mlx5_0,mlx5_1,mlx5_2,mlx5_3" # Use specific RoCE HCAs
- name: NCCL_IB_GID_INDEX
value: "3" # Match the RoCE v2 GID
- name: NCCL_IB_TIMEOUT
value: "22" # Adjust for network retries
- name: NCCL_ALGO
value: "Ring" # Or "Tree" depending on the cluster size
The "Invisible" Killer: Soft Failures
In a 100k GPU cluster, things break every hour. But "hard failures" (a dead switch) are easy to detect. "Soft failures" are the nightmare.
Imagine a fiber optic cable that is slightly bent. It’s dropping 0.001% of packets. In a standard network, you’d never notice. In a RoCE fabric, this triggers PFC. Because of the way PFC works, that one bad cable can "back-pressure" the entire spine, slowing down thousands of GPUs that have nothing to do with that cable.
Telemetry at the Edge
To solve this, we implement Streaming Telemetry. We pull counters from every switch and NIC every second:
pfc_requests_tx: Is this switch telling others to stop?pfc_requests_rx: Is this switch being told to stop?ecn_marked_packets: Is congestion building up?
By correlating these metrics, our orchestration engine can automatically quarantine a node or a link. If we see a spike in PFC frames on a specific port without a corresponding spike in throughput, we know we have a "Gray Failure" and the scheduler automatically drains that node.
The Rise of the Ultra Ethernet Consortium (UEC)
If you think this sounds complicated, you’re right. This complexity is exactly why the Ultra Ethernet Consortium (UEC) was formed. The goal of UEC (backed by AMD, Broadcom, Google, Meta, and others) is to evolve Ethernet into something that handles RDMA natively, without the brittle mess of PFC and DCQCN.
The UEC is working on a new transport protocol that:
- Allows Packet Spraying (sending packets of the same flow across all available paths) natively.
- Handles Out-of-Order delivery at the transport layer.
- Implements Grant-based flow control (similar to InfiniBand's credit system) rather than the "Stop/Start" reactive nature of PFC.
We are currently in a transition period. We are using "Hackable Ethernet" (RoCE v2) to bridge the gap between the InfiniBand era and the UEC era.
Practical Engineering: The Checkpoint Problem
Let’s talk about a real-world scenario we encountered. We had a 2,048 GPU cluster training a Llama-3 variant. Every 2 hours, the training would stall for 4 minutes.
The Diagnosis: The "Checkpointing" process was saving the model weights to an S3-compatible object store. The storage traffic was using the same physical links as the RoCE training traffic. Even though they were on different Virtual Lanes (VLs), the total bandwidth of the 400G link was being saturated.
The switch, trying to be helpful, used its shared buffer to absorb the burst. But the buffer filled up. PFC kicked in on all VLs because the switch's shared buffer pool was exhausted.
The Fix:
- Strict Buffer Reservation: We reconfigured the switch buffer allocation to reserve 30% of the buffer exclusively for the RoCE Priority, preventing the Storage Priority from ever touching it.
- Egress Shaping: We limited the storage traffic at the NIC level to 50 Gbps, ensuring it could never saturate the 400G link and cause a "micro-burst" that would trigger PFC.
# Limiting the storage interface (eth0) to prevent RoCE interference
tc qdisc add dev eth0 root handle 1: htb default 11
tc class add dev eth0 parent 1: classid 1:1 htb rate 50Gbit
The Economics of RoCE at Scale
Why go through all this pain? Why not just pay NVIDIA for InfiniBand?
- Vendor Lock-in: InfiniBand is essentially a single-vendor ecosystem. With RoCE, you can mix and match switches (Arista, Cisco, Broadcom-based whitebox) and NICs (NVIDIA, AMD/Pensando, Intel).
- Unified Fabric: Large labs don't just have GPUs. They have CPU clusters for data preprocessing and storage clusters. Running one unified Ethernet fabric for everything reduces OpEx significantly.
- Cost: At the 100k GPU scale, the savings of Ethernet over InfiniBand can be measured in the hundreds of millions of dollars. That’s more budget for more GPUs.
Summary of the Engineering Playbook
If you are tasked with building a multi-tenant RoCE fabric for LLM training, here is your checklist:
- Fabric: Use a non-blocking 1:1 Spine-Leaf topology.
- Routing: Enable Adaptive Routing (AR) and ensure your NICs support Hardware Out-of-Order placement.
- Isolation: Map RDMA/NCCL traffic to a dedicated DSCP/Priority. Use PFC for losslessness but keep the thresholds tight.
- Congestion Control: Implement DCQCN and prepare to spend weeks tuning the
AlphaandBetaparameters for your specific cable lengths and hop counts. - Orchestration: Use a topology-aware Kubernetes scheduler. Do not let the scheduler treat the network as a "black box."
- Monitoring: Treat PFC frames as an "Error" condition. Any sustained PFC activity is a sign of a misconfigured fabric or a failing component.
The path to RoCE at hyperscale is paved with buffer tuning and telemetry, but it is the only way to build the massive, open AI infrastructure of the future. The "InfiniBand Moat" is shrinking, and the future of the AI data center is looking remarkably like... really, really well-tuned Ethernet.
Happy scaling.
