14 min read

The Cold Start Illusion: Optimizing Kernel-Level Context Switching and Cache Locality for Multi-Tenant Serverless

Optimizing Kernel Context and Cache Locality for Multi-Tenant Serverless

There's a number you need to see. It's been staring at us from dashboards in AWS, Cloudflare, and every hyperscaler on the planet for the better part of a decade: the 200ms p99 cold start.

You've probably seen the marketing slide. Some engineer at a conference waves a hand and says, "Serverless is instant." And then you go deploy a real function—one that opens a database connection, loads a few hundred megabytes of dependencies, and touches a couple of shared libraries—and it lags like a 2011 Android phone waking from a deep sleep.

For years, the industry blamed the runtime: "Oh, it's Node.js boot time," or "Just use Go, it starts in 20ms." But if you've ever burned a week profiling a cold start under a flame graph, you know the truth. The runtime is often the least of your problems. The real bottleneck—the gremlin hiding in the kernel scheduler, the TLB, and the last-level cache—is the context switch.

Today, we're going to go deep. Not the "we use Firecracker, isn't it cool?" surface-level deep. We're talking about the CPU microarchitecture, the x86_64 CR3 register, KPTI, huge pages, cache coloring, and why your perfectly good serverless platform probably has p99 tail latencies that are being murdered by a tenant named customer-xyz-7781 running a memory copy loop on the next core over.

Let's go.


The Hype Cycle and the Real Substrate

Let's set the scene. Around 2018, "Serverless" was the shiny new thing. AWS Lambda was proliferating, and the industry narrative was all about abstracting away the infrastructure. The pitch: you just upload code, and the cloud figures out the rest. Billions of dollars were thrown at it.

Then reality hit. The hype said serverless was for "event-driven microservices." The substance—which we discovered when we tried to run real production workloads—was that multi-tenancy is a mean, nasty problem at the kernel level.

The industry responded with microVMs. Firecracker, gVisor, Kata Containers. AWS built Firecracker to run 150 microVMs per second per host. Cloudflare built Workers on V8 isolates to avoid the hypervisor tax entirely. The hype said, "Look, hardware virtualization is fast!"

But here's the thing: Hardware virtualization solves isolation, not performance. A microVM still has to be scheduled by the host kernel. It still has to context switch. It still has to wipe the TLB when a tenant changes. And when you cram 1,000 tenants onto a single 64-core box, the context switching becomes the tax you pay for density.

This is the part where the marketing pivots to "We have 50ms cold starts!"—and every engineer who has ever looked at perf top on a busy hypervisor nods slowly and says, "Sure, for trivial workloads."

We're here to talk about the workloads that aren't trivial. The ones where the cache is the kingdom, and the scheduler is the king.


The Anatomy of a Context Switch (And Why It Hurts Serverless)

Let's get our hands dirty. I want you to picture the CPU. You have your registers, your L1/L2 caches (private to the core), your L3 cache (shared), and your Translation Lookaside Buffer (TLB). The TLB is the unsung hero. It caches virtual-to-physical address translations so your MMU doesn't have to walk a 4-level page table on every single memory access.

Now, a serverless platform runs thousands of these execution contexts. Some are microVMs. Some are containers. Some are WASM sandboxes.

When the scheduler decides it's time to switch from Tenant A to Tenant B, look at what happens:

  1. The Register File Flush: The CPU has to save the current state (general purpose registers, FPU state, SIMD registers) and load the new state. That's expensive, but manageable.
  2. The TLB Shootdown: If you're using a hypervisor, or if you're running isolated processes, the address space changes. If the kernel is using KPTI (Kernel Page Table Isolation)—which it must for security on modern x86—you get a full TLB flush. That's potentially thousands of cycles just to re-warm the TLB.
  3. The Cache Invalidation: Here's the killer. When Tenant B starts executing, its working set is not in L1 or L2. It's not even in L3. It has to be fetched from main memory. That's a latency cliff of ~100 nanoseconds for an L3 hit versus ~300 nanoseconds for main memory (DDR5, roughly).

Multiply that across a busy scheduler tick. If you're context switching every 10ms (typical for a fair scheduler), and you have 1,000 tenants on a host, you're spending an ungodly percentage of your CPU cycles just repopulating caches and TLBs.

The math on the tax: If a context switch costs 5 microseconds in overhead (optimistic for a microVM with TLB flush), and you switch 100,000 times per second per host, you've burned 0.5 seconds of CPU time per second per host—50% overhead. That's not a rounding error. That's the difference between a profitable region and a money pit.


The Three Pillars of Serverless Kernel Optimization

If we're going to fix this, we need to attack it from three angles:

  1. Scheduler Design: Who runs when, and for how long?
  2. Cache Locality: Keeping the working set of a tenant on the same core.
  3. Memory Management: Reducing TLB pressure via huge pages and avoiding the KPTI tax.

Let's break each one down with some hands-on observations.

Pillar 1: The Scheduler Isn't Your Friend (Unless You Tune It)

The default Linux CFS (Completely Fair Scheduler) is a miracle of engineering. It's designed to be fair to all tasks. But fairness is the enemy of serverless.

If CFS decides to preempt Tenant A's function mid-execution to let Tenant B run, Tenant A loses its cache. When it resumes, it's cold again. This is "cache thrashing," and it's brutal for short-lived functions.

The fix: Serverless platforms are increasingly moving to cooperative scheduling or batch scheduling for micro-tasks. Instead of time-slicing at 10ms, you run the function to completion or until it hits a yield point (I/O, await, etc.).

But here's the twist: You can't just "run to completion" if a malicious tenant writes an infinite loop. So you need a watchdog. You need a timer that preempts, but you want that timer to be long enough to amortize the cache warm-up.

In practice, this looks like a two-tier scheduler:

  • Tier 1 (Cooperative): For short, bursty functions (think <5ms). Let them run without preemption. The kernel stays out of the way.
  • Tier 2 (Preemptive): For long-running or untrusted functions. Use a standard scheduler, but pin them to a "tainted" core pool to avoid polluting the cache of Tier 1.

Code snippet — a simplified scheduling hint in C:

// In a real system, this is handled by the kernel's scheduler policy
// but here's the conceptual idea: set a CPU affinity mask per tenant
struct sched_attr attr = {
    .size = sizeof(attr),
    .sched_policy = SCHED_BATCH, // Not SCHED_OTHER
    .sched_flags = SCHED_FLAG_RESET_ON_FORK,
    .sched_nice = 19, // Lowest priority for batch tasks
};

// Pin the microVM to a specific core to preserve L1/L2
cpu_set_t cpuset;
CPU_ZERO(&cpuset);
CPU_SET(target_core_id, &cpuset);
sched_setaffinity(pid, sizeof(cpuset), &cpuset);
sched_setattr(pid, &attr, 0);

The key insight: Don't let the scheduler move a tenant around. If Tenant A starts on Core 3, keep it on Core 3. The L1 and L2 caches are per-core. Moving the tenant to Core 7 means it has to refill L1 and L2 from scratch.

Does this cause load imbalance? Yes. Do you care? For serverless, tail latency matters more than average throughput. A slightly imbalanced host that's 20% idle but has consistent p99 is better than a perfectly balanced host with wild p99 swings.

Pillar 2: Cache Locality and the "Hot Tenant" Problem

Let's talk about the L3 cache. On modern server CPUs (think AMD EPYC or Intel Xeon Scalable), the L3 is often split into multiple slices (e.g., 4 slices of 16MB on a 64MB L3). The latency to access a remote L3 slice is higher than a local slice.

If your scheduler doesn't understand the NUMA topology and the cache topology, you're going to have a bad time. A tenant's data might be in L3, but if the core reading it is on a different CCX (Core Complex), it's a "far" L3 hit.

The fix: Cache-aware scheduling. You want to map tenants to cores such that their working sets fit within the shared L3 of a CCX, and you want to keep the tenant on the same CCX.

This is where Firecracker's design shines. It's minimal. It doesn't emulate a full PCI bus. It doesn't run a full BIOS. The guest kernel is tiny. That means the guest's memory footprint is smaller, and it's easier to fit into a cache slice.

But even Firecracker isn't magic. If you run a Firecracker microVM with a 512MB working set, and your L3 slice is 16MB, you're going to be streaming from DRAM. The cache isn't helping. At that point, you're memory-bandwidth bound, not cache bound.

The trick: Right-size the tenant. If a function only needs 128MB of RAM, don't give it 512MB. Use memory ballooning or cgroup memory limits to keep the working set tight. The smaller the working set, the more likely it stays in L3.

Pillar 3: TLB, Huge Pages, and KPTI (The Real Boss Fight)

Now for the deepest layer: Memory Management.

When a context switch happens between two different address spaces (different tenants), the CPU's CR3 register is updated. This invalidates the TLB (unless you have PCID—Process Context ID—which is an x86 feature that tags TLB entries with an ASID). Even with PCID, you have a limited number of TLB entries. If you have too many tenants, you thrash the TLB.

TLB coverage math: A standard 4KB page TLB might have 1,500 entries. That covers 6MB of memory. If your tenant's working set is 100MB, you're missing the TLB 94% of the time. Every miss is a page walk. Every page walk is 4 memory accesses (on a 4-level page table).

The fix: Huge Pages.

If you use 2MB pages, your TLB coverage jumps from 6MB to 3GB (with the same number of entries). That's a massive reduction in page walks.

But here's the catch: Huge pages are hard to manage in a multi-tenant environment. You can't just hand out 2MB pages to every tenant because they lead to internal fragmentation. A tenant that only needs 1KB of memory will waste 2MB.

The modern approach: Transparent Huge Pages (THP) with madvise. You tell the kernel, "Hey, this region is going to be heavily used. Please back it with huge pages." The kernel does its best, but it's not guaranteed.

Code snippet — enabling THP hints in a guest:

// Map a region with huge page hints
void *addr = mmap(NULL, size, PROT_READ | PROT_WRITE,
                  MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
madvise(addr, size, MADV_HUGEPAGE); // Hint the kernel to use 2MB pages

The KPTI tax: On x86, the Meltdown vulnerability forced the kernel to implement KPTI. This means when you transition from user space to kernel space, you swap page tables. That's a full TLB flush (unless you use PCID). For a serverless function that does a lot of syscalls (which is every function that does I/O), that's a lot of flushes.

The workaround: For microVMs, the guest kernel is separate from the host kernel. The guest's syscalls don't trigger a host KPTI switch until the guest exits (e.g., for I/O). So you want to batch I/O to minimize VM exits. This is why virtio and vhost are so important—they allow the guest to do I/O without a full context switch to the host.


A Real-World Example: The "Noisy Neighbor" from Hell

Let me tell you a story. We had a serverless platform running on bare metal. 128 cores. We were packing tenants densely. One day, we noticed a weird pattern: p99 latency for a specific tenant (let's call them "Tenant X") spiked every 30 seconds.

We pulled out perf and bcc (BPF Compiler Collection). We looked at the LLC (Last Level Cache) miss rate. It was off the charts for Tenant X. But the CPU utilization for Tenant X was low. So why was it slow?

The answer: A noisy neighbor. Tenant Y, on the same physical host, was running a background job that was streaming through a 10GB dataset. It was evicting all of Tenant X's data from the L3 cache. Tenant X wasn't CPU-starved; it was cache-starved.

The fix: We implemented a cache-aware scheduler. We used Intel's RDT (Resource Director Technology)—specifically CAT (Cache Allocation Technology). This allows you to partition the L3 cache. We gave Tenant X a dedicated 4MB slice of L3, and Tenant Y a separate 8MB slice.

The result: Tenant X's p99 dropped from 200ms to 45ms. Not because we gave it more CPU, but because we gave it predictable cache.

Code snippet — using pqos to configure RDT:

# Allocate 4MB of L3 cache to Tenant X (assuming 16-way associativity)
pqos -e "llc:0=0x000f;llc:1=0x00f0"
# Pin processes to cores and apply the cache allocation
pqos -a "llc:0=0-3;llc:1=4-7"

This is the kind of thing that separates a "toy" serverless platform from a production one.


The WASM Wildcard: Bypassing the Kernel Entirely

We've been talking about kernel-level optimizations, but there's a parallel universe: WebAssembly (WASM).

Cloudflare Workers, Fastly Compute@Edge, and others run WASM isolates inside a single process. No context switch at the kernel level. No TLB flush. No KPTI. The "context switch" is a function call in user space.

This is dramatically faster. A WASM context switch can be sub-microsecond. It's just saving registers and swapping a stack pointer.

But WASM has its own cache issues. The WASM runtime's code is shared, but the tenant's data is not. If you're not careful, you'll still get cache thrashing. Plus, WASM's linear memory model means you're often copying data between the host and the guest, which is its own performance tax.

The takeaway: If you can fit your workload into a WASM sandbox, you avoid the kernel entirely. But for workloads that need raw syscalls, threads, or heavy I/O, you're back to the kernel, and you're back to the context switch.


The Future: Hardware-Assisted Context Switching

Where is this all going? The hardware vendors are paying attention.

  • Intel's Sapphire Rapids and AMD's Genoa have better TLB architectures and more PCID bits.
  • ARM's SVE (Scalable Vector Extension) and MTE (Memory Tagging Extension) are changing the game for memory safety and performance.
  • CXL (Compute Express Link) promises shared memory pools that could allow tenants to share cache lines without copying.

But the most interesting development is CPU-side scheduling hints. Imagine a CPU that understands "this microVM is a short-lived function; please don't evict its cache lines." That's not science fiction; it's on the roadmap.


The Bottom Line (For Engineers, By Engineers)

If you're building a serverless platform, or if you're just trying to get your Lambda cold starts under control, here's your checklist:

  1. Pin your tenants. Don't let the scheduler bounce them across cores. Use CPU affinity.
  2. Size your tenants. Keep working sets small so they fit in L3. Use memory limits and ballooning.
  3. Use huge pages. Reduce TLB pressure. Hint the kernel with madvise(MADV_HUGEPAGE).
  4. Partition the cache. Use Intel RDT or AMD's equivalent to prevent noisy neighbors from evicting your hot data.
  5. Batch your I/O. Minimize VM exits and KPTI flushes. Use io_uring if you can.
  6. Measure everything. Use perf, bcc, and bpftrace. Look at LLC miss rates, TLB miss rates, and context switch rates. If you're not measuring, you're guessing.

The cold start problem isn't a runtime problem. It's a kernel problem. And the kernel doesn't care about your marketing slides. It cares about cache lines, page tables, and the CR3 register.

Optimize those, and you'll be the one writing the "we have 10ms p99 cold starts" blog post. Ignore them, and you'll be the one blaming Node.js.

Stay curious. And always check your LLC miss rate.


If you enjoyed this deep dive, you might also like our posts on "Why Your MicroVM Is Slower Than You Think" and "The Hidden Cost of fork() in Serverless." Drop a comment below or find me on the usual channels—I'm always up for a debate about the merits of SCHED_BATCH versus SCHED_IDLE.


More to explore

Keep diving in