Trending:

Cloudflare trims RAM usage on Pingora-based services by 25% and reclaims 100TB globally, aided by math and Rust

Illustration of memory chips with Cloudflare logo
TechStaged-owned

Summary

  • Cloudflare reduced the memory footprint of a Pingora-based service by 25% through a memory-layout optimization of the hashing structure.
  • This memory reduction contributed to reclaiming more than 100TB of RAM globally, in addition to 100TB shed by the DNS team last month.
  • Pingora Backend Router (PBR) uses pingora-ketama for consistent hashing to route cacheable requests.

Cloudflare describes a memory-usage improvement in a Pingora-based service that led to reclaiming more than 100TB of RAM globally. The effort builds on prior memory reductions achieved by the DNS team last month.

WHAT HAPPENED AND WHY IT MATTERED

The investigation started with excessive memory usage in the Pingora Backend Router (PBR) related to the pingora-ketama component, which implements consistent hashing used to route cacheable requests to servers by URL. TechStaged has also covered Policy window open as Lehane calls for stronger AI safety evidence, shared standards, and durable policy action.

Consistent hashing maps tasks and servers along a numeric range. The post explains how the 32-bit hash outputs can create imbalances and how adding multiple hashes per server can help even out workloads, while weights can reflect disk space or CPU/GPU capacity.

HOW THE MEMORY SAVINGS WERE ACHIEVED

A Rust-related memory layout issue limited how small the in-memory representation of the hashing structure could be. A safer optimization involved storing the hash and index as a raw byte array and accessing them with getters, which reduced memory usage for consistent hashing by about 25%.

The post notes that shrinking the number of hashes per server was key to reducing memory, relying on math to show that fewer hashes could be used without appreciable error under their workloads.

SCALE, RISK, AND THE PATH FORWARD

Cloudflare emphasizes the scale of its operations—thousands of servers with petabytes of RAM and millions of CPU cores—and that small percentage improvements can have large aggregate effects.

To minimize cache churn, the team avoided a single, global switch to a new hash ring. Instead, they maintained both versions of the cacheable load during the transition.

WHAT HAPPENS NEXT

The post frames this as part of ongoing performance optimization, with memory efficiency continuing to be a priority as the Pingora-based services evolve and as resource sharing across teams is managed.

Reporting by Owen Blackridge; editing by TechStaged editors

Editorial disclosure: This article was prepared with AI assistance from a source-limited research package and passed TechStaged's automated factual, originality, licensing, and publication checks.

Our Standards: The TechStaged Editorial Principles.

Suggested Topics: Software Business Software
f in

Owen Blackridge

Owen Blackridge

Technology Editor

Owen covers platform shifts, AI launches, and the practical impact of emerging technology on small teams.