Showing posts with label BBU. Show all posts
Showing posts with label BBU. Show all posts

Sunday, September 13, 2026

Working with GPUs - Part10 - Energy Storage Systems, BBUs, and CBUs explained

As AI clusters transition from tens of kilowatts to over 200 kw per rack, traditional server power architectures are hitting their physical limits. High-density GPU accelerators like Nvidia Hopper, Blackwell, and liquid-cooled OCP ORV3 HPR racks introduce extreme dynamic power loads that challenge legacy data center distribution.

To prevent continuous voltage droops, nuisance trips, and power supply shutdowns under bursting workload conditions, the data center industry is shifting its approach to power delivery. Centralized power management is giving way to a multi-tiered Energy Storage System (ESS) network distributed across every level of the facility - from containerized yards down to sidecar racks, rack-level Battery Backup Units (BBUs) and Capacitor Backup Units (CBUs).

In this post, we’ll explore the paradigm shift in data center energy storage, examine why AI changes the stakes, break down multi-tier ESS deployment levels, and highlight the architectural planning required for safe, efficient operation.

Why Energy Storage Matters Beyond Central UPS

Historically, data center backup power was straightforward: a massive, centralized room-level Uninterruptible Power Supply (UPS) paired with diesel generators handled every outage. Today, that model is breaking down.

  • Distributed Deployment: Energy storage is no longer confined to a single, centralized grey-space room. Modern high-density architectures distribute storage across multiple tiers: facility-level yards, row-level rows/sidecars, and rack-level modules.
  • Proximity to the IT Load: Moving energy storage closer to the compute load fundamentally changes the engineering conversation. Instead of pushing massive transient currents over long busbars from a distant UPS, local buffering absorbs shocks/ spikes right where they happen.
  • Flexible Footprints: Today, ESS components can live wherever physical efficiency dictates - whether in dedicated equipment rooms, active white space, grey service corridors, or outdoor container yards.

Why AI Changes the Stakes for Energy Storage

Traditional cloud workloads follow relatively stable, predictable power utilization curves. Artificial intelligence and machine learning training jobs, however, behave entirely differently.

  • Extreme Power Dynamics: Large AI clusters experience rapid load changes, massive peak currents, and intense local power densities. When thousands of tensor cores synchronize to process a massive matrix multiplication, the instantaneous jump in current draw can cause catastrophic voltage droops.
  • Multi-Tiered Time Bands: Different energy storage technologies operate across vastly different time scales:
    • Nanoseconds to Milliseconds: Supercapacitors absorb high-frequency ripples and microsecond spikes.
    • Seconds to Minutes: Rack-level lithium-ion BBUs provide ride-through power for short grid glitches or graceful checkpoint shutdowns.
    • Hours: Facility-level battery energy storage systems and generators cover extended utility outages.
  • The "One-Size-Fits-All" Fallacy: No single energy storage technology can cover every time band efficiently. High-power density demands a hybrid approach where specialized technologies handle specific layers of the power delivery network.

Image credits: academy.opencompute.org

Levels of ESS Deployment

In modern high-performance AI data centers, energy storage is deployed across distinct, coordinated tiers:


Image credits: academy.opencompute.org



Rack-Level BBUs: Standardized via the OCP ORv3 HPR architecture, rack-level battery module enclosures slide directly into the vertical 48V DC busbar. They provide 2 to 4 minutes of local ride-through power.

Supercapacitor Buffers (CBUs): Placed right at the rack as physically separate modular enclosures, supercapacitors neutralize sub-millisecond voltage droops and transient spikes caused by rapid GPU/TPU execution shifts, shielding battery chemistry from repetitive thermal stress. CBU play a critical role in power smoothing within modern high-density AI infrastructure as the multi-node GPU clusters shift from low-power states to intense tensor execution almost instantaneously, they generate violent current spikes and high-frequency voltage ripples. CBUs smooth out these power instabilities and the power supply, and BBU units are shielded from the "bursty" nature of AI workloads. The CBU acts as a localized shock absorber, turning erratic, volatile power demands into a smooth, continuous load profile for the upstream rectifiers and lithium-ion batteries.

Sidecar Enclosures: For extreme 100 kw+ racks, rectifiers, BBUs, and CBUs are frequently grouped into dedicated vertical power sidecars attached to the side of the compute rack.

Centralized UPS & Containerized BESS: Massive lithium-ion or advanced chemistry battery arrays deployed in dedicated power rooms or outdoor modular yards. These act as the primary bridge between sudden grid failure and emergency generator startup.

Onsite Generation Integration: Modern resilient data centers increasingly integrate localized microgrid components - such as solar arrays, fuel cells, or advanced generators - into the facility ESS loop to stabilize incoming power and reduce reliance on fragile public grids.

The Recharge Challenge: Managing Grid Limits and PUE


Deploying distributed energy storage introduces a critical engineering consideration: every ESS must eventually recharge.
  • The Coordination Problem: If a large fleet of rack-level BBUs or facility BESS units all initiate high-current recharging simultaneously after a minor grid glitch or peak-shaving event, they can inadvertently breach facility max-demand thresholds.
  • Impact on PUE: Uncoordinated recharging spikes total facility load, driving up power consumption and negatively impacting Power Usage Effectiveness (PUE) metrics.
  • Architectural Integration: Recharging cycles must be algorithmically scheduled and integrated directly into the site's power management firmware, ensuring that energy replenishment happens smoothly during off-peak windows or throttled alongside current IT workload demands.

As AI infrastructure scales into the next generation of high-density computing, treating power delivery and energy storage as an intelligent, distributed ecosystem is no longer optional - it is the baseline for uncompromised uptime.

References

Saturday, March 7, 2026

Working with GPUs – A Practical Blog Series

This blog series captures practical learnings from working with GPUs in real‑world environments, with a focus on operations, reliability, and scale. Each post deep‑dives into specific aspects of GPU systems based on hands‑on experience, incidents, and operational challenges. Together, these articles aim to share actionable insights, highlight common pitfalls, and help teams build more robust and predictable GPU operations.


Part 01: Using nvidia-smi
Part 02: Memory fault indicators
Part 03: Using dcgmi
Part 04: Thermal issues
Part 05: XID errors
Part 06: H100 SXM5 architecture
Part 07: GPU has fallen off the bus
Part 08: Powering and cooling AI accelerators with OCP ORV3 HPR
Part 09: Power shelves in OCP ORV3
Part 10: Energy storage systems, BBUs, and CBUs