Sunday, September 13, 2026

Working with GPUs - Part10 - Energy Storage Systems, BBUs, and CBUs explained

As AI clusters transition from tens of kilowatts to over 200 kw per rack, traditional server power architectures are hitting their physical limits. High-density GPU accelerators like Nvidia Hopper, Blackwell, and liquid-cooled OCP ORV3 HPR racks introduce extreme dynamic power loads that challenge legacy data center distribution.

To prevent continuous voltage droops, nuisance trips, and power supply shutdowns under bursting workload conditions, the data center industry is shifting its approach to power delivery. Centralized power management is giving way to a multi-tiered Energy Storage System (ESS) network distributed across every level of the facility - from containerized yards down to sidecar racks, rack-level Battery Backup Units (BBUs) and Capacitor Backup Units (CBUs).

In this post, we’ll explore the paradigm shift in data center energy storage, examine why AI changes the stakes, break down multi-tier ESS deployment levels, and highlight the architectural planning required for safe, efficient operation.

Why Energy Storage Matters Beyond Central UPS

Historically, data center backup power was straightforward: a massive, centralized room-level Uninterruptible Power Supply (UPS) paired with diesel generators handled every outage. Today, that model is breaking down.

  • Distributed Deployment: Energy storage is no longer confined to a single, centralized grey-space room. Modern high-density architectures distribute storage across multiple tiers: facility-level yards, row-level rows/sidecars, and rack-level modules.
  • Proximity to the IT Load: Moving energy storage closer to the compute load fundamentally changes the engineering conversation. Instead of pushing massive transient currents over long busbars from a distant UPS, local buffering absorbs shocks/ spikes right where they happen.
  • Flexible Footprints: Today, ESS components can live wherever physical efficiency dictates - whether in dedicated equipment rooms, active white space, grey service corridors, or outdoor container yards.

Why AI Changes the Stakes for Energy Storage

Traditional cloud workloads follow relatively stable, predictable power utilization curves. Artificial intelligence and machine learning training jobs, however, behave entirely differently.

  • Extreme Power Dynamics: Large AI clusters experience rapid load changes, massive peak currents, and intense local power densities. When thousands of tensor cores synchronize to process a massive matrix multiplication, the instantaneous jump in current draw can cause catastrophic voltage droops.
  • Multi-Tiered Time Bands: Different energy storage technologies operate across vastly different time scales:
    • Nanoseconds to Milliseconds: Supercapacitors absorb high-frequency ripples and microsecond spikes.
    • Seconds to Minutes: Rack-level lithium-ion BBUs provide ride-through power for short grid glitches or graceful checkpoint shutdowns.
    • Hours: Facility-level battery energy storage systems and generators cover extended utility outages.
  • The "One-Size-Fits-All" Fallacy: No single energy storage technology can cover every time band efficiently. High-power density demands a hybrid approach where specialized technologies handle specific layers of the power delivery network.

Image credits: academy.opencompute.org

Levels of ESS Deployment

In modern high-performance AI data centers, energy storage is deployed across distinct, coordinated tiers:


Image credits: academy.opencompute.org



Rack-Level BBUs: Standardized via the OCP ORv3 HPR architecture, rack-level battery module enclosures slide directly into the vertical 48V DC busbar. They provide 2 to 4 minutes of local ride-through power.

Supercapacitor Buffers (CBUs): Placed right at the rack as physically separate modular enclosures, supercapacitors neutralize sub-millisecond voltage droops and transient spikes caused by rapid GPU/TPU execution shifts, shielding battery chemistry from repetitive thermal stress. CBU play a critical role in power smoothing within modern high-density AI infrastructure as the multi-node GPU clusters shift from low-power states to intense tensor execution almost instantaneously, they generate violent current spikes and high-frequency voltage ripples. CBUs smooth out these power instabilities and the power supply, and BBU units are shielded from the "bursty" nature of AI workloads. The CBU acts as a localized shock absorber, turning erratic, volatile power demands into a smooth, continuous load profile for the upstream rectifiers and lithium-ion batteries.

Sidecar Enclosures: For extreme 100 kw+ racks, rectifiers, BBUs, and CBUs are frequently grouped into dedicated vertical power sidecars attached to the side of the compute rack.

Centralized UPS & Containerized BESS: Massive lithium-ion or advanced chemistry battery arrays deployed in dedicated power rooms or outdoor modular yards. These act as the primary bridge between sudden grid failure and emergency generator startup.

Onsite Generation Integration: Modern resilient data centers increasingly integrate localized microgrid components - such as solar arrays, fuel cells, or advanced generators - into the facility ESS loop to stabilize incoming power and reduce reliance on fragile public grids.

The Recharge Challenge: Managing Grid Limits and PUE


Deploying distributed energy storage introduces a critical engineering consideration: every ESS must eventually recharge.
  • The Coordination Problem: If a large fleet of rack-level BBUs or facility BESS units all initiate high-current recharging simultaneously after a minor grid glitch or peak-shaving event, they can inadvertently breach facility max-demand thresholds.
  • Impact on PUE: Uncoordinated recharging spikes total facility load, driving up power consumption and negatively impacting Power Usage Effectiveness (PUE) metrics.
  • Architectural Integration: Recharging cycles must be algorithmically scheduled and integrated directly into the site's power management firmware, ensuring that energy replenishment happens smoothly during off-peak windows or throttled alongside current IT workload demands.

As AI infrastructure scales into the next generation of high-density computing, treating power delivery and energy storage as an intelligent, distributed ecosystem is no longer optional - it is the baseline for uncompromised uptime.

References

Sunday, August 9, 2026

Working with GPUs - Part9 - Power shelves in OCP ORV3 High Power Rack

The ORV3 HPR Power Shelf is a standardized, ultra-high-density power distribution system designed under the Open Compute Project (OCP) framework. Essentially, it acts as a centralized power hub for a server rack. Instead of every individual server having its own power supply unit (PSU) throwing off heat and hogging space, the power shelf sits in the rack, takes high-voltage AC utility power and converts it into a single, massive pool of 48V DC power. This power is then delivered to the entire rack via a heavy-duty vertical copper backplane (busbar).


AI and LLM workloads are notoriously bursty; a GPU cluster can spike from an idle state to maximum power draw in microseconds. The ORV3 HPR power shelf features high pulse-load capabilities (often supporting up to 150% load capacity for transient windows) and active current sharing. This smooths out dynamic loading and prevents voltage sags without tripping upstream data center breakers. Equipped with an integrated Power Management Controller (PMC), these shelves expose real-time metrics, black-box fault logging, and granular thermal monitoring via standard APIs like DMTF Redfish over Gigabit Ethernet. This allows infrastructure teams to optimize power provisioning, balance loads accurately, and proactively manage hot spots in the data center.

Power shelf is a group c, third party component present in Nvidia GB200/300 NVL72 rack. 

  • Currently Nvidia supports Delta and LiteOn power shelves. 
  • They are 33KW EIA 1 RU (6 x 5.5KW PSUs) units with additional bulk capacitors and 60A whip support. 
  • Nvidia GB300 NVL72 MGX rack has single bus bar in the middle, and total 6 power shelves located at top and bottom of the rack. 
  • The rack power consumption is approximately 120kW, and the 6 power shelves will provide N+2 redundancy.

Power Shelf firmware

  • PMC firmware is based on Open BMC.
  • The firmware directly governs how power is managed, balanced, and protected across the entire rack.
  • It controls the internal switching frequencies, power factor correction (PFC), and voltage regulation loops of the individual PSUs.
  • It ensures that if you have six PSUs in a shelf, they all pull their weight equally. If one PSU lags, the firmware recalibrates the others in microseconds to prevent overloading a single unit.
  • It dictates how the shelf handles massive, sudden spikes in power when GPUs transition from idle to 100% utilization.
  • It hosts the communication protocols (like Modbus, PMBus, or Redfish over Ethernet) used by the Power Management Controller (PMC) to talk to your rack-level orchestrators.

Accessing Power shelf

  • Power Shelf has a PMC (Power Management Controller) - connected via ethernet port.
  • This PMC usually gets connected to OOB network.
  • You can access it via web UI, SSH, or Redfish.
  • Also supports SNMP.

Updating Power shelf platform firmware

  • We need to update two things:
    • PMC firmware
    • PSU firmware
  • This can be done using nvfwupd utility.
  • Notes:
    • LiteOn PSUs may only be updated one at a time. To select the PSU to update, a special JSON file containing the “LiteOnPowerDeviceId” value is required.
    • Delta PSUs update simultaneously and no special JSON file is required.
    • After the update completes, PowerShelf components automatically activate the new firmware.
    • Starting with NVFWUPD 2.1.0, OnReset activation is supported for Delta and LiteOn PowerShelf platforms. By default, all updates use immediate activation.
    • OnReset activation applies only to PMC firmware updates on Delta and LiteOn PowerShelf platforms. PSU firmware updates always use immediate activation.

Operational reality and firmware update

  • You don't need to patch power shelf firmware monthly. Usually, a biannual or annual cadence or aligning updates with major hardware maintenance windows is standard.
  • When expanding clusters or adding new generations of compute nodes to existing racks, updating the power shelf firmware ensures the power delivery system is fully compatible with the power sequencing behaviors of the newer servers.
  • Modern ORV3 shelves support hitless/lossless firmware updates. This means you can flash the firmware on the Power Management Controller (PMC) or individual PSUs sequentially while the rack remains fully powered and live, eliminating the need to take your compute offline just to update the power system.

References

Sunday, July 26, 2026

Working with GPUs - Part8 - Powering and cooling AI accelerators with OCP ORV3 HPR

When we talk about working with GPUs at scale, we usually focus on the software stack, CUDA tuning, or optimizing workloads across NVLink, etc. But if you are an infrastructure SRE managing bare-metal clusters, you quickly run into a much harsher physical reality: Power and Thermals.

As we deploy next-generation AI platforms like Nvidia’s GB200/ 300 NVL72 architectures - the power demands are obliterating traditional infrastructure. Standard server racks are physically hitting a wall. To keep these high-density clusters running without melting down, the industry is rapidly transitioning to the Open Compute Project’s (OCP) Open Rack Version 3 (ORV3) High Power Rack (HPR).

Here is what you need to know about how this architecture feeds and cools modern GPU nodes.

The GPU infrastructure bottleneck

Traditional data center racks rely on a 19-inch width standard with standard 44.45 mm Rack Units (RU). In that legacy model, every individual server chassis houses its own AC-to-DC power supply units (PSUs).

If you try to stuff a modern cluster of high-TDP GPUs into that traditional 19-inch racks, you run into immediate problems:

  • Cable chaos: The back of the rack becomes choked with heavy AC power cords, blocking vital airflow.
  • Efficiency loss: Converting AC to DC at every single node generates massive heat and power waste.
  • Weight limits: GPU nodes are incredibly dense and heavy; standard frames simply aren't structurally rated for them.

How ORV3 changes the game for GPU compute


Developed collaboratively by hyper-scalers like Meta, Google, and Microsoft, the ORV3 standard throws out the legacy playbook to accommodate modern accelerator demands.

Instead of treating the rack like a cabinet for isolated servers, ORV3 treats the entire rack as a single, unified compute machine:
  • Native 21-inch bays and Open Units (OU): Provides a wider internal bay (21 inches wide vs. the traditional 19 inches). It replaces RUs with Open Units (OU = 48mm), offering more structural space for complex GPU heat sinks and optimized front-to-rear airflow.
21-inch-wide rack

44 OU

  • Centralized 48V DC busbar: Individual server power supplies are completely gone. Instead, 3-phase AC or High Voltage DC (HVDC) enters a centralized power shelf, which converts it to 48V DC. This power is run down a copper busbar mounted at the rear center of the rack.
48V busbar

Busbar BarKlip connector

  • Power shelves: The ORV3 HPR power shelf acts as a centralized power hub that converts incoming 3-phase AC (from the overhead busways) into 48V DC power distributed via the rear busbar. Housing high-density 5.5 kW Power Supply Units (PSUs), a single 1U power shelf delivers up to 33kW of total output. Racks deploy multiple shelves in N+N or N+1 redundant configurations - complete with integrated Power Monitor Modules (PMM) to supply reliable, cable-free energy to high-TDP GPU clusters.

AC input overhead busways

AC input - Power shelf - DC output - Busbar

Power shelf with multiple PSUs inside it
GB200 NVL72 ORV3 HPR rack

  • Blind-mate infrastructure: When you slide a heavy GPU compute node into the rack, it connects directly to the 48V DC busbar via copper clips (blind-mate connections). No power cables required. The specialized copper clip/ jaw that physically clamps onto the vertical busbar blade to transmit high-current DC power is called a Busbar BarKlip connector. Blind-mate connectors are widely used for liquid cooling as well in modern high-density data centers. In the OCP ORV3 HPR architecture, the concept of "blind-mating" applies to both power and coolant distribution. Instead of manually hooking up coolant hoses to the back of a server, liquid-cooled blind-mate connectors (often referred to as BMQC or Blind Mate Quick Connectors) allow fluid lines to engage automatically as the compute tray slides into the rack.
Liquid cooling blind-mate connectors and manifold

  • Heavy duty chassis support: The frame is built to support up to 1400 kg, meaning it won't buckle under a full stack of liquid-cooled accelerators.


The ORV3 HPR specification


While standard ORV3 configurations top out at 18 kW to 36 kW per rack, modern AI workloads easily blow past those thresholds. To keep up with platforms like the Nvidia GB200 pushing rack limits to 140kW and beyond, the OCP community introduced the ORV3 HPR (High Power Rack) variant which is an extension of ORV3.
  • Power density: 92 kW to 140 kW+ / rack
  • Power shelf capacity: 5.5 kW PSUs (33 kW total per power shelf)
  • Cooling architecture: Blind-Mate Direct Liquid Cooling (DLC) Manifolds
To handle massive electrical currents without thermal runaway, the HPR upgrades to a massive 80 kg busbar with deeper tracking and aggressive grounding. More importantly, it addresses the massive heat generated by high-TDP GPUs by integrating blind-mate liquid cooling manifolds right into the chassis. Just like the power clips, the liquid cooling loops engage automatically when the node is seated.

What’s next: The 1-Megawatt sidecar


As we look forward, GPU power requirements show no signs of slowing down. As clusters head toward 1 Megawatt (MW) per rack, the OCP community is already developing Project Mount Diablo. This next step introduces a disaggregated Power Rack Sidecar, moving the massive rectifiers completely outside of the main compute rack so we can fill every square inch of the primary frame with pure, liquid-cooled GPU compute.



Working with GPUs at scale means understanding the infrastructure that keeps them alive. Without open standards like ORV3 HPR solving the physical limitations of power delivery and fluid dynamics, the next leap in AI compute wouldn't even be able to turn on.

References


Hope it was useful. Cheers!

Sunday, June 14, 2026

Working with GPUs - Part7 - Unable to determine the device handle for GPU

One of the more common GPU failures encountered in large-scale AI clusters is a situation where the Nvidia driver can no longer communicate with one or more GPUs. In these cases, nvidia-smi may partially work, showing healthy GPUs while reporting errors for the affected devices. This post walks through a real-world example involving Nvidia H100 GPUs on an 8 GPU Supermicro server where two GPUs became inaccessible and generated XID 79 - GPU has fallen off the bus errors.


Symptoms

The first indication of the problem was that nvidia-smi could not communicate with all GPUs in the system.

Running the following command: nvidia-smi -L


Notice that GPUs 0-5 are detected correctly, while two GPUs fail during enumeration because the driver can no longer obtain a valid device handle. This means the Nvidia Management Library (NVML) loaded and detected the GPU on the PCIe bus but failed when attempting to initialize communication with the GPU.


A common observation in this state is:

  • Queries against healthy GPUs continue to work.
  • Queries against the affected GPUs fail immediately.
  • Workloads that require all GPUs may fail or become stuck.

Troubleshooting


Step 1: Check for XID Errors


The next step is to look for Nvidia XID events.
  • dcgmi dmon -e 230 --count 1
  • dmesg | egrep -i "xid|fallen"


In this case, you can see logs indicating GPU has fallen off the bus.
# dmesg | egrep -i "xid|fallen"
[6780696.284301] NVRM: Xid (PCI:0000:bf:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
[6780696.284304] NVRM: GPU 0000:bf:00.0: GPU has fallen off the bus.
[6780696.284647] NVRM: Xid (PCI:0000:e4:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
[6780696.284648] NVRM: GPU 0000:e4:00.0: GPU has fallen off the bus.

XID 79 indicates that the Nvidia driver lost communication with the GPU over PCIe. From the operating system's perspective, the device is no longer responding correctly to configuration or memory transactions.

Potential causes include:
  • PCIe link failures
  • GPU hardware issues
  • Power-related problems
  • Motherboard or PCIe switch issues
  • Firmware or driver defects
  • Unexpected device resets

Once this occurs, workloads using the affected GPU are typically unable to continue.

Step 2: Verify PCIe Device Health


Next, inspect the PCIe devices directly.
  • lspci | grep -i nvidia
  • lspci -vvv -s <pcie device id>
# lspci | grep -i nvidia
05:00.0 Bridge: NVIDIA Corporation Device 22a3 (rev a1)
06:00.0 Bridge: NVIDIA Corporation Device 22a3 (rev a1)
07:00.0 Bridge: NVIDIA Corporation Device 22a3 (rev a1)
08:00.0 Bridge: NVIDIA Corporation Device 22a3 (rev a1)
19:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
2d:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
3f:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
66:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
9b:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
ae:00.0 3D controller: NVIDIA Corporation Device 2330 (rev a1)
bf:00.0 3D controller: NVIDIA Corporation Device 2330 (rev ff)
e4:00.0 3D controller: NVIDIA Corporation Device 2330 (rev ff)

# lspci -vvv -s bf:00.0
bf:00.0 3D controller: NVIDIA Corporation Device 2330 (rev ff) (prog-if ff)
        !!! Unknown header type 7f
        Kernel driver in use: nvidia
        Kernel modules: nvidiafb, nouveau, nvidia_drm, nvidia

# lspci -vvv -s e4:00.0
e4:00.0 3D controller: NVIDIA Corporation Device 2330 (rev ff) (prog-if ff)
        !!! Unknown header type 7f
        Kernel driver in use: nvidia
        Kernel modules: nvidiafb, nouveau, nvidia_drm, nvidia
A healthy GPU normally returns detailed PCIe configuration information. When lspci reports: !!! Unknown header type 7f - it usually means the operating system can still "see" something at that PCIe address, but it cannot successfully read the device's PCI configuration space.

This is a strong indication that the device is no longer responding correctly and aligns with the earlier XID 79 messages.

Step 3: Run DCGM PCIe Diagnostics


Nvidia DCGM provides additional validation capabilities.

Run: dcgmi diag -r 3


This confirms that CUDA is unable to initialize the affected GPU, further validating that the device is no longer accessible from the driver stack. At this point, multiple layers of the stack indicate the same issue, and these findings strongly suggest that the GPUs are no longer responding properly over the PCIe fabric.

Step 4: Check syslog


Take a look at the syslog, and you should be able to see NVRM or XID events in it. Following is a sample log snippet:
kernel	[19209.003187] pcieport 0000:18:00.0: pciehp: Slot(1-1): Card not present
kernel	[19209.003184] pcieport 0000:18:00.0: pciehp: Slot(1-1): Link Down
kernel	[19209.003187] pcieport 0000:18:00.0: pciehp: Slot(1-1): Card not present
kernel	[19209.003184] pcieport 0000:18:00.0: pciehp: Slot(1-1): Link Down
kernel	[19209.003246] NVRM: Attempting to remove device 0000:19:00.0 with non-zero usage count!
kernel	[19209.003246] NVRM: Attempting to remove device 0000:19:00.0 with non-zero usage count!
kernel	[19209.174506] NVRM: GPU at PCI:0000:19:00: GPU-16xx3295-1242-bdc9-b0bf-0d540dxxxxx
kernel	[19209.174511] NVRM: GPU Board Serial Number: xx551xx014xx7.
kernel	[19209.174516] NVRM: GPU 0000:19:00.0: GPU has fallen off the bus.
kernel	[19209.174524] NVRM: A GPU crash dump has been created. If possible, please run
kernel	[19209.174524] NVRM: nvidia-bug-report.sh as root to collect this data before
kernel	[19209.174506] NVRM: GPU at PCI:0000:19:00: GPU-16xx3295-1242-bdc9-b0bf-0d540dxxxx
kernel	[19209.174513] NVRM: Xid (PCI:0000:19:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
kernel	[19209.174516] NVRM: GPU 0000:19:00.0: GPU has fallen off the bus.
kernel	[19209.174524] NVRM: A GPU crash dump has been created. If possible, please run
kernel	[19209.174524] NVRM: nvidia-bug-report.sh as root to collect this data before
kernel	[19209.174524] NVRM: the NVIDIA kernel module is unloaded.
kernel	[19209.174513] NVRM: Xid (PCI:0000:19:00): 79, pid='<unknown>', name=<unknown>, GPU has fallen off the bus.
kernel	[19209.174517] NVRM: GPU 0000:19:00.0: GPU serial number is xx551xx014xx7.
kernel [19209.174524] NVRM: the NVIDIA kernel module is unloaded. kernel [19209.174511] NVRM: GPU Board Serial Number: xx551xx014xx7 kernel [19209.174517] NVRM: GPU 0000:19:00.0: GPU serial number is xx551xx014xx7.

Step 5: Verify BMC


You can also take a look the server health event logs in the BMC for events like GPU or HBM missing.

Resolution


Following are the most common steps you can try:
  • Power cycle the server. In most cases this will resolve the issue.
  • If a power cycle is not resolving the issue, you may try reprovisioning the node by reinstalling the OS and other required software components.
  • Last option is to request for RMA of the server/ affected part by working with the vendor.

Hope it was useful. Cheers!

Sunday, May 17, 2026

Working with GPUs - Part6 - H100 SXM5 architecture

Now that we have a foundational understanding of the core utilities used to monitor and manage GPUs, let's dive into the hardware architecture of the NVIDIA H100 SXM. To truly understand GPU computing, it is essential to visualize how data flows through the silicon. The following overview maps out the internal components of the H100, providing a clear frame of reference so you can easily correlate key architectural terms such as Streaming Multiprocessors (SMs), Tensor Cores, L2 Cache, High Bandwidth Memory (HBM3), etc.

Overview

  • H100 is released around Sep 2022
  • Based on the Hopper architecture 
  • It comes in two form factors
    • PCIe (300W)
    • SXM (700W)
  • Has 80 GB HBM3 memory (3.35 TB/s)
  • 132 SMs
  • 528 Tensor cores (4 per SM)
  • 80B transistors on a custom 4N process node

Architecture


Image ref: NVIDIA Hopper Architecture In-Depth | NVIDIA Technical Blog

HBM3 - High Bandwidth Memory

  • This is the off-chip 80 GB device memory.
  • Divided in 5 stacks and connected via 10 independent 512-bit memory controllers.
  • Data flow: SM - L1 - L2 - memory partition/ crossbar - memory controller - HBM stack
  • H100 SXM5 has 5 HBM3 stacks.
  • HBM3 stack is DRAM.

L2 cache

  • 50 MB of L2 cache, divided into two 25 MB partitions.
  • L2 cache is SRAM.

Unified shared memory + L1 cache

  • 256 KB size 33 TB/s bandwidth per SM divided into 32 banks, each 32 bits (4 bytes) wide.
  • These are SRAM.

Registers

  • Every thread gets a private set of on-chip registers. 
  • They have very high bandwidth, and very low latency.
  • 256 KB per SM.

Gigathread engine

  • This is the hardware that takes a kernel launch and hands out thread blocks to SMs.
  • It tracks which thread blocks are not yet started, running, and finished.
  • When an SM has capacity to run another thread block, the Gigathread engine assigns the next thread block to that SM.
  • This ensures intelligent work distribution for optimal GPU utilization.

SM - Streaming Multiprocessor

  • SMs are the fundamental execution unit of the GPU which executes thread blocks of a CUDA kernel. 
  • H100 SXM5 GPU has 132 SMs.
  • Following are the components of SM:
    • FP32 CUDA cores, Int/FP64 units
    • 4th gen Tensor cores
    • Shared memory/ L1 cache
    • L1 instruction cache
    • Warp scheduler
    • Dispatch units
    • Registers
    • L0 instruction cache
  • Each SM is divided into 4 identical sub-divisions called Quadrants or SMSPs (SM Sub Partitions).

TMA - Tensor Memory Accelerator

  • Each SM has a TMA unit.
  • Offloads the tensor copy operations from the SMs.

Tensor core

  • They are really fast units for performing MMA operations (Matrix Multiply Accumulate).
  • 4 tensor cores per SM.

GPC - Graphics Processing Cluster

  • It is a group of 18 SMs.
  • There are 8 GPCs in a H100.
  • Each GPC is connected to its own chunk of L2 cache.
  • GPCs also enable the use of distributed shared memory between the SMs.

TPC - Texture Processing Unit

  • Single TPC holds 2 SMs.
  • Job of TPC is to have shared SM block, so that communication between the two SMs is really fast.

NVLink 

  • 4th gen NVLink.
  • 18 NVLink 4.0 lanes which gives 900 GB/s of GPU-GPU bandwidth.

References

Hope it was useful. Cheers!

Sunday, April 12, 2026

Working with GPUs - Part5 - XID errors

If you are running large-scale AI training or LLM inference, you already know that managing a GPU cluster is less about "if" things break, and more about "when" and "why". In this post, we’ll demystify Nvidia XID errors, interpret them in the context of H100 NVLink systems, and outline a practical approach to triage and remediation.


XID (short for eXception ID) errors are diagnostic messages emitted by the Nvidia kernel driver (NVRM) when a GPU encounters an abnormal condition or fault. While some point to minor software glitches, others signal catastrophic hardware failures. With the H100 equipped with High Bandwidth Memory (HBM3) and NVLink interconnects - understanding these errors are critical to minimizing downtime.

How do we identify if any GPUs has XID errors


DCGMI diag


One way to identify XID errors is from DCGMI level 3 tests. If a critical or fatal hardware XID fires while the Level 3 tests are actively running (or if a sticky hardware error state was already present), the test will fail and output specific error strings. Here is an example:
# dcgmi diag -r 3
Successfully ran diagnostic for group.
+---------------------------+------------------------------------------------+
| Diagnostic                | Result                                         |
+===========================+================================================+
|-----  Metadata  ----------+------------------------------------------------|
| DCGM Version              | 3.1.8                                          |
| Driver Version Detected   | 550.90.07                                      |
| GPU Device IDs Detected   | 2330,2330,2330,2330,2330,2330,2330,2330        |
|-----  Deployment  --------+------------------------------------------------|
| Denylist                  | Pass                                           |
| NVML Library              | Pass                                           |
| CUDA Main Library         | Pass                                           |
| Permissions and OS Blocks | Pass                                           |
| Persistence Mode          | Pass                                           |
| Environment Variables     | Pass                                           |
| Page Retirement/Row Remap | Pass                                           |
| Graphics Processes        | Pass                                           |
| Inforom                   | Pass                                           |
+-----  Integration  -------+------------------------------------------------+
| PCIe                      | Pass - All                                     |
+-----  Hardware  ----------+------------------------------------------------+
| GPU Memory                | Pass - All                                     |
| Diagnostic                | Pass - GPUs: 1, 2, 3, 4, 5, 6, 7               |
|                           | Fail - GPU: 0                                  |
| Warning                   | GPU 0 Found 56954234 faulty memory elements o  |
|                           | n GPU 0 Run a field diagnostic on the GPU.     |
| Info                      | GPU 0 Allocated space for 137 output matricie  |
|                           | s from 75937126809 bytes available., GPU 0 Ru  |
|                           | nning with precisions: FP64 1, FP32 1, FP16 1  |
|                           | , GPU 0 GPU 0 calculated at approximately 230  |
|                           | 2.72 gigaflops during this test                |
+-----  Stress  ------------+------------------------------------------------+
| Targeted Stress           | Pass - All                                     |
| Targeted Power            | Pass - GPUs: 1, 2, 3, 4, 5, 6, 7               |
|                           | Fail - GPU: 0                                  |
| Warning                   | GPU 0 Detected 43 xid_errors for GPU 0         | < xid_error
| Info                      | GPU 0 GPU 0 power average: 161 W               |
| Info                      | GPU 1 GPU 1 power average: 170 W               |
| Info                      | GPU 2 GPU 2 power average: 164 W               |
| Info                      | GPU 3 GPU 3 power average: 158 W               |
| Info                      | GPU 4 GPU 4 power average: 154 W               |
| Info                      | GPU 5 GPU 5 power average: 169 W               |
| Info                      | GPU 6 GPU 6 power average: 158 W               |
| Info                      | GPU 7 GPU 7 power average: 159 W               |
| Memory Bandwidth          | Pass - All                                     |
| EUD Test                  | Skip - All                                     |
+---------------------------+------------------------------------------------+
Note that while DCGMI is exceptionally good at flagging structural hardware faults, dcgmi diag will generally not flag application-level or user-space errors, even if they generate XIDs.

DCGMI dmon


Here is another way to look for XID errors using dcgmi dmon.
# dcgmi dmon -e 230 --count 1
#Entity   XIDER
ID
GPU 7     0
GPU 6     0
GPU 5     0
GPU 4     0
GPU 3     0
GPU 2     0
GPU 1     0
GPU 0     43
  • -e 230 is the filed id that shows the XID errors. The value shown under XIDER column is the specific XID error.
  • Note: If there are non‑zero values, that would mean one or more GPUs had logged Xid errors, and we need to cross‑reference the specific Xid codes in the kernel log and documentation to understand the nature of the fault.
  • Ref: Field Identifiers — NVIDIA DCGM Documentation latest documentation

OS logs


You will also find these XID errors in OS kernel logs, and syslog. Following are some examples:
# dmesg | grep -i xid
[  747.179157] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 1): Out Of Range Address
[  747.180533] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a7b0=0x100000e 0x56a7b4=0x20 0x56a7a8=0x1f81fb60 0x56a7ac=0x1174
[  747.209815] NVRM: Xid (PCI:0000:19:00): 43, pid=8548, name=nvvs, Ch 00000009
[ 1250.627528] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 3, SM 0): Out Of Range Address
[ 1250.629013] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x568730=0x100000e 0x568734=0x20 0x568728=0x1f81fb60 0x56872c=0x1174
[ 1250.657381] NVRM: Xid (PCI:0000:19:00): 43, pid=10603, name=nvvs, Ch 00000009
[45911.449627] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 0): Out Of Range Address
[45911.451042] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a730=0x103000e 0x56a734=0x20 0x56a728=0x1f81fb60 0x56a72c=0x1174
[45911.479823] NVRM: Xid (PCI:0000:19:00): 43, pid=79302, name=nvvs, Ch 00000009
# journalctl -k | grep -i xid May 04 20:30:00 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 1): Out Of Range Address May 04 20:30:00 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a7b0=0x100000e 0x56a7b4=0x20 0x56a7a8=0x1f81fb60 0x56a7ac=0x1174 May 04 20:30:00 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 43, pid=8548, name=nvvs, Ch 00000009 May 04 20:38:23 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 3, SM 0): Out Of Range Address May 04 20:38:23 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x568730=0x100000e 0x568734=0x20 0x568728=0x1f81fb60 0x56872c=0x1174 May 04 20:38:23 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 43, pid=10603, name=nvvs, Ch 00000009 May 05 09:02:43 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 0): Out Of Range Address May 05 09:02:43 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a730=0x103000e 0x56a734=0x20 0x56a728=0x1f81fb60 0x56a72c=0x1174 May 05 09:02:43 xx110-r113-node-02 kernel: NVRM: Xid (PCI:0000:19:00): 43, pid=79302, name=nvvs, Ch 00000009 # grep -i xid /var/log/syslog May 4 20:30:00 xx110-r113-node-02 kernel: [ 747.179157] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 1): Out Of Range Address May 4 20:30:00 xx110-r113-node-02 kernel: [ 747.180533] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a7b0=0x100000e 0x56a7b4=0x20 0x56a7a8=0x1f81fb60 0x56a7ac=0x1174 May 4 20:30:00 xx110-r113-node-02 kernel: [ 747.209815] NVRM: Xid (PCI:0000:19:00): 43, pid=8548, name=nvvs, Ch 00000009 May 4 20:38:23 xx110-r113-node-02 kernel: [ 1250.627528] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 3, SM 0): Out Of Range Address May 4 20:38:23 xx110-r113-node-02 kernel: [ 1250.629013] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x568730=0x100000e 0x568734=0x20 0x568728=0x1f81fb60 0x56872c=0x1174 May 4 20:38:23 xx110-r113-node-02 kernel: [ 1250.657381] NVRM: Xid (PCI:0000:19:00): 43, pid=10603, name=nvvs, Ch 00000009 May 4 20:40:44 xx110-r113-node-02 drpcli[4139]: Starting xid error detection test... May 4 20:40:44 xx110-r113-node-02 drpcli[4139]: [MANDATORY] test_gpu_xid_errors: PASS, GPU XID error check passed. No errors found. May 5 09:02:43 xx110-r113-node-02 kernel: [45911.449627] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics SM Warp Exception on (GPC 6, TPC 5, SM 0): Out Of Range Address May 5 09:02:43 xx110-r113-node-02 kernel: [45911.451042] NVRM: Xid (PCI:0000:19:00): 13, pid='<unknown>', name=<unknown>, Graphics Exception: ESR 0x56a730=0x103000e 0x56a734=0x20 0x56a728=0x1f81fb60 0x56a72c=0x1174 May 5 09:02:43 xx110-r113-node-02 kernel: [45911.479823] NVRM: Xid (PCI:0000:19:00): 43, pid=79302, name=nvvs, Ch 00000009

Common XID errors in H100 NVL GPUs


Application/ CUDA errors: XID 11/25/32/37/69/80 are often caused by application bugs. These are typically recoverable after the application restart.

Memory/ ECC errors: XID 48/64/94/95/140 are caused by GPU memory/ ECC/ remapping related errors or events. Immediate action is to reset the GPU, and if the problem persists, contact your hardware vendor.

NVLink fabric fault: XID 74 indicates a problem with a connection from the GPU to another GPU or NVSwitch over NVLink. A GPU reset or node reboot is needed to clear this error. This event may indicate a hardware failure with the link itself or may indicate a problem with the device at the remote end of the link. For example, if a GPU fails, another GPU connected to it over NVLink may report an XID 74 simply because the link went down as a result. The nvidia-smi nvlink command can provide additional details on NVLink errors, and connection information on the links. If this error is seen repeatedly and GPU reset or node reboot fails to clear the condition, contact your hardware vendor for support.

Here is the full list of XIDs, including their applicability across platforms (H100, B100, GB200, etc.): Analyzing Xid Errors with the Xid Catalog — XID Errors

References

Saturday, March 7, 2026

Working with GPUs – A Practical Blog Series

This blog series captures practical learnings from working with GPUs in real‑world environments, with a focus on operations, reliability, and scale. Each post deep‑dives into specific aspects of GPU systems based on hands‑on experience, incidents, and operational challenges. Together, these articles aim to share actionable insights, highlight common pitfalls, and help teams build more robust and predictable GPU operations.


Part 01: Using nvidia-smi
Part 02: Memory fault indicators
Part 03: Using dcgmi
Part 04: Thermal issues
Part 05: XID errors
Part 06: H100 SXM5 architecture
Part 07: GPU has fallen off the bus
Part 08: Powering and cooling AI accelerators with OCP ORV3 HPR
Part 09: Power shelves in OCP ORV3
Part 10: Energy storage systems, BBUs, and CBUs