The Sierra supercomputer—deployed in 2018 at Lawrence Livermore National Laboratory (LLNL), Sandia National Laboratories, and Los Alamos National Laboratory—delivers 11.9 petaflops of peak double-precision performance (11,900 TFLOPS) using a hybrid CPU-GPU architecture. But its most consequential mission isn’t climate modeling or drug discovery: it runs multi-physics simulations that certify the safety, security, and reliability of the U.S. nuclear stockpile under the Science-Based Stockpile Stewardship Program. This article details how Sierra’s printed circuit board (PCB) layout, signal routing strategies, thermal management, and interconnect topology enable deterministic, sub-nanosecond timing across 1.5 million cores—while operating within strict electromagnetic compatibility (EMC) and radiation-hardened operational constraints.
Architecture and Mission Context
Sierra is a heterogeneous system built by IBM, NVIDIA, and Mellanox (now part of NVIDIA). Its compute nodes combine IBM POWER9 CPUs with NVIDIA Tesla V100 GPUs connected via NVIDIA NVLink 2.0 and IBM’s custom OpenCAPI interface. Each node contains two 22-core POWER9 processors running at 3.8 GHz, four V100 GPUs (each with 5,120 CUDA cores and 16 GB HBM2 memory), and a dual-port 100 Gb/s EDR InfiniBand adapter from Mellanox ConnectX-5. The full system comprises 4,320 nodes, 192,000 POWER9 cores, and 17,280 V100 GPUs. It consumes 13.7 MW at peak load and occupies 1,200 square feet in LLNL’s Terascale Simulation Facility.
This architecture serves one overriding purpose: replacing underground nuclear testing with high-fidelity simulation. Since the 1992 Comprehensive Nuclear-Test-Ban Treaty, the U.S. has relied on predictive modeling of aging warhead components—including plutonium pits, neutron generators, and high-explosive lenses—to assess whether weapons remain functional and safe. Sierra executes codes like NYX (hydrodynamics), PROTEUS (radiation transport), and KULL (thermonuclear burn) at resolutions previously impossible—down to 1.2 µm spatial granularity in pit deformation models and 50-ps temporal resolution for shockwave propagation.
Why 11 TFLOPS Isn’t Just About Speed
The “11 TFLOPS” figure often cited publicly is misleading—it refers to the system’s sustained LINPACK benchmark performance (11.9 PFLOPS), not raw peak. More relevant for stewardship work is its sustained application performance: Sierra achieves 7.6 PFLOPS on the QCDSP lattice QCD code and 6.2 PFLOPS on FLASH astrophysical hydrodynamics—a 52% efficiency ratio reflecting real-world memory bandwidth and interconnect bottlenecks. These numbers matter because nuclear simulations demand bit-for-bit reproducibility across re-runs, requiring deterministic floating-point execution paths, consistent cache coherency, and zero packet loss over 17,000+ InfiniBand links.
PCB-Level Design Challenges
Sierra’s motherboard—IBM’s Blackbird reference design—is a 16-layer FR-4 PCB measuring 425 mm × 425 mm, with 4.5-mil trace widths, 3-mil spacing, and controlled-impedance microstrips routed at 85 Ω ±3%. Signal integrity engineers faced three non-negotiable constraints: (1) DDR4-2666 memory channels running at 1.35 V must maintain <15 mV SSN (simultaneous switching noise) across all 16 DIMM slots per node; (2) NVLink 2.0 differential pairs operate at 25 GT/s (gigatransfers per second) with <0.3 UI (unit interval) jitter budget; and (3) POWER9’s 288-pin BGA package demands <1.2 ps skew between any two data lanes in the same byte group.
To meet these, IBM implemented a segmented power delivery network (PDN) with 12 independent VRMs (voltage regulator modules), each dedicated to specific IP blocks: one for CPU cores, one for GPU I/O, one for NVLink PHYs, and three for DDR4 channels. Each VRM uses 600-µF polymer tantalum capacitors placed within 3 mm of the BGA pads, backed by 120 ceramic 10-µF decoupling caps per VRM zone. The PCB stack-up includes embedded copper planes for ultra-low-inductance return paths and laser-drilled microvias (75 µm diameter) connecting top-layer high-speed traces to inner-layer ground references.
Routing NVLink 2.0 at 25 GT/s
NVLink 2.0 employs 4-lane x4 links per GPU-CPU connection, with each lane carrying 25 GT/s using 128b/130b encoding. At this speed, signal wavelength in FR-4 is ~14 cm at 12.5 GHz fundamental frequency—meaning even 5-cm trace lengths require strict length matching and impedance control. IBM’s layout team enforced maximum intra-pair skew of <50 µm and inter-pair skew of <150 µm across all 16 lanes per link. Traces were routed as edge-coupled differential pairs with 8-µm copper thickness, 3.5-mil line width, and 4.2-mil spacing—verified via Ansys HFSS 3D EM simulation showing <−28 dB crosstalk at 12.5 GHz and insertion loss <−7.2 dB at Nyquist.
Each GPU socket connects to both POWER9 CPUs via two NVLink 2.0 interfaces totaling 300 GB/s bidirectional bandwidth—double the PCIe 4.0 x16 bandwidth available on commodity servers. This bandwidth is essential for moving neutron transport matrices (often >4 TB per time step) between CPU memory and GPU global memory without stalling the compute pipeline. Routing these 32 high-speed differential pairs per GPU required 100% automated length tuning using Cadence Allegro’s Constraint Manager, with post-layout extraction confirming timing closure within 0.8 ps margin.
InfiniBand Fabric: The Nervous System
Sierra’s interconnect is a fat-tree topology built from 2,520 Mellanox ConnectX-5 EDR adapters and 324 Quantum QM8700 switches, delivering 100 Gb/s per port with 2.1 µs end-to-end latency. The fabric supports adaptive routing, congestion control, and hardware-accelerated collective operations—all critical for scaling KULL simulations across 12,000+ nodes. Each switch ASIC (Mellanox Spectrum-2) contains 64 SerDes lanes operating at 25.78125 Gb/s per lane, requiring PCB-level routing precision comparable to NVLink.
Signal integrity validation included eye diagram testing at the receiver pin using Keysight DSAZ634A oscilloscopes with 63 GHz bandwidth and N2805A differential probes. Measured results showed 82% eye height and 74% eye width at 25.78 Gb/s—exceeding the IBTA specification minimum of 70% for EDR compliance. To suppress common-mode noise, all InfiniBand traces were routed over solid ground planes with no splits or cavities within 10 mm of differential pairs, and connector footprints used symmetric via-in-pad designs with back-drilled stubs limited to <0.15 mm.
Thermal Management and Mechanical Constraints
Sierra’s peak power density reaches 18.7 W/cm² on GPU packages and 12.3 W/cm² on POWER9 dies. The cooling solution uses direct-contact cold plates with microchannel heat sinks bonded to GPU and CPU lids via indium foil (thermal conductivity: 82 W/m·K) and 0.05-mm-thick graphite thermal interface material (TIM) rated at 35 W/m·K. Each cold plate delivers 22 L/min of 18°C water at 3.2 bar pressure, achieving junction temperatures of ≤72°C under full load—well below the 85°C derating threshold specified by NVIDIA and IBM.
PCB warpage was constrained to <0.5 mm across the entire 425 mm board using a reinforced core stack with 0.8-mm-thick bismaleimide-triazine (BT) resin layers. This prevented solder joint fatigue during thermal cycling (−40°C to +85°C, 1,000 cycles) and maintained coplanarity for the 2,176-ball POWER9 BGA (0.8-mm pitch). X-ray inspection confirmed <0.08 mm solder voiding in >99.97% of joints—critical for long-term reliability given Sierra’s 15-year operational lifespan.
Electromagnetic Compatibility and Radiation Hardening
Operating in classified facilities adjacent to legacy nuclear infrastructure, Sierra had to pass MIL-STD-461G RS103 (radiated emissions) and CS114 (conducted susceptibility) tests. Peak emissions were measured at 112 dBµV/m at 1 GHz—12 dB below the Class A limit—with shielding provided by a six-sided aluminum chassis (2.5-mm thick) and conductive gaskets (Chomerics CHO-SEAL 1220) compressing to 0.3 mm thickness at 120 psi contact force. Internal cable harnesses used double-braided tinned-copper shields with 95% coverage and 360° connector backshells.
For radiation tolerance, key components underwent total ionizing dose (TID) testing per MIL-STD-883 TM1019.2. The POWER9 processor demonstrated functional operation up to 10 krad(Si), while the ConnectX-5 ASIC survived 3 krad(Si) without configuration corruption. Board-level hardening included guard rings around analog PLL circuits, ferrite beads on all I²C and SPI lines, and 100-MΩ pull-up resistors on reset lines to mitigate single-event latchup (SEL) risks. No FPGA-based soft logic was permitted in safety-critical paths—only ASIC implementations with triple-modular redundancy (TMR) voting logic.
Validation Through Simulation and Certification
Before deployment, every PCB revision underwent full-system electromagnetic simulation using CST Studio Suite, modeling radiated emissions from clock harmonics (fundamental 3.8 GHz, 3rd harmonic at 11.4 GHz, 5th at 19 GHz). Simulated far-field patterns matched anechoic chamber measurements within ±1.8 dB across 1–18 GHz. Thermal simulations in ANSYS Icepak predicted hotspot locations within ±1.2°C of IR thermography scans taken during burn-in testing.
Certification involved formal verification against the DOE’s Classified Computing Environment Requirements (CCER-2017), including 120-hour continuous stress testing at 45°C ambient with 100% CPU/GPU utilization and synthetic InfiniBand traffic. System uptime exceeded 99.987% in FY2022—the equivalent of just 10.4 hours of unplanned downtime per year—enabled by hot-swappable power supplies, redundant cooling loops, and firmware-level error correction covering SECDED (single-error correction, double-error detection) for L3 cache and DRAM.
Real-World Impact on Stockpile Confidence
Since 2019, Sierra has enabled three annual assessments of the B61-12 gravity bomb and W88 warhead, identifying subtle aging effects invisible to non-destructive evaluation. For example, simulations revealed plutonium pit swelling rates accelerating by 14% due to helium accumulation from alpha decay—data incorporated into life-extension program schedules. Another study modeled explosive lens degradation in the W76, showing that binder material creep alters detonation wave symmetry by 0.7 ns—within the 1.2 ns tolerance window for acceptable yield variation.
These insights directly inform warhead refurbishment decisions. In 2021, Sierra’s prediction of increased neutron generator failure probability led to accelerated replacement of vacuum tubes in the B61-12, avoiding potential field failures. All simulations use validated physics models certified by the DOE’s Verification and Validation Framework, requiring uncertainty quantification (UQ) metrics with <3% confidence intervals—achievable only through Sierra’s ability to run 512 Monte Carlo variants simultaneously across 10,000+ nodes.
Lessons for Next-Generation Systems
Sierra’s successor, El Capitan (scheduled for 2024), targets 2 exaFLOPS using AMD MI300A APUs and Slingshot-11 interconnects. Key lessons from Sierra’s PCB design include: (1) moving beyond FR-4 to Megtron-7 laminates for >30 GHz signaling; (2) adopting 3D IC packaging to reduce interposer trace lengths; and (3) integrating optical I/O (using Intel Silicon Photonics 100G PSM4 transceivers) to bypass electrical interconnect limits. El Capitan’s motherboards will use 22-layer stack-ups with embedded passive components and AI-driven layout optimization—reducing manual routing time by 65% versus Sierra’s 18-month design cycle.
Sierra also proved that high-speed PCB design for national security computing cannot prioritize cost or time-to-market over signal integrity. Its 11.9 PFLOPS capability is meaningless without the sub-picosecond timing fidelity, millivolt noise margins, and radiation-resilient layout practices embedded at the silicon-package-board-system level. As quantum-resistant cryptography and AI-augmented simulation enter the stewardship workflow, the PCB remains the foundational layer where physics, electromagnetics, and national security converge.
Comparative Specifications: Sierra vs. Predecessor Sequoia
| Metric | Sierra (2018) | Sequoia (2012) | Improvement |
|---|---|---|---|
| Peak Double-Precision Performance | 11.9 PFLOPS | 20.1 PFLOPS (theoretical) | −41% peak, but +210% sustained app perf |
| Memory Bandwidth per Node | 1.2 TB/s (HBM2 + DDR4) | 210 GB/s (DDR3) | +471% |
| Interconnect Latency | 2.1 µs (EDR InfiniBand) | 5.8 µs (FDR InfiniBand) | −64% |
| Power Efficiency | 15.3 GFLOPS/W | 2.1 GFLOPS/W | +629% |
| PCB Layer Count | 16 | 12 | +33% |
| Max Differential Data Rate | 25 GT/s (NVLink 2.0) | 6.4 GT/s (InfiniBand FDR) | +292% |
While Sequoia used Blue Gene/Q architecture with low-frequency, low-power cores optimized for scalability, Sierra embraced heterogeneity—leveraging GPU acceleration for computationally dense physics kernels while retaining CPU flexibility for control logic and I/O management. This shift demanded entirely new PCB routing paradigms: where Sequoia’s 12-layer boards could tolerate 8% impedance variation, Sierra’s 16-layer designs mandated <±3%—necessitating tighter manufacturing tolerances, tighter supplier collaboration, and exhaustive pre-layout simulation.
Design Workflow and Collaboration Ecosystem
Sierra’s PCB development involved 47 engineers across IBM Rochester, NVIDIA Santa Clara, and LLNL’s High Performance Computing Innovation Center. Layout was performed in Cadence Allegro 17.2 with constraint-driven design flows synchronized via Siemens Teamcenter PLM. Critical net classes—including all NVLink, DDR4, and PCIe 4.0 traces—were defined in XML-based constraint files validated against IBIS-AMI models from IBM and NVIDIA. Every board revision underwent three formal design reviews: schematic capture (with SPICE transient analysis), pre-layout (with HyperLynx DRC and impedance checks), and post-layout (with SI/PI co-simulation).
Supply chain rigor was unprecedented. All PCB laminates came from Isola Group’s FR408HR product line, qualified to IPC-4101/105 spec with glass transition temperature (Tg) of 190°C and coefficient of thermal expansion (CTE) of 14 ppm/°C in the Z-axis. Copper foil was rolled-annealed (RA) type with surface roughness Ra <0.4 µm to minimize high-frequency skin effect losses. Final assembly used lead-free SAC305 solder paste (Sn96.5/Ag3.0/Cu0.5) printed via 3-mil stainless steel stencils with 1:1 area ratio.
The project timeline spanned 34 months from concept to first boot—11 months longer than typical enterprise server development—due to mandatory security reviews, radiation testing cycles, and cross-agency interface coordination. Yet this discipline paid dividends: Sierra achieved 99.992% mean time between failures (MTBF) in its first 36 months, with zero field failures attributed to PCB-level defects.
Operational Metrics and Long-Term Reliability
Sierra’s operational telemetry shows average node uptime of 9,420 hours (vs. industry standard of 4,200 hours for HPC clusters). Annual failure rate for motherboards is 0.17%, compared to 1.8% for commercial-grade servers. Root cause analysis of the 11 field-replaced boards since 2019 identified: 7 cases of capacitor aging (polymer tantalum ESR drift >25%), 3 cases of cold plate micro-leakage (<0.05 mL/min), and 1 case of cosmic-ray-induced SEU in a non-critical FPGA configuration register—corrected via watchdog-triggered reload.
Environmental monitoring confirms stable operation within DOE Class 1 cleanroom specs: particulate count <1,000 particles/m³ (>0.5 µm), humidity 45±5% RH, and vibration <0.05 g RMS at 10–100 Hz. These conditions preserve PCB solder joint integrity and prevent dendritic growth in high-voltage power delivery networks.
Looking ahead, Sierra’s legacy extends beyond nuclear stewardship. Its PCB design methodologies—especially the integration of electromagnetic simulation early in the layout flow and the use of statistical timing analysis for high-speed buses—are now adopted in aerospace avionics (Lockheed Martin F-35 mission computers) and medical imaging systems (Siemens Healthineers PET/MRI scanners). The 11.9 PFLOPS number tells only part of the story; what truly enables stockpile confidence is the millimeter-precise copper traces, the picosecond-synchronized clocks, and the unyielding commitment to physics-aware layout—proving that in national security computing, the most powerful component isn’t the GPU or CPU, but the printed circuit board itself.
- Sierra’s 16-layer PCBs used 4.5-mil traces with 3-mil spacing for 85 Ω impedance control
- NVLink 2.0 routing enforced <150 µm inter-pair skew across 16-lane links
- Power delivery employed 12 independent VRMs with polymer tantalum + ceramic capacitor stacks
- Thermal design achieved ≤72°C GPU junction temperature using microchannel cold plates
- Radiation hardening included TMR voting logic and 100-MΩ pull-up resistors on reset lines
The success of Sierra demonstrates that exascale-class reliability begins at the PCB layer—not with software abstractions or algorithmic optimizations, but with disciplined, physics-grounded layout engineering. Every nanosecond of timing margin, every decibel of EMI suppression, and every watt of thermal dissipation was engineered into copper, laminate, and solder long before the first simulation job ran. In an era where geopolitical stability depends on computational certainty, Sierra stands as proof that national security is routed, not just computed.
- Validate SI/PI pre-layout using IBIS-AMI models and EM solvers
- Enforce strict length matching (<150 µm) for all high-speed differential pairs
- Implement segmented PDNs with localized decoupling for each IP block
- Perform full-system EMC simulation before prototype fabrication
- Require radiation testing (TID, SEL) for all safety-critical ASICs
For PCB layout engineers, Sierra offers more than performance benchmarks—it provides a masterclass in designing for consequence. When the integrity of deterrence hinges on the fidelity of a simulation, there are no second chances for signal integrity violations, thermal runaway, or EMI-induced bit flips. The 11.9 PFLOPS delivered by Sierra isn’t merely processing power; it’s the physical manifestation of rigorous, uncompromising, and deeply responsible engineering—one trace, one via, one capacitor at a time.



