The Power Struggle

Nearly a year of hard GPU crashes turned into a hardware investigation with one recurring problem: the graphics card was always where the machine failed first, but it was not the part that was broken. The fault was somewhere in the power-delivery path, which could look healthy right up until it did not.

RX 7900 XTXSapphire Nitro+ Vapor-X
Ryzen 7 5800XGaming System
1 HzTelemetry Logging
305 WGPU Power Limit

00The symptom

The machine did not politely crash. It hard-froze: display stuck, system unresponsive, usually little or nothing useful after the reboot. The Sapphire Nitro+ RX 7900 XTX was the obvious suspect. It was the hottest and most power-hungry component in the Ryzen 7 5800X system, and it was present in every failure.

That was enough to start with. It was not enough to conclude anything.

01The heat hypothesis

There was a real cooling problem. Before tuning, the edge temperature was around 78 C while GPU junction had reached 110 C. That difference made the cooler a legitimate suspect, so the first job was to turn the GPU into an instrument rather than a mystery.

MangoHud recorded edge, junction and memory temperatures, board power, fan speed, PWM, clocks and voltage every second. LACT supplied a junction-based fan curve and reduced the card to its minimum supported power limit of 305 W from a 305-389 W range.

thermal logger · loaded sessions
before 305 W cap junction avg/max: 99.6 C / 103 C hotspot delta avg/max: 25 C / 29 C fan average: 3,034 RPM after 305 W cap duration: 12 min junction avg/max: 95.9 C / 99 C edge maximum: 74 C hotspot delta avg/max: 23.5 C / 29 C fan average: 3,008 RPM

The power cap made the card cooler. It did not explain why the entire machine sometimes stopped responding.

02Trying to make it fail

An intermittent failure needs repeatable workloads. FurMark 2 ran at the monitor's native 3440x1440 resolution with its artifact scanner enabled; every result below is from the same card that had been blamed for the freezes.

30 min · FurMark 2 · 3440x1440 · clean

385,144 frames at 214 FPS average and 100 percent GPU load. Board power averaged 296.9 W and sampled at 341 W. Junction averaged 96.7 C and peaked at 99 C; edge averaged 72.1 C and peaked at 73 C. VRAM averaged 75.3 C and peaked at 76 C. No artifacts, amdgpu reset, ring timeout or kernel error.

2 x 20 min · hot/cold cycles · clean

Two complete full-load runs followed by 15-minute cooldowns, plus most of a third hot phase. Both completed runs cooled to 57 C junction and ended clean.

Windows partition · ordinary use · freeze

The same hard freeze happened outside Linux and without a major stress workload. Reseating the card made it disappear temporarily. That weakened a driver-only explanation and shifted attention to the PCIe path and power delivery.

PCIe and storage checks · all clean
link: PCIe Gen4 x16 AER correctable: 0 AER non-fatal: 0 AER fatal: 0 LaneErrStat: 0 Btrfs / NVMe SMART: no relevant errors

A card that can survive 30 minutes at native resolution, multiple hot/cold cycles and clean PCIe checks is not cleared. But a simple thermal shutdown or permanently damaged PCIe link becomes much less convincing.

03Making cooling boring

The thermal work still mattered. I replaced the GPU interface material with PTM and kept the 305 W cap, 2,700 MHz maximum core clock and the custom fan curve. A final guarded FurMark pass provided a baseline for a properly cooled, fully loaded card.

post-PTM FurMark · 120.2 seconds · 238 half-second samples
board power avg/max: 303.7 W / 330 W junction avg/max: 92.6 C / 96 C edge avg/max: 71.1 C / 75 C hotspot delta avg/max:21.5 C / 25 C memory maximum: 80 C

The card was no longer a thermal wildcard. The random freeze still had no convincing explanation.

04The power path

The cross-OS freeze and temporary reseat effect left a wider suspect list: PCIe edge and slot contact, card sag, a damaged trace, GPU power connectors, extension cables and the PSU. The next suspect was the power-delivery path.

The replacement was intentionally a clean intervention. The GPU was left seated. Every ATX and PCIe extension cable came out. The replacement Corsair PSU used only its supplied cables, and the 7900 XTX received three separate PCIe 8-pin runs rather than sharing a daisy-chained cable.

first post-swap test · Witcher 3 streaming
duration: 16 min 52 sec telemetry cadence: 1 reading per second GPU instantaneous PPT:385 W maximum 3.3 V rail (new PSU): 3.264 V min / 3.279 V avg / 3.296 V max VCC at >=250 W GPU: 3.264 V minimum edge / junction / VRAM maximum: 68 C / 96 C / 72 C amdgpu / PCIe / AER / NVMe / Btrfs / watchdog / reset / lockup: none PCIe link: Gen4 x16 LaneErrStat / AER: 0 / empty

The new 3.3 V reading stayed close to nominal: its lowest observed value was 36 mV below 3.3 V, even while the GPU was drawing at least 250 W. The equivalent readings with the old PSU had sat materially farther from 3.3 V. This is motherboard telemetry rather than an oscilloscope trace, but it was another recorded change that pointed at power delivery rather than the GPU.

One clean 17-minute run was not proof on its own because the old fault was intermittent. It was the first normal result after changing power delivery, without reseating the GPU that had spent nearly a year taking the blame.

05The verdict

The evidence points to a power-delivery fault. The GPU had a real thermal issue and it deserved fixing, but it could also survive controlled 305 W loads. The failures crossed operating systems, left the PCIe counters clean and stopped following the power-supply replacement and cable cleanup. Because the PSU, extensions and power cables changed together, this test cannot isolate one component. The GPU was reporting a power problem.

06What this actually taught me

Hardware: Sapphire Nitro+ RX 7900 XTX Vapor-X / Ryzen 7 5800X · Logging: MangoHud 1 Hz CSV + LACT · Test workloads: FurMark 2 and The Witcher 3


</post>