The Power Struggle
2026-09-13Nearly a year of hard GPU crashes turned into a hardware investigation with one recurring problem: the graphics card was always where the machine failed first, but it was not the part that was broken. The fault was somewhere in the power-delivery path, which could look healthy right up until it did not.
00The symptom
The machine did not politely crash. It hard-froze: display stuck, system unresponsive, usually little or nothing useful after the reboot. The Sapphire Nitro+ RX 7900 XTX was the obvious suspect. It was the hottest and most power-hungry component in the Ryzen 7 5800X system, and it was present in every failure.
That was enough to start with. It was not enough to conclude anything.
01The heat hypothesis
There was a real cooling problem. Before tuning, the edge temperature was around 78 C while GPU junction had reached 110 C. That difference made the cooler a legitimate suspect, so the first job was to turn the GPU into an instrument rather than a mystery.
MangoHud recorded edge, junction and memory temperatures, board power, fan speed, PWM, clocks and voltage every second. LACT supplied a junction-based fan curve and reduced the card to its minimum supported power limit of 305 W from a 305-389 W range.
The power cap made the card cooler. It did not explain why the entire machine sometimes stopped responding.
02Trying to make it fail
An intermittent failure needs repeatable workloads. FurMark 2 ran at the monitor's native 3440x1440 resolution with its artifact scanner enabled; every result below is from the same card that had been blamed for the freezes.
385,144 frames at 214 FPS average and 100 percent GPU load. Board power averaged 296.9 W and sampled at 341 W. Junction averaged 96.7 C and peaked at 99 C; edge averaged 72.1 C and peaked at 73 C. VRAM averaged 75.3 C and peaked at 76 C. No artifacts, amdgpu reset, ring timeout or kernel error.
Two complete full-load runs followed by 15-minute cooldowns, plus most of a third hot phase. Both completed runs cooled to 57 C junction and ended clean.
The same hard freeze happened outside Linux and without a major stress workload. Reseating the card made it disappear temporarily. That weakened a driver-only explanation and shifted attention to the PCIe path and power delivery.
A card that can survive 30 minutes at native resolution, multiple hot/cold cycles and clean PCIe checks is not cleared. But a simple thermal shutdown or permanently damaged PCIe link becomes much less convincing.
03Making cooling boring
The thermal work still mattered. I replaced the GPU interface material with PTM and kept the 305 W cap, 2,700 MHz maximum core clock and the custom fan curve. A final guarded FurMark pass provided a baseline for a properly cooled, fully loaded card.
The card was no longer a thermal wildcard. The random freeze still had no convincing explanation.
04The power path
The cross-OS freeze and temporary reseat effect left a wider suspect list: PCIe edge and slot contact, card sag, a damaged trace, GPU power connectors, extension cables and the PSU. The next suspect was the power-delivery path.
The replacement was intentionally a clean intervention. The GPU was left seated. Every ATX and PCIe extension cable came out. The replacement Corsair PSU used only its supplied cables, and the 7900 XTX received three separate PCIe 8-pin runs rather than sharing a daisy-chained cable.
The new 3.3 V reading stayed close to nominal: its lowest observed value was 36 mV below 3.3 V, even while the GPU was drawing at least 250 W. The equivalent readings with the old PSU had sat materially farther from 3.3 V. This is motherboard telemetry rather than an oscilloscope trace, but it was another recorded change that pointed at power delivery rather than the GPU.
One clean 17-minute run was not proof on its own because the old fault was intermittent. It was the first normal result after changing power delivery, without reseating the GPU that had spent nearly a year taking the blame.
05The verdict
The evidence points to a power-delivery fault. The GPU had a real thermal issue and it deserved fixing, but it could also survive controlled 305 W loads. The failures crossed operating systems, left the PCIe counters clean and stopped following the power-supply replacement and cable cleanup. Because the PSU, extensions and power cables changed together, this test cannot isolate one component. The GPU was reporting a power problem.
06What this actually taught me
- Collect telemetry before changing hardware. The one-hertz logs separated a real heat problem from the cause of the freezes.
- Stress tests prove a narrow claim. A 30-minute clean benchmark only proves the machine survived that workload at that moment.
- Cross-OS failures are powerful evidence. Windows freezing too made a Linux driver theory much less useful.
- Power delivery is part of the GPU. The card, its connectors, extensions, cables and PSU form one failure domain under load.
Hardware: Sapphire Nitro+ RX 7900 XTX Vapor-X / Ryzen 7 5800X · Logging: MangoHud 1 Hz CSV + LACT · Test workloads: FurMark 2 and The Witcher 3
</post>