I was about fifteen minutes into a local model parsing run on the Latitude when the fan hit max and just stayed there. I had htop open in one terminal and watch -n 1 sensors looping in another, because at some point you stop trusting the first number you see and start watching it move. Core package temp: 97°C. Then 98. I put my hand on the chassis near the hinge and pulled it back without really deciding to — that's not "warm laptop" hot, that's don't-touch-it hot. And right there on screen, the clock speed didn't step down gracefully. It fell off a cliff, from boosted all the way down into numbers that made the whole run pointless.
The Latitude wasn't broken. It was doing exactly what it's built to do — protect itself the second it runs out of places to put the heat.
That's the part I didn't want to admit for a while: this isn't a software problem. No driver update fixes it, no thermal policy tweak talks your way around it. It's physics, and it comes down to two completely different answers to the same question — how do you get heat off a package fast enough to keep it from cooking itself. The Latitude answers that with one thin heat pipe and a fan not much bigger than a coffee coaster. The Precision answers it with two fans and a vapor chamber built for exactly this kind of sustained load.
I watch enterprise laptops hit thermal limitations that have nothing to do with RAM or swap. The fan's already at full tilt within a minute of the workload starting. The chassis is warm enough to notice through the aluminum. And the tokens per second, whatever they started at, keep drifting down the longer the workload keeps running.
That's not a memory story. That's a heat story, and it plays by a different set of rules entirely.
The Physics Nobody Budgets For
Local inference isn't light work for a CPU. Matrix multiplication across every layer, every generated token, leans hard on the vector execution units — the same AVX-style instruction paths that spike power draw far higher than an ordinary integer-heavy workload ever would. Sustained vector load pulls more current than almost anything else you'll ever throw at that chip.
More current means more heat, and that heat has to go somewhere. Package temperature climbs fast under sustained inference, and most modern mobile CPUs hit their thermal ceiling somewhere around 100°C — the point where the silicon itself is at real risk if nothing steps in to intervene.
That's where the embedded controller takes over. It's not a bug, and it's not optional — it's the chip protecting itself the only way it knows how. The controller pulls core voltage down and cuts the clock multiplier, trading raw speed for a temperature the package can actually survive indefinitely.
There's a second layer under this that most people never think about: OEMs configure two separate power ceilings into the firmware, not one. A short boost limit lets the chip run fast for a few seconds on a burst. A much lower sustained limit is what the chip settles into for anything longer than that — and that sustained number isn't set by what the silicon is capable of. It's set by how much heat the chassis around it can physically remove.
How far the clock actually drops isn't fixed by the chip alone. It depends on one variable above everything else: how fast the cooling system can move that heat back out of the case. That's the whole equation, and it's exactly where a thin corporate laptop and a mobile workstation stop being the same category of machine.
The Teardown: Two Machines, Two Philosophies
A standard Latitude wasn't designed with sustained AI inference in mind. It was designed to be thin, light, and quiet in a conference room — a single fan, a thin copper heat pipe running to a modest fin stack, sized for bursts of office work, not forty-five straight minutes of continuous vector math.
That cooling setup handles short spikes just fine. Open a spreadsheet, compile something small, it recovers cleanly between bursts without ever hitting a wall. Sustained, uninterrupted load is a completely different problem, because there's no recovery window built in — the load never actually lets up long enough for the chassis to catch its breath.
A Precision is a different machine wearing a superficially similar shell. Dell builds its mobile workstation line around dual-fan cooling, and in the higher configurations, a copper vapor chamber in place of a simple heat pipe. That's not a marketing footnote on a spec sheet — it's a fundamentally different heat-transport mechanism doing the actual work.
A heat pipe moves heat along a line, point to point. A vapor chamber spreads it across a plane, using the same phase-change principle over a much larger contact area, which means more of the chip's output gets picked up and carried away before it has anywhere to bottleneck. Two fans also mean more sustained airflow volume moving through that plane, which is what lets the firmware safely configure a higher sustained power limit instead of walking straight into the thermal ceiling within minutes.
The result is exactly what the physics predicts. Under a short burst, the two machines look nearly identical — boost clocks are boost clocks, and both chassis can coast on stored thermal mass for a few seconds. Under a sustained inference loop running real work instead of a synthetic benchmark, the Latitude settles into a noticeably lower equilibrium clock than the Precision manages to hold, simply because the Precision's cooling system can actually keep pace with the heat the chip keeps generating.
This is exactly why Monday's memory work and this piece aren't two separate problems — they're the same underlying constraint showing up at two different layers of the stack. Quantizing your weights and capping your context stops the machine from choking on memory bandwidth. None of that touches how much heat the chip throws off once it's actually computing, and on a thin chassis, that heat finds its own way to slow you back down regardless of how clean your memory budget looks on paper.
The Silicon Thermal Clock Decay Map
Thermal Power Budget Calculator
Compare peak boost power limits against sustained chassis thermal budgets.
Spec Sheet Parameters
Power Budget Metrics
Optional: Paste Logged Stress Test Data
Paste raw CSV data (e.g., from HWiNFO64, Intel XTU, or ThrottleStop). Expected format per line: Time(s), Clock(GHz) or Time(s), Clock(MHz).
Where This Leaves You
None of this makes a Latitude a bad machine. It makes it the wrong machine for sustained inference workloads specifically, which is a different judgment entirely, and one worth making before you buy instead of after your fans have been screaming for three hours straight.
Match the hardware to the job. A conference-room laptop and a compute workstation were never solving the same problem, even when they're sitting in the same product catalog wearing similar aluminum.
Now that the model's running efficiently and the hardware thermals are accounted for I Built an Offline Python Pipeline to Convert Bulk PDFs into Clean Markdown
