Jul 19, 2026 · 6 min read
The AI Boom Runs on Memory, and Last Week the Market Noticed
Founded in 2018 and led by Leah Goldblum, Founder & Creative Director.
Jul 19, 2026 · 6 min read
Founded in 2018 and led by Leah Goldblum, Founder & Creative Director.
If you build products with AI in them, or you sell AI-powered work to clients, last week gave you a rare gift: the physical bottleneck under the whole industry became a headline. I want to translate that headline into something useful for people who ship interfaces, not chips.
Here is the one-sentence version of the thesis I keep repeating to clients: AI capability is not an infinitely scalable cloud resource. It is rationed physical infrastructure. And the ration is currently set by a component called high-bandwidth memory, or HBM.
Start with the week itself. On Friday, July 10, SK Hynix, the South Korean company that held roughly 58 percent of HBM revenue in the most recent quarter per Counterpoint Research, listed American depositary receipts on the Nasdaq under the ticker SKHY. It priced at $149 and raised about $26.5 billion. Reporting called it the largest US listing ever by a foreign company, topping Alibaba’s $25 billion debut in 2014, and the second-largest share sale in US history, trailing only SpaceX’s roughly $86 billion offering the month before. Shares opened at $170 and closed their first day at $168.01, up about 13 percent. For a company most consumers have never heard of, that is an extraordinary amount of money.
Then it reversed. On July 14, SK Hynix’s Korean shares fell 15.4 percent, their worst single day in nearly two decades, and dragged the broader Korean market down enough to trigger circuit breakers. The US-listed shares slid more than 13 percent in a session. Micron fell as much as 13 percent in a single day. Nvidia, AMD, and Broadcom all dropped. The reported trigger was a story that SK Hynix might slow its HBM expansion and shift some capacity toward more conventional DDR5 memory, which investors read as a crack in the “demand is infinite” story. Whether that read is correct is a question for people with a Bloomberg terminal. What matters for us is why a memory chip can move a trillion-dollar market.
So, what is HBM and why is it the choke point? A modern AI accelerator, the expensive chip doing the actual math, is only as fast as the memory feeding it data. HBM solves that by stacking DRAM chips vertically and wiring them together with through-silicon vias, tiny vertical connections that run straight through the silicon, then packaging the whole stack right next to the processor. It is a feat of manufacturing: advanced DRAM, precise vertical stacking, sophisticated packaging, high yields, and a brutal customer qualification process before a company like Nvidia will accept it. You cannot spin up more of it in a quarter. That is the definition of a supply constraint.
The market for it is exploding. Micron has forecast that the total HBM market grows from about $35 billion in 2025 to around $100 billion by 2028, a compound growth rate near 40 percent. On Micron’s fiscal Q1 2026 earnings call in December 2025, CEO Sanjay Mehrotra told investors, “We have completed agreements on price and volume for our entire calendar 2026 HBM supply, including Micron’s industry-leading HBM4.” Read that again: a whole year of a critical AI component, spoken for before it is made.
Market share tells you how concentrated the risk is. Per Counterpoint Research, SK Hynix held around 57 to 62 percent of the HBM market through 2025 depending on the quarter, with Micron near 21 percent and Samsung between 17 and 22 percent. Three companies, on two continents, make essentially all of it. When one of them sneezes, every AI product downstream catches a cold.
That downstream is where you and I live. The demand pulling on this supply is staggering. The four biggest US hyperscalers, Microsoft, Amazon, Alphabet, and Meta, are collectively guiding to roughly $700 billion in capital spending in 2026, up about 77 percent from around $410 billion in 2025. Microsoft has said that roughly $25 billion of its 2026 capital budget is just component price inflation. And yet the returns are still thin. Several analyses this year noted that AI-related services generated only a small fraction of a dollar of revenue for every dollar of infrastructure spending, and Goldman Sachs has flagged that keeping those returns sustainable would eventually require far higher AI profits than current consensus expects. That gap between spending and revenue is exactly the anxiety that showed up in last week’s selloff.
Meanwhile the money is trying to build its way out of the shortage. On July 9, Micron poured the first concrete at its megafab in Clay, New York, more than a quarter ahead of schedule, and raised its planned US investment to more than $250 billion through 2035, up from $200 billion. The New York complex alone is planned at up to $100 billion for four fabs, supported in part by CHIPS Act funding and state incentives, with tens of thousands of jobs projected. But a fab is a multi-year build. None of it relieves the ration this year.
Now the payoff, the reason a UX person should care. Every constraint you have felt in an AI product this year is a symptom of this shortage. When ChatGPT’s image feature went viral in March 2025, Sam Altman posted that “our GPUs are melting” and capped the free tier to three generations a day. In August 2025, Anthropic added weekly usage limits on top of its existing five-hour rolling window for Claude, and its CEO has repeatedly described the company as compute-constrained. Rate limits, tiered pricing, the quiet swap to a smaller “turbo” model at peak times, the gap between a flawless demo and a throttled production experience: none of that is a UX failure by your team. It is the physical ration reaching your interface.
What I tell clients to do with this:
Design for the throttle, not the demo. Assume capacity will be rationed and make the degraded state graceful. A clear “we’re at capacity, here is a lighter option” beats a spinner that never resolves.
Price like memory is scarce, because it is. If your product’s cost of goods is dominated by inference, flat unlimited pricing is a bet against physics.
Do not promise the demo. The demo runs on unthrottled capacity. Production does not. Set client expectations to the ration, and you will look honest instead of broken.
The AI you are shipping is not magic and it is not infinite. It is a stack of memory chips, made by three companies, sold out a year in advance. Once you see it that way, a lot of confusing product decisions start to make sense.
Sources