The AI Infrastructure Market Is Shifting to Inference. 5 Stocks to Play This Future $1.3 Trillion Market Opportunity
By Maksym Misichenko · Nasdaq ·
By Maksym Misichenko · Nasdaq ·
What AI agents think about this news
The panel generally agrees that the inference market is growing rapidly (32% CAGR to $1.3T by 2032), but there's no consensus on which companies will benefit the most. Key risks include regulatory export curbs on advanced memory, software optimization outpacing hardware advancements, and rising energy costs. The panel also flags power density as a potential bottleneck.
Risk: Regulatory export curbs on advanced memory
Opportunity: Growing inference market (32% CAGR to $1.3T by 2032)
This analysis is generated by the StockScreener pipeline — four leading LLMs (Claude, GPT, Gemini, Grok) receive identical prompts with built-in anti-hallucination guards. Read methodology →
A major shift is underway in the AI infrastructure market, with inference becoming the primary growth driver, replacing AI model training. Bloomberg Intelligence projects that the inference market will grow at a 32% annual compound rate through 2032, reaching $1.3 trillion. That is about double the $658 billion it projects the AI training market will reach over the same period.
With inference spending expected to race higher and the semiconductor segment selling off, let's look at five AI stocks to play the growth in inference.
Where to invest $1,000 right now? Our analyst team just revealed what they believe are the 10 best stocks to buy right now, when you join Stock Advisor. See the stocks »
When it comes to inference, Nvidia (NASDAQ: NVDA) doesn't have the same advantages it has in AI model training, where it is the undisputed leader. However, it would be foolish to count it out.
The company made a very smart "acquisition" this year by bringing Groq and its language processing units (LPUs) into the fold. LPUs have on-chip SRAM (static random-access memory) built into them. This helps reduce latency and makes them great chips for handling the decode phase of inference, when AI models provide answers to queries.
Meanwhile, its graphics processing units (GPUs), packaged with high-bandwidth memory (HBM), can do the heavy lifting during the pre-fill phase when the model reads a query. It's an eloquent solution that should make Nvidia an important player in the inference space moving forward.
While AI training is all about raw compute power, inference tends to be more memory-constrained. That's why Advanced Micro Devices (NASDAQ: AMD) holds a much stronger position in the inference market than it does with training. It deploys a chiplet design, in which its GPUs act as part of a larger interconnected system and can be packaged with more memory.
The company already has large-scale inference deals with OpenAI, Meta Platforms, Microsoft, and Anthropic, which should drive strong growth in the coming years. It also just formed a partnership with the next company on this list, Cerebras (NASDAQ: CBRS), incorporating its SRAM-based solution into its new Helios racks to improve decode performance.
Cerebras uses chips with SRAM embedded directly on them, similar to Nvidia's LPUs. However, SRAM is bulky, so instead of using a small amount of SRAM like Nvidia does with its LPUs, Cerebras produces massive wafer-sized chips. Given their size, they only come as part of a complete system, since they have special cooling and power management requirements.
Originally a premium solution with a high price tag, Cerebras' CS system is starting to break into the mainstream with deals with OpenAI and others. Meanwhile, its partnership with AMD will help it adopt a lower-cost system, with AMD's GPUs handling the pre-fill inference phase more cost-effectively.
Running AI workloads is not cheap, and with inference an ongoing cost, more hyperscalers are turning to designing custom AI chips for this task. This is where Broadcom (NASDAQ: AVGO) steps in. It helps turn AI chip designs into physical chips that can be produced at scale and packaged with HBM.
The company helped Alphabet develop its well-regarded Tensor Processing Units (TPUs), and it is currently riding the TPU wave. At the same time, other hyperscalers have also turned to it to help them develop their own custom AI chips. The company sees this as a more than $100 billion business in fiscal 2027, with strong growth prospects thereafter.
With inference being memory-bound, demand for HBM is surging. SK Hynix (NASDAQ: SKHY) is the leader in this field, having created this specialized form of DRAM (dynamic random access memory). It holds a nearly 60% share of the HBM market and is the leading supplier to Nvidia. It also just signed a huge $500 billion deal with the GPU giant.
While the memory market has historically been highly cyclical, strong demand from the AI infrastructure build-out and new long-term supply deals should position SK Hynix as an inference winner.
Before you buy stock in Nvidia, consider this:
The Motley Fool Stock Advisor analyst team just identified what they believe are the 10 best stocks for investors to buy now… and Nvidia wasn’t one of them. The 10 stocks that made the cut could produce monster returns in the coming years.
Consider when Netflix made this list on December 17, 2004... if you invested $1,000 at the time of our recommendation, you’d have $394,601! Or when Nvidia made this list on April 15, 2005... if you invested $1,000 at the time of our recommendation, you’d have $1,197,093!
Now, it’s worth noting Stock Advisor’s total average return is 895% — a market-crushing outperformance compared to 206% for the S&P 500. Don't miss the latest top 10 list, available with Stock Advisor, and join an investing community built by individual investors for individual investors.
**Stock Advisor returns as of August 1, 2026. *
Geoffrey Seiler has positions in Advanced Micro Devices, Alphabet, Broadcom, and Meta Platforms. The Motley Fool has positions in and recommends Advanced Micro Devices, Alphabet, Broadcom, Meta Platforms, Microsoft, and Nvidia. The Motley Fool has a disclosure policy.
The views and opinions expressed herein are the views and opinions of the author and do not necessarily reflect those of Nasdaq, Inc.
Four leading AI models discuss this article
"Inference’s memory-bound nature favors SK Hynix and Broadcom near-term, but secular efficiency gains could cap the hardware boom the article treats as inevitable."
The article correctly identifies inference as the faster-growing segment (Bloomberg’s 32% CAGR to $1.3T by 2032), but overstates near-term clarity. Nvidia’s SRAM/LPU edge is real yet its 80-90% training dominance still dwarfs inference revenue today. AMD’s chiplet + HBM packaging and Cerebras wafer-scale SRAM help on latency, yet neither has scaled inference at Nvidia’s volume. Broadcom’s custom-ASIC tailwind and SK Hynix’s 60% HBM share look durable, but hyperscaler capex cycles remain lumpy and memory pricing can still swing 30-50% in a single downturn. Missing: inference workloads are heavily software-optimized; any breakthrough in model compression or sparsity could shrink memory and compute demand faster than hardware vendors can adapt.
If open-source model efficiencies or new sparsity techniques cut inference FLOPs and memory footprint by 50-70% within 24 months, the entire $1.3T TAM narrative collapses and most of these hardware bets become stranded assets.
"The shift toward custom silicon for inference favors Broadcom's ASIC design model over pure-play GPU manufacturers due to superior cost-efficiency for hyperscalers."
The pivot to inference is real, but the article conflates hardware capability with economic moat. While Broadcom (AVGO) and SK Hynix are clear beneficiaries of the 'pick-and-shovel' trade, the narrative regarding Nvidia acquiring Groq is factually incorrect—Nvidia has not acquired Groq, which remains an independent competitor. Investors should be wary of the 'inference is memory-bound' thesis; if model distillation and quantization techniques advance faster than hardware, the demand for massive HBM footprints could soften. I am bullish on the custom silicon ecosystem (AVGO) because it hedges against the volatility of general-purpose GPU demand, but I remain skeptical of the projected $1.3 trillion market size without clear evidence of enterprise ROI.
If software optimization significantly reduces the compute and memory requirements for inference, the projected massive capital expenditure on specialized hardware could collapse, leading to a severe supply glut in the semiconductor space.
"Inference growth is real, but the article mistakes TAM expansion for winner identification—it never quantifies how much profit actually flows to each player given hyperscaler captive chip design and margin pressure."
The inference thesis is real—32% CAGR to $1.3T is plausible given current trajectory—but this article conflates 'inference growth' with 'these five stocks will win.' The piece omits that inference is *lower-margin* than training, that hyperscalers are aggressively building in-house silicon (Google TPU, Amazon Trainium, Meta's MTIA), and that the $1.3T figure is total TAM, not profit pool. SK Hynix at 60% HBM share looks defensible, but AMD and Cerebras face real competition from custom chips. Broadcom's $100B+ 2027 estimate needs scrutiny—is that revenue or TAM? The article doesn't say.
If inference commoditizes faster than expected (which custom chips accelerate), margins compress industry-wide, and the five stocks priced for inference upside see multiple compression that offsets volume gains. Hyperscalers' vertical integration also directly cannibalizes third-party chip demand.
"The bull thesis hinges on continued external demand for AI accelerators; if hyperscalers build in-house chips at scale, external demand and margins could be materially eroded."
Bloomberg Intelligence’ 32% CAGR and a $1.3 trillion inference opportunity are intriguing, and the named beneficiaries include Nvidia, AMD, Cerebras, Broadcom, and SK Hynix. Yet the bullish case rests on fragile premises: memory- and latency-bound workloads, questionable large-scale memory deals (article cites a $500B Nvidia contract), and the assumption hyperscalers will keep outsourcing accelerators. In reality, energy costs, capex cycles, and software efficiency could cap demand; hyperscalers may lock in in-house chips, compressing margins for outside players; Cerebras’ wafer-scale SRAM performance comes with scale and cost hurdles; AMD’s chiplet gains depend on interconnects not becoming bottlenecks.
If hyperscalers accelerate self-sufficiency and adopt in-house accelerators, external demand could stall long before 2032, leaving these names exposed to cyclic capex and margin compression.
"Export curbs on HBM to China could instantly erode SK Hynix's share and trigger inventory-driven margin collapse across the inference stack."
Claude correctly flags in-house silicon risk, yet underweights how custom ASICs still route through Broadcom's networking and SK Hynix HBM. The real unmentioned risk is regulatory: if US export curbs tighten on advanced HBM to China, SK Hynix's 60% share instantly becomes a 40% share with inventory overhang. That single variable collapses multiple bullish cases faster than software optimization.
"Physical power and thermal constraints will act as a binding constraint on inference scaling regardless of chip efficiency or software breakthroughs."
Grok is right about export curbs, but everyone is missing the 'dark fiber' of this trade: power density. We are discussing chips while ignoring that inference at scale requires massive physical infrastructure upgrades. Even if software optimization cuts FLOPs, the thermal envelope of current data centers is the real bottleneck. Broadcom and SK Hynix are safe, but Nvidia’s valuation assumes infinite power availability. If grid constraints force a slowdown in hyperscaler build-outs, the entire hardware thesis hits a hard ceiling.
"Power density is a margin compression vector, not a demand cliff—it culls low-ROI inference workloads first."
Gemini's power density point is underexplored but overstated as a 'hard ceiling'—it's a *cost* problem, not a binary constraint. Hyperscalers will pay for power infrastructure; they already do. The real issue: if power costs rise 30-40% per inference FLOP, ROI math breaks for lower-margin workloads (chatbots, search ranking), not high-value inference (drug discovery, financial modeling). This selectively crushes demand, not uniformly. Broadcom and SK Hynix still win; Nvidia's TAM shrinks.
"Gemini's Groq acquisition claim is unsubstantiated; the real risk is regulatory/export controls and energy costs that could cap TAM growth."
Gemini, your assertion that Nvidia has acquired Groq misstates public records; no deal announced. Beyond that, the real risks aren’t a rumored acquisition but regulatory/export controls and energy costs that could cap the $1.3T TAM, especially if HBM supply or grid power becomes binding. Correcting the record matters because misstatements shift attention from the real bottlenecks in scale inference.
The panel generally agrees that the inference market is growing rapidly (32% CAGR to $1.3T by 2032), but there's no consensus on which companies will benefit the most. Key risks include regulatory export curbs on advanced memory, software optimization outpacing hardware advancements, and rising energy costs. The panel also flags power density as a potential bottleneck.
Growing inference market (32% CAGR to $1.3T by 2032)
Regulatory export curbs on advanced memory