Investment thesis · T·03

Semiconductors

Compute becomes infrastructure.

I’m not a semiconductor engineer, and this thesis doesn’t require me to be

I can’t tell you why one lithography approach beats another, and I’m not going to pretend otherwise. What I can do is follow a chain of reasoning about where the world is heading and work out which parts of it are hard to replace.

That chain starts simply. The technologies I’ve written about in T·01 and T·02 all run on computation. Intelligence that scales needs compute. Intelligence that acts in the physical world needs compute in the machine and compute behind it. Everything I believe about the next twenty years assumes an enormous and growing amount of it.

So the question I care about is narrow: will the world need far more computation than it has, and is that need difficult enough to satisfy that the companies satisfying it earn something for their trouble?

The obvious objection, which I want to deal with first

My starting intuition was straightforward. AI needs advanced chips, demand is growing fast, and advanced fabs are extraordinarily difficult, expensive and slow to build. Scarcity plus demand equals pricing power.

The problem is that this makes the thesis depend on chips staying expensive, and I don’t believe they will. The entire history of this industry is falling cost per unit of capability. If I’m relying on scarcity, then the technology succeeding is what breaks my case.

So the harder version of the question: if computation keeps getting dramatically cheaper, does the opportunity shrink, or does cheapness create demand faster than it destroys revenue? I don’t have a settled personal answer. It’s the part of this thesis where I’ve had to lean on evidence rather than instinct.

Cheaper compute has expanded the market, not shrunk it

The cost declines are real and fast. Epoch AI examined how quickly the price of reaching a given level of model performance has fallen, and found the price of matching GPT-4 on PhD-level science questions dropping by roughly 40 times per year, with rates across different capability thresholds ranging from 9 to 900 times annually. Epoch noted the fastest of those declines were also the most recent, so I wouldn’t extrapolate any of them indefinitely. The point that survives is simpler: compute capability is getting dramatically cheaper while total usage rises at the same time.

If falling prices destroyed this market, we’d see it by now. Instead demand has grown faster than cost has fallen. Google reported processing more than 3.2 quadrillion tokens per month as of May 2026, roughly seven times higher than a year earlier, and said that over the preceding twelve months more than 375 of its cloud customers had each processed over a trillion tokens. That second number matters more, because it shows the demand is broad rather than one company’s own products inflating a headline. An analysis by Epoch AI researchers estimates global inference capacity is more than tripling each year while proxies for demand suggest closer to tenfold annual growth, which would mean demand outpacing supply. They’re explicit that these are rough proxies, and that efficiency cuts both ways: cheaper tokens create new uses, and smaller models displace larger ones.

That gap matters more than either number alone. Supply is growing at a rate most industries would consider absurd, and it still isn’t keeping up.

The mechanism seems to be that cheaper compute makes new things economically possible. When a capability costs a hundred times less, applications that made no commercial sense become obvious. Reasoning models spend more compute per question. Agents run for hours instead of seconds. Video and multimodal work is heavier again. Each of these consumes what efficiency saves, and then some. Economists call this the Jevons Paradox, where something getting cheaper to use raises total consumption instead of lowering it, and I’d treat that as a description of what has happened so far rather than a law about what has to happen next.

This is the point where a technology stops being a product and starts being infrastructure. Not because it’s boring, but because everything else begins to depend on it.

Infrastructure is a promise and a warning

I want to be careful with that word, because it cuts both ways and most people only use the flattering half.

Infrastructure means indispensable. It also means, quite often, a poor investment. Roads, railways, power grids and telecom networks are all indispensable and have all destroyed enormous amounts of capital for the people who built them. Being essential is not the same as being profitable.

So “compute becomes infrastructure” is not by itself a reason to own semiconductor companies. It’s a reason to ask a second question: which parts of this stack are genuinely hard to replicate, and are they the parts that capture the money? That’s the same distinction I drew in T·01 between value created and value captured, and it applies here with more force, because this industry has a long history of building capacity into a boom and eating the consequences afterwards.

The bottleneck keeps moving

Here’s what I found most useful in the research, and it isn’t what I expected.

The constraint on AI compute is not simply that there aren’t enough GPUs. Through 2026, advanced packaging rather than wafer fabrication has repeatedly been described as one of the major constraints. TSMC has been expanding its CoWoS packaging capacity aggressively, which is what turns fabricated dies into working accelerators. High-bandwidth memory is a second chokepoint, concentrated in very few suppliers. Power is a third: SpaceX said in its own filing that it believes the key constraints on continued AI growth are physical, naming chip manufacturing, data centre infrastructure and power generation.

A fourth gets less attention and explains more. It isn’t only that fabs are hard to build. The machines that build the chips are among the hardest manufactured objects in existence. Leading-edge chips require extreme ultraviolet lithography, and ASML is the only company on earth that makes those machines. An EUV system depends on an unusually specialised global supplier network, including critical optics that ASML sources exclusively from ZEISS. So the chain runs: AI demand creates chip demand, which creates fab demand, which creates demand for equipment only a handful of companies can produce. That’s why money alone doesn’t buy leading-edge capacity, and it’s the best evidence for my original intuition that this is hard.

The pattern I take from all this is that the constraint moves. It was wafer capacity. Then it was accelerators. Now packaging, memory and electricity are all in the frame. As constraints shift, bargaining power and attractive economics can shift with them, and companies that looked indispensable eighteen months ago can become merely important.

That’s the strongest argument for how I hold this theme rather than what I hold. If I can’t predict where the pressure point sits in five years, I shouldn’t concentrate as though I can.

Why I don’t try to pick the winner here

As of August 2026 I see NVIDIA as the obvious leader, and that isn’t a controversial view today. What matters is that “leader” and “permanent winner” are different claims, and only one is supported.

The case for NVIDIA is real. The software ecosystem around CUDA took fifteen years to build and can’t be replicated by a competitor shipping better silicon. The company sells systems rather than chips, including networking, which makes substitution harder than swapping one part. And it has secured priority access to constrained packaging capacity, which is a supply-chain advantage rather than a technical one.

The case against is equally real. NVIDIA’s revenue has become concentrated in a small number of very large direct customers, and several of those customers are also building their own chips. Google’s TPUs, Amazon’s Trainium, Meta’s and Microsoft’s programmes and Broadcom’s work with AI labs all exist to give their owners an alternative. Those programmes span training and inference rather than sitting neatly in one camp, and the logic is straightforward: a company running its own enormous workloads can design silicon specifically for them, potentially lowering cost, improving efficiency and reducing how much it depends on anyone else. Reported concentration figures vary depending on whether you count direct customers or the end buyers behind them, and NVIDIA doesn’t necessarily know every ultimate user. The tension is clear without a precise number: some of its largest customers have both the money and the reason to need it less.

I’m not predicting how that resolves. It’s the reason I hold semiconductors across several strong companies rather than as a single-company bet. The designers, the foundry, the equipment makers, the memory suppliers and the packaging capacity are all different businesses facing different competition, and I don’t know which of them ends up keeping the money.

Being right about compute is not the same as being right about a chip company.

If another company develops materially better or cheaper technology and ends up better positioned, I’ll move capital toward it. My conviction is in the growth of compute demand, not in any single company’s ability to keep capturing it.

Terafab, and what it does and doesn’t tell me

One thing that sharpened my view is Elon Musk’s Terafab project, though it’s easy to misuse.

Where it stood as of August 2026: announced in March 2026 as a joint Tesla and SpaceX chip manufacturing initiative, confirmed on 6 August 2026 as a site in Grimes County, Texas, with an initial investment of $16.8 billion, more than 100 million square feet planned and at least 3,000 employees. It’s designed to bring logic, memory, packaging and testing together in one place, with Intel involved, and the chips are intended for Tesla’s Optimus robots and Cybercabs and for the space-based data centres SpaceX wants to build. Musk’s reasoning was blunt: both companies will need far more chips than current and future global production can supply.

Three details keep me honest. The headline number at the March announcement was around $25 billion. The $16.8 billion first phase is well below the roughly $55 billion first phase described in SpaceX’s May 2026 IPO filing, which shows how much these numbers move. And that same filing called the arrangement a general framework with no binding commitments and no obligation for either company to stay involved. Tesla’s actual near-term chips remain contracted to Samsung and TSMC.

For context on the difficulty, TSMC now describes total planned Arizona investment of around $265 billion, covering six logic fabs already planned, two advanced-packaging facilities, an R&D centre, and an intention announced in July 2026 to add several more fabs. It has explicitly declined to attach a timeline, saying construction pace will follow demand. That’s a company with decades of experience and every incentive to move fast, spreading the work across many years.

So Terafab is a signal about expected demand, not a prediction about supply. Someone with unusually good visibility into his own future compute needs looked at the global supply curve and decided he’d rather attempt one of the hardest manufacturing problems in the world than wait in line. That tells me how large he expects the need to be. It doesn’t tell me the fab gets built, the economics work, or that today’s semiconductor companies lose anything. If Terafab never opens, the demand signal still stands.

Where this could go wrong

The failure mode I take most seriously isn’t that AI disappoints. It’s that this industry does what it has always done.

Semiconductors are cyclical in a way that has humbled better investors than me. Capacity gets built for peak demand, arrives well after the decision, and meets a market that has cooled. Memory in particular has a long history of violent price swings in both directions.

The numbers cut both ways, which is exactly why I trust them. SEMI projects global 300mm fab equipment spending of $133 billion in 2026, rising to $151 billion in 2027, $155 billion in 2028 and $172 billion in 2029, with installed capacity growing around 6 to 7% a year through that period. Read one way, that’s the industry putting extraordinary sums behind a belief in structural demand. Read the other way, it’s an extraordinary amount of new capacity arriving at once, and if demand disappoints, today’s shortage becomes tomorrow’s glut. SEMI’s own extended forecast has spending falling in 2030. Worth noting too that SEMI’s figure for 2026 was $116 billion when they published in October 2025 and $133 billion by April 2026, which tells you how quickly these expectations move and how much weight to put on any of them.

The fact that today’s constraint is real doesn’t tell me the capacity being built for it is correctly sized. If hyperscaler capital expenditure decelerates, or AI infrastructure spending fails to convert into returns for the companies doing the spending, the correction runs straight through this thesis.

The efficiency question also isn’t closed. Everything above says cheaper compute has expanded demand so far. It doesn’t prove it always will. If model efficiency improves faster than new use cases appear, aggregate demand growth could slow even while the technology keeps working. I don’t think that’s the likely path, but I hold it as a real possibility, and it’s the specific thing I’d want evidence on before adding.

Then the structural risks. Custom silicon could compress the economics of merchant accelerators. Compute could commoditise at the layer where the money currently sits. Power could physically cap deployment regardless of how many chips exist. And valuations across this sector have been priced for a great deal to go right, which means the industry can be correct while the shares still disappoint.

Geopolitics deserves a mention and not more. Leading-edge manufacturing is concentrated in Taiwan to a degree with no parallel in any other critical industry, and US-China export restrictions continue to reshape who can buy and build what. Diversification into the US, Japan and Europe is real but slow and expensive, and duplicating capacity raises costs. It’s a tail risk I can’t price. A reason for humility rather than a reason to avoid the sector.

How this actually sits in my portfolio

Semiconductors are not one of my largest exposures, and I’d rather say that than write a page implying more conviction than I hold.

My philosophy is to own pieces of the companies I think are building the future. AI software, physical AI, cloud, and the compute layer underneath all of it are different expressions of the same broad view. Semiconductors are one layer of that future rather than the whole of it, and I size them accordingly. Where I have higher conviction I’m willing to concentrate; here my conviction sits in the direction of compute demand rather than any single company’s durability, so the exposure is spread.

That’s the honest shape of it. I believe the world will need vastly more computation than it currently produces, I believe supplying it is genuinely difficult, and I don’t know which companies will still be capturing the economics in fifteen years. Owning the layer rather than the name is how I invest around that particular kind of uncertainty.