An internal Meta memo, first reported by Reuters on July 9, says the company will begin manufacturing Iris in September.

Iris Enters Production in September, Built Only for Inference
The chip is designed in-house with help from Broadcom and manufactured by TSMC, and the memo notes that testing took only six weeks and turned up no major problems – fast progress for an effort that had struggled for years.
Iris is not a one-off. It sits inside Meta’s Training and Inference Accelerator program, or MTIA, and in March 2026 the company laid out four variants – the 300, 400, 450, and 500 – on a roughly six-month release cadence, about twice the pace of a normal chip roadmap. The Broadcom partnership runs through 2029, and later parts are expected to be among the first custom AI chips built on a 2-nanometer process.
The purpose is narrow on purpose. Meta says it will keep buying merchant GPUs for the heaviest model training, while Iris handles the daily inference work – the recommendation and ranking that decide what shows up in Facebook and Instagram feeds. That is the highest-volume, most repetitive compute Meta runs, and it is exactly the kind of workload a chip tuned for one job can do more cheaply than a general-purpose GPU.
Why Meta Is Targeting Inference First
Meta was often framed as the last large US hyperscaler to establish custom AI silicon, but that framing is now outdated: Meta says hundreds of thousands of MTIA chips already serve inference workloads, while Google and Amazon have older programs. The relevant difference is fleet scale and release cadence, not whether Meta has crossed a first-production line. The interesting part is not that Meta built a chip – it is which job the chip was built for.

Inference is the recurring workload: every feed refresh, ranked post and recommendation requires another execution. Running more of that volume on hardware tuned to Meta’s own software could lower cost per operation, while merchant GPUs remain available for frontier training and workloads that custom silicon cannot yet handle.
The strategic objective is therefore narrower than replacing Nvidia: reduce the cost of the workload Meta runs most often.
This is why “Meta is ditching Nvidia” misreads the memo. Meta is running two tracks at once: rent Nvidia for the frontier training it cannot yet do in-house, and run Iris for the inference it does constantly. The real bottleneck is not chip design, which Broadcom and TSMC can supply to anyone with a checkbook. It is power and leading-edge fab capacity – the same scarce foundry capacity that has companies racing to lock up TSMC lines.
A 14GW target is a statement about watts, and a custom chip is how you get more useful work out of each one.
The caveats are real, and worth stating plainly. In-house silicon at Meta has floundered before, so “enters production” is not the same as “runs the fleet.” The six-week test figure and the 14GW plan both come from an internal memo, not a shipped result. And custom silicon does not remove dependence so much as move it: Meta still needs TSMC’s 2nm capacity and Broadcom’s design help, and it still needs Nvidia for training. What changes is leverage. A credible in-house inference chip is the strongest card any large buyer can hold the next time it negotiates GPU supply and pricing.
Fourteen Gigawatts Sets the Scale
The 14-gigawatt plan sets the economic scale. Meta is treating power and useful work per watt as scarce resources, and custom inference silicon as one way to protect operating margin as model capabilities become less differentiated. Watch whether Iris actually reaches volume in the fall and whether it measurably bends Meta’s GPU spending; if it does, expect every hyperscaler still fully renting inference to feel pressure to answer.
Fleet Size and Roadmap Cadence Tell Different Stories
The framing of a final holdout crossing over does not survive a look at the deployment numbers. Custom accelerators across the large operators are put at roughly 1.9 million units for 2026 – around 900,000 Google TPUs, 600,000 AWS Trainium, 250,000 Microsoft Maia and 180,000 Meta MTIA. Meta sits fourth on that list. What it owns is not primacy but pace: three chip generations moving through 2026, with MTIA v3 entering production mid-year and a v4 sampling late, which is a more aggressive cadence than anyone else is attempting.

The share numbers underneath are what make this more than a procurement story. ASIC-based AI servers are projected at 27.8% of shipments for 2026, with custom silicon growing about 44.6% year over year against 16.1% for merchant GPUs – nearly triple the rate. One forecast has Nvidia’s share of inference specifically falling from the ninety-percent range toward 20 to 30% by 2028. Inference is the workload that runs every hour of every day, which is precisely why it is the one worth owning and the one a merchant supplier loses first.

Two details complicate the independence reading. Google has split its seventh-generation programme across suppliers – Broadcom on the training part, MediaTek on the inference-focused variant – so even the most mature custom effort is buying design capability rather than possessing it outright. And TSMC fabricates for all five. The hyperscalers are diversifying away from one vendor’s architecture into a single foundry’s capacity, which relocates the dependency without removing it.
Production Volume Is the Test
Volume in the autumn is the article’s stated test, and it is the right one. Three more would settle the rest.

- A disclosed cost-per-inference comparison. Meta has never published what a ranking query costs on GPU versus MTIA. Without it, “bends GPU spending” stays a directional claim that capital expenditure alone cannot confirm, since a flat GPU line during rising traffic means something different from a falling one.
- The v3-to-v4 cadence holds. An aggressive roadmap is a liability if a generation slips: silicon committed against a workload that has already moved on is a write-down. Two on-time generations would show the roadmap is real rather than announced.
- Someone rents theirs out. Maia serves Azure and OpenAI traffic but cannot be rented externally. The first hyperscaler to sell custom inference capacity to outside customers turns an internal cost programme into a competing product – and that is the move that would actually take share from a merchant supplier rather than merely declining to buy from one.
Whether Iris Replaces Nvidia
No. Iris is built for inference – the ranking and recommendation work behind Facebook and Instagram feeds – while Meta says it will keep buying Nvidia GPUs for frontier model training. The chip is meant to cut the cost of the workload Meta runs most often, not to end its Nvidia relationship.
Sources
- Meta, MTIA roadmap
- Meta and NVIDIA partnership
- cnbc.com — CNBC carrying Reuters’ exclusive report on the internal Meta memo – Iris production start, six-week test cycle, 14GW compute plan; direct fetch returned HTTP 403 on repeated attempts, date carried from the article’s own published dateline and this job’s original Sources annotation. (2026-07-09)
- presenc.ai — Presenc AI’s hyperscaler custom-silicon tracker, fetched and confirmed: ~1.9M total 2026 custom accelerators (~900k Google TPU, ~600k AWS Trainium, ~250k Microsoft Maia, ~180k Meta MTIA), and Microsoft Maia 2 serving OpenAI/Azure inference and internal workloads without third-party rental. Page states “last updated May 2026”; no more precise day is shown. (2026-05)
View all sources
- tomshardware.com — Tom’s Hardware analysis of Meta’s three MTIA generations moving through 2026 (v3 production mid-year, v4 sampling late), Google’s seventh-gen TPU split between Broadcom and MediaTek, and TSMC fabricating for all five hyperscalers; repeated direct fetches returned only site navigation, not the article body, so content and date are carried from this job’s original Sources annotation, not independently re-confirmed here. (2026-05)
- introl.com — Introl trade analysis, fetched and confirmed: custom ASICs growing at 44.6% CAGR against 16.1% for GPUs, and Nvidia’s inference-specific share projected (per New Street Research) to fall to 20-30% by 2028. The article’s byline date is 2026-02-23, earlier than this job’s original “(2026)” annotation; this fetch located the growth-rate and inference-share figures verbatim but did not independently re-locate the specific “27.8% of 2026 shipments” figure the draft cites – the page’s visible ASIC data is framed as revenue ($165B by 2033 vs $290B for GPUs) rather than a 2026 shipment-share percentage. (2026-02-23)
- money.usnews.com — U.S. News carrying the same Reuters exclusive on the internal memo (six-week test, doubling compute capacity); direct fetch timed out on repeated attempts, date carried from the article’s own URL dateline and this job’s original Sources annotation. (2026-07-09)
- techcrunch.com — TechCrunch, fetched and confirmed: September 2026 production start, Broadcom design partnership, TSMC manufacturing, Samsung/SanDisk/Sumitomo component agreements, 7GW-to-14GW capacity plan, and roughly $125-145B in related 2026 capex. (2026-07-09)
- mlq.ai — MLQ News coverage of the MTIA 300/400/450/500 roadmap, the Broadcom partnership through 2029, and 2nm-node parts; direct fetch returned HTTP 403 on repeated attempts, date carried from this job’s original Sources annotation. (2026-07-09)
- cryptobriefing.com — Crypto Briefing, fetched and confirmed: Iris built with Broadcom as design partner and TSMC as manufacturer, September 2026 production start, and framing that each successful MTIA generation reduces Meta’s marginal dependence on external GPU suppliers without eliminating Nvidia/AMD purchases outright. (2026-07-09)
- techerati.com — Techerati, fetched and confirmed: the 7GW-to-14GW-by-2027 target framed as requiring coordinated manufacturing/supply-chain planning, plus Samsung/SanDisk/Sumitomo Electric agreements; Iris positioned as complementing rather than replacing Nvidia/AMD GPUs. This fetch’s byline reads “Friday, July 10, 2026,” one day later than this job’s original “(2026-07-09)” annotation. (2026-07-10)
- finance.yahoo.com — Yahoo Finance carrying the Reuters exclusive, fetched and confirmed: 7GW deployed in 2026 (1GW in H1, targeting 2.5GW by year-end) doubling to 14GW in 2027, Iris as part of a four-generation MTIA project, six-week clean test cycle, Broadcom design plus TSMC production, and roughly $145B in 2026 AI infrastructure spend. This fetch did not explicitly surface an inference-vs-training workload split in the memo text it extracted, though the dependence-reduction framing is consistent with it. (2026-07-09)
This article is for informational and educational purposes only and does not constitute investment, financial, or legal advice.