1. The Thesis: Inference Becomes the Main Event
As explored in The Custom AI Silicon Era - Jarsy Research #22, AI’s growing scale and cost are creating opportunities for specialized processors. This report begins a closer look at several startups pursuing that opportunity, starting with Etched and its concentrated bet on frontier inference.
Training develops a model; inference runs it whenever a user or another AI agent asks for an answer. As applications generate more tokens and agents make repeated model calls, inference is becoming an increasingly important cost. GPUs offer programmability and the mature CUDA ecosystem, but that flexibility carries overhead in silicon, power and data movement. Etched believes inference has become predictable enough to justify deeper specialization and deliver more useful tokens per dollar and watt. Investors have embraced that thesis: in July 2026, Sequoia led a $300 million Series C valuing Etched at $10.3 billion.
2. The Founders: A Dorm-Room Idea and a Second Chance
Gavin Uberti, now CEO, had already spent years close to the hardware-software boundary, working on AI compiler infrastructure and developing a Cortex-M backend for Apache TVM. Writing low-level code showed him how much of a general-purpose chip was devoted to moving data and preserving flexibility rather than performing the calculations AI models repeatedly needed.
At Harvard, Uberti found a close technical partner in Chris Zhu, a mathematics and computer-science student with experience in combinatorics and high-performance computing. While Uberti finished a gap year, Zhu remained on campus and the two developed their dorm-room project over Zoom calls every other day. Working together in person the following spring convinced them to leave Harvard and build Etched full-time.
The third founder brought a different kind of conviction. Robert Wachen (Uberti’s college roommate and now Etched’s president) was 16 when a bump on his back was diagnosed as stage-four bone cancer, with less than a 30% chance of survival. Surgery offered better odds but risked leaving him unable to walk; radiation preserved mobility but reduced his chance of surviving. He chose surgery, ultimately needed radiation as well, and spent nearly two years undergoing treatment and relearning to walk. Wachen has said the experience left him asking what he would do with his life if he were given it back.

Etched co-founders Robert Wachen, Gavin Uberti, and Chris Zhu. Image Credits: CNBC, Etched
That question became Etched's origin. Wachen had fallen for AI when GPT-3 first spoke back to him; when OpenAI shipped GPT-4V, the first model that could see an image, he pulled up a photo of his own back from before anyone knew what the bump was and asked the model to reason like a doctor. Within seconds, it suggested the bump could be a tumour and recommended an MRI, the conclusion his doctors had taken months to reach. Then he ran out of image credits. To Wachen, the moment showed both AI’s potential and the shortage of computing infrastructure needed to make it widely available. His conclusion matched Uberti’s: intelligence could become abundant, but the infrastructure to serve it did not yet exist.
What made Etched unusual was not simply the founders’ youth, but their lack of conventional semiconductor experience. While Cerebras, Groq and SambaNova were founded by industry veterans, Etched emerged from three Harvard students. They founded the company in 2022, left Harvard to pursue it full-time and became one of the few complete founding teams to receive Thiel Fellowships together in 2024.
Their timing was early and uncomfortable. In 2023, few investors believed transformers were stable enough to justify a highly specialized chip, much less that a group of college dropouts could build one. Nearly every major investor rejected the founders’ 30-page technical thesis, leaving Etched operating month to month and close to running out of cash.
Primary Venture Partners saw the opportunity differently. It led Etched’s $5.36 million seed round in March 2023, providing roughly three-quarters of the capital at a $34 million valuation. A year later, Primary and Positive Sum Ventures co-led a $120 million Series A to fund Sohu’s tape-out and manufacturing.
3. What Etched Is Building: From Chip to Inference Cluster
Etched began with a radical proposition: if transformers became the dominant AI architecture, a chip designed around their recurring operations could remove much of the flexibility and overhead of a GPU. Its original answer was Sohu, presented in 2024 as a transformer-only ASIC.
3.1 From chip to complete system
Etched’s 2026 product is broader than the original Sohu chip. Its frontier inference cluster combines specialized processors, memory, compute cards, interconnects, liquid cooling and software in a complete rack-scale system. Etched says it supports transformer families including Llama, Qwen and DeepSeek, as well as Mamba, a state-space model. The platform remains highly specialized, but is no longer easily described as transformer-only.

Table 1: Etched’s product has four connected layers. © 2026 Jarsy Research
A fast processor is useful only if it can access memory, communicate with neighbouring chips, remain cool and run production models reliably. This is why former Cypress CTO Mark Ross’s early advice “People buy systems, not chips” became central to Etched’s strategy.

Chip (left), Compute card (middle), Rack (right). Etched is building a rack-scale inference system rather than selling a standalone processor. Image Credits: Etched
3.2 Solving two inference bottlenecks
Inference has two main stages. Prefill processes the prompt and can perform many calculations in parallel. Decode generates the response token by token and is often limited by memory access and communication between chips.
Etched has designed one core technique for each bottleneck:

Table 2: Etched’s techniques for inference bottlenecks. © 2026 Jarsy Research
Low-Voltage Inference
The switching power of a CMOS* circuit is approximately:
Pdynamic=αCV2f
The important term is V2: power rises with the square of voltage. If voltage falls from 1.0 V to 0.5 V while everything else remains equal:
(0.5/1.0)2=0.25
In theory, switching power falls to one-quarter, creating room to operate more arithmetic units without exceeding the rack’s power and cooling limits. Etched says its mathematical units run at less than half the voltage of most AI accelerators. Achieving this requires the circuits, scheduling, power delivery, packaging and cooling system to be designed together.
The company claims this allows trillion-parameter sparse mixture-of-experts models to sustain more than 80% of peak FLOPs without thermal throttling. That remains an Etched-reported figure rather than an independently reproduced benchmark, although the underlying voltage-power relationship is established in semiconductor physics.
*CMOS stands for complementary metal-oxide-semiconductor. It uses two complementary types of microscopic transistors NMOS and PMOS as switches that turn electrical signals on and off, representing the binary values 1 and 0.
Cluster-Scale Memory
Decode often spends more time moving data than performing calculations. Its performance can be simplified as:
Performance ≤ min (Peak Compute, Memory Bandwidth × Arithmetic Intensity)
This is the logic of the Roofline Model: once memory bandwidth becomes the bottleneck, adding more computing power produces little benefit. For example, a 70-billion-parameter model stored at FP8 requires approximately:
70B parameters × 1 byte ≈ 70 GB
During decode, the system must repeatedly access these weights and the KV cache containing information from previous tokens. Longer contexts and larger batches increase the memory requirement further.
Etched’s CSM (Cluster-Scale Memory) combines HBM and SRAM with a proprietary high-bandwidth, low-latency interconnect. It is intended to let multiple chips access a shared memory pool, combining HBM’s greater capacity with some of SRAM’s responsiveness. Etched has not disclosed CSM’s bandwidth, latency, SRAM capacity or network topology. It should therefore be presented as the company’s proposed solution to the memory bottleneck, not yet as a proven one.

Inside an Etched rack. Image Credit: Etched
3.3 From working silicon to production
Etched’s A0 (the first silicon revision) returned from TSMC’s N4P process in early 2026, and the company reports achieving first-pass success. It is now validating rack-scale systems with customers and says it has more than $1 billion in customer contracts.
The company’s original Sohu materials projected more than 500,000 tokens per second on Llama-70B using an eight-chip server. That figure predates the working A0 silicon and has not been independently reproduced. Etched’s more recent quantitative claim is that its system can sustain over 80% of peak FLOPs on trillion-parameter sparse-MoE workloads, but the underlying customer-test results have not been published.
Etched has also built a 10-megawatt lab, datacentre, test house and prototyping facilities in the San Jose area, alongside a Taiwan manufacturing operation. It says production is underway and its first racks are scheduled to ship during summer 2026.
4. The Believers: Team, Culture, and Capital
Primary Venture Partners’ early conviction rested less on a finished product than on the founders’ ability to build one. The firm had independently been studying Nvidia’s dominance, the GPU bottleneck and the potential for hardware designed around large language models, but Etched’s lack of semiconductor experience remained a significant risk.
Primary asked industry experts to scrutinize Uberti’s white paper and question him directly. Several reportedly began as sceptics but came away impressed by his reasoning, adaptability and speed of execution. Primary therefore framed its first test simply: could these young founders recruit the experienced team required to turn an unconventional architecture into a production system?
4.1 From three founders to 400 engineers
A rack-scale accelerator is not a one-chip project. It requires expertise across silicon, firmware, packaging, compilers, cooling, manufacturing and complete datacentre systems. Etched says it has assembled more than 400 people, including engineers from Nvidia, Google’s TPU group, Broadcom, SK Hynix and TSMC. Its leadership combines the founders’ willingness to rethink the architecture with executives who have shipped complex hardware before:
Mark Ross, CTO: former CTO of Cypress Semiconductor, with experience delivering several successful first-pass silicon systems.
Brian Loiler, VP of Platform: spent 22 years at Nvidia and helped build its HGX and DGX platforms.
David Munday, VP of Software: built Google’s TPU software team and worked across TPU generations v1 through v5.
Wayne Cao, VP of Production: led production and supply-chain programmes for products including the original iPhone, MacBook Air, Pixel and Chromebook.
Saptadeep Pal, VP of ASIC and Architecture: previously worked on Nvidia’s H100, A100 and V100 architectures.
Recruiting this level of experience is an important execution signal. It does not prove that Etched can manufacture reliable systems at scale, but it reduces the risk that a technically ambitious design remains only a promising prototype.
4.2 An in-person, vertically integrated culture
Etched works primarily in person, allowing its silicon, packaging, software, cooling and platform teams to develop the system together. It has built datacentre and prototyping facilities in San Jose and opened a Taiwan facility to work continuously with manufacturing partners.
This expensive, unusually integrated approach reflects Mark Ross’s early advice: customers buy complete systems, not theoretical chip performance.
4.3 Capital follows the team
As Etched recruited experienced engineers and moved from a white paper to working silicon, its investor base expanded from one unusually convicted seed fund to major venture firms, strategic semiconductor investors and trading companies.

Table 3: Etched Funding Rounds. Sources: CNBC, Bloomberg, Reuters, Etched. © 2026 Jarsy Research
5. Competition and Outlook
Etched is entering a market in which specialization is no longer unusual. Nvidia and AMD are developing complete rack-scale platforms, while startups are redesigning processors, memory and interconnects around inference. For more information, refer to our earlier research report The Custom AI Silicon Era.

Table 4: Competitions for Etched. © 2026 Jarsy Research
These companies compete with Etched through different combinations of speed, efficiency, flexibility and system integration. Nvidia remains the benchmark because of its software ecosystem and installed base, while d-Matrix is the closest specialized startup comparison. SambaNova and Cerebras offer more established alternative architectures.vThe wider field includes MatX, Fractile and Positron, while Google, Amazon and Microsoft reduce the available market by developing custom processors for their own clouds.
Etched must now turn working silicon and customer commitments into reliable production deployments. Its opportunity is substantial, but success depends on delivering reproducible gains in performance, cost and power efficiency while keeping pace with changing model architectures. It does not need to replace GPUs everywhere, only prove that a large and valuable share of inference runs better on its systems.
Further Reading:




