What separates Etched's inference racks from Nvidia GPUs
Etched builds its inference racks to handle only the final stage of running large AI models, while Nvidia GPUs were designed for a wider range of training and general computing tasks. The difference shows up in how Etched divides the work. It splits each request into a prefill phase that loads the context and a decode phase that generates each new token. This split lets the hardware run at lower voltage and share memory across chips without the overhead that general-purpose GPUs carry.
The shared memory layout removes repeated data movement between separate memory banks. Low-voltage chips reduce power draw while still keeping the decode step fast. Together these choices cut the time it takes to produce the first useful output and keep later tokens flowing at a steady rate. Jane Street has already placed the racks into live trading systems, where small gains in response speed matter.
Because the racks are fixed to this inference pattern, they avoid the extra circuitry and software layers that Nvidia GPUs need to support training jobs or other workloads. The result is hardware that uses fewer watts per generated token and requires less complex cooling in a rack. Companies that run high volumes of inference can therefore fit more capacity into the same data-center space and lower their power bills without changing their model code.
The core lesson is that hardware tuned to one narrow but common step in AI work can outperform general chips on that step alone.
How prefill and decode phases cut response times in practice
Etched builds rack systems that run AI inference, the step where a trained model turns a user prompt into an answer. Inference splits into two distinct phases that shape how fast responses appear. The prefill phase reads the full prompt in one pass and loads the necessary context into memory. The decode phase then produces each new token one after the other until the answer is complete.
Etched's chips and shared memory architecture keep data close to the compute units during both phases. This reduces the repeated memory transfers that occur on general purpose GPUs built for training as well as inference. Because the hardware stays focused on these two operations, the time between prompt submission and the first visible token drops, and the overall generation finishes sooner.
Teams running customer facing AI services see the difference in daily operations. Lower latency per request means they can handle more traffic on the same number of racks or cut the number of racks needed for a given load. Contracts already signed by Etched show customers value this efficiency at scale.
The practical result is that specialization for inference removes a common source of delay without changing the model itself.
Why Jane Street became both investor and first customer
Jane Street led the $700 million funding round that pushed Etched to a $21 billion valuation. The quantitative trading firm had already backed the chip startup earlier and chose to increase its commitment while also becoming the first customer for the specialized inference hardware.
Trading firms run large language models constantly to analyze market data and execute trades. Speed and cost matter directly because every extra millisecond or dollar per token affects profits in competitive markets. Etched designs chips that target exactly those constraints, promising higher tokens per dollar and per watt than general-purpose alternatives.
The decision to both invest and buy the systems early reflects a practical calculation. Jane Street can test the chips in live trading environments while holding equity in the company building them. This arrangement gives the firm influence over product direction and early access to hardware that could lower its own infrastructure expenses.
Etched already has a working chip and more than 400 employees. The combination of real silicon and clear demand from a sophisticated user helped drive the rapid valuation increase from the $10.3 billion level reached just weeks earlier.
The pattern shows how one large user with heavy inference needs can accelerate both product development and investor interest in a new chip design.
What low-voltage chips and shared memory actually change
Etched reached its new valuation after Jane Street took delivery of the first cluster and ran it on live trading workloads. The hardware uses lower voltage operation and a single memory pool visible to every processor, two choices that directly affect how much power an inference job consumes and how quickly data moves between chips.
Lower voltage reduces the energy cost per calculation, which matters when models run nonstop. Shared memory removes the need to copy large tensors back and forth across separate memory banks, cutting both latency and the number of chips required to keep utilization high. Jane Street's engineers measured these effects in practice and chose to lead the next funding round, moving the company's value to 21 billion dollars inside a month.
Teams that run repeated inference see the difference in daily operations. Power draw per token drops, so the same rack can handle more traffic without new electrical upgrades. Training jobs that once required extra machines to hide memory latency finish in fewer hours because data stays resident. Architecture planning changes as well: instead of sizing clusters around peak bandwidth limits, engineers can allocate capacity based on actual model size.
The pattern is straightforward. When a specialized design solves the exact constraints of one class of workloads, the first customer to prove it in production can shift market expectations faster than broad announcements ever do.
How Etched secured over a billion dollars in contracts already
Etched moved from a $10.3 billion valuation to $21 billion in under a month once Jane Street shifted from investor to customer and took delivery of its Sohu chips. The trading firm tested the transformer-only inference hardware and placed an order, turning a funding round into proof of real demand. That single transaction helped drive a $700 million raise at the new price, a figure that equals roughly half the market cap gains AMD posted from its AI chip business over the past year.
The move stands out because Etched has barely started shipping product. Most chip startups spend years proving they can build at scale before they land enterprise orders. Here, the hardware reached a production customer fast enough to double the company's worth in weeks. Investors appear to treat the order as evidence that specialized chips can handle inference workloads more efficiently than general-purpose designs.
The real test lies ahead. Etched must still show it can manufacture chips in volume, serve several large customers at once, and adapt as transformer models change. One firm running the chips in production marks a beginning. Multiple customers with live workloads would turn the startup into a credible alternative to broader platforms. Reaching that point would require steady output and ongoing engineering work, not just early wins.
The lesson is straightforward. A paying customer with hardware in hand carries more weight for valuation than promises alone, even when the company is still early in its production cycle.
Why the valuation moved from 5 billion to 21 billion so fast
Etched moved from a 5 billion dollar valuation in December 2025 to 21 billion dollars by August. The largest part of that increase happened in under a month. In July the company closed a 300 million dollar round at 10.3 billion dollars. Weeks later Jane Street led a new round that took the valuation to 21 billion dollars after committing 700 million dollars.
The trigger was not another set of slides or benchmark numbers. Jane Street shifted from investor to customer and took delivery of Etched's rack-scale inference system. That single change turned a paper promise into working hardware already running in a real environment. Most AI chip startups raise money on designs and test results. Few reach the point where a trading firm accepts physical racks and puts them into production use.
The difference matters for how capital flows next. When a buyer pays for delivered systems rather than future plans, later investors treat the risk profile as lower. Sequoia and Andreessen Horowitz had already backed the earlier round. Jane Street's decision to buy hardware on top of investing signaled that the product had cleared internal tests at a scale that matters for high-frequency workloads.
Three years earlier the company was still three founders working from a Thiel Fellowship bet that general-purpose GPUs would lose on cost and speed for large deployed models. The recent round shows that bet now carries a concrete order instead of a forecast. For other chip startups the lesson is direct: valuation jumps accelerate once the first customer accepts hardware rather than a demo.
What the 400-person team and San Jose base signal about scaling
Etched now employs more than 400 people and operates from a California facility while holding a working chip. That combination points to a shift from design sketches to actual production capacity. Most early chip startups stay small until they prove silicon in the lab. Etched moved past that stage quickly, which aligns with the jump from a $10.3 billion valuation in July to $21 billion after raising another $700 million.
A team of this size supports parallel work on tape-out, testing, and customer integration rather than sequential hand-offs. The San Jose location places engineers near foundry partners and supply-chain specialists, shortening feedback loops when issues arise during volume ramp. Investors appear to treat these operational details as proof that inference hardware can reach data centers at scale, where performance is judged by tokens per dollar and per watt.
The result affects how buyers plan deployments. Companies evaluating AI chips now see a vendor with enough staff to handle custom optimizations and enough physical presence to manage logistics. This reduces the risk that a promising design stays stuck in low-volume samples.
The clear lesson is that valuation multiples for hardware startups track team size and facility readiness as closely as benchmark numbers once customers begin taking delivery.
Which other investors joined Kleiner Perkins and Sequoia
Etched closed its latest round at a $21 billion valuation after raising $700 million. Jane Street led the round and continued its earlier support. Kleiner Perkins and Sequoia took part, and they were joined by Andreessen Horowitz and Tiger Global along with several other firms.
The list shows how quickly established names moved once the company proved it had working silicon and real revenue. Etched already holds more than $1 billion in signed contracts with public and private AI users plus cloud providers. Jane Street itself took delivery of the first rack last month and began running production workloads on it.
For a company that reached only $10.3 billion in July, the new backers supplied both capital and signals of demand. The round brought Etched's total funding to $1.9 billion and gave it a larger war chest while it scales from more than 400 employees. Investors appear to be pricing the bet on chips that can deliver higher tokens per dollar and per watt than general-purpose alternatives.
The practical result is that Etched now has both fresh money and a broader set of relationships with firms that operate large AI systems. That combination shortens the path from working prototype to volume deployment.
How inference-only hardware affects power and cost for teams
Teams that run large AI models spend heavily on electricity and hardware once training finishes. General chips must handle both training and inference, which means they carry extra circuits and memory systems that stay idle during response generation. Inference-only designs drop those unused parts.
Etched builds systems that focus solely on running trained models. Without training logic, the chips use fewer transistors and simpler memory layouts. This cuts the power draw per query and lowers the cooling load inside data centers. The result shows up directly in monthly bills, since each watt saved reduces both energy costs and the size of backup power systems needed.
For engineering teams, the shift changes project math. A cluster sized for inference can process more requests in the same rack space, or it can run on smaller power contracts. Deployment timelines shorten because teams no longer provision hardware that will sit underused. Budgets move from capital purchases of versatile but expensive chips toward operating expenses that scale with actual usage.
The practical lesson is straightforward. Hardware built only for inference removes overhead that training features create, so power consumption and total cost per generated response both drop. Teams gain headroom to serve more users without buying extra capacity or signing larger electricity agreements.
What this round reveals about demand for Nvidia alternatives
The jump to a $21 billion valuation in just weeks shows clear investor interest in hardware built for AI inference without relying on Nvidia. Etched packages its chips into full systems called frontier inference clusters, and the latest round drew backing from firms testing the hardware directly in their own data centers.
What stands out is how Etched separates the two main stages of inference. Its prefill chip runs at low voltage so more transistors fit on the die without extra heat. For the decode stage the company added cluster-scale memory and a custom interconnect. This setup lets many chips share one memory pool at low latency, which cuts the time needed to generate each new token. The result is faster output at lower power draw than standard high-end designs.
Investors see this as a practical path to lower costs when running large models at scale. Jane Street ran early tests and reported that the approach delivered the precision required for demanding workloads, enough to install its own rack. Etched also moved away from its original plan of hard-wiring chips to single models, so the systems now handle any frontier model.
The funding round therefore points to a growing market for inference hardware that trades general-purpose flexibility for speed and efficiency on specific tasks. Companies that need high token throughput at controlled cost now have a concrete alternative to evaluate.

