# Frontier Chips for Frontier AI Labs, with Walter Goodwin, Founder/CEO of Fractile

No Priors: AI, Machine Learning, Tech, & Startups · 2026-10-02

<https://nopriors.podhood.com/3f46492d-281b-4b78-be96-6b4d5f791c45>

Fractile founder and CEO Walter Goodwin tells Sarah Guo that full-stack chip design and a memory bandwidth bet—not flops—will drive the next leap in inference. Keeping architecture through physical design in-house at Fractile's 150-person team avoids Broadcom handoffs that leave hyperscaler ASICs like NVIDIA and AMD GPUs. Fractile pivoted from SRAM to a chip ramping in the second half of next year pairing DRAM's cost with Groq-like speed, aimed at long-running agents, not chatbots. Compressing the design cycle means more shots on goal—AI shrinks front-end work despite 3–5 month fab cycles—and a 3–6 month structural edge wins deployments. He sees bandwidth scaling laws enabling sparser MoEs and argues labs can't all-in on proprietary silicon: a rival breakthrough could kill them in nine months.

## Questions this episode answers

### What does Fractile do and why did Walter Goodwin start the company?

Walter Goodwin says Fractile builds very fast inference chips for the world's largest models. Started in summer of 2022, the bet was driven by foundation models generalizing across everything and wise people arguing compute should be poured into models at test time, like AlphaGo's rollout becoming superhuman, applied to language modeling.

[1:24](https://nopriors.podhood.com/3f46492d-281b-4b78-be96-6b4d5f791c45?t=84000)

### What are Fractile's most important technical bets?

Walter Goodwin says the first bet was on inference and speed. Fractile initially built an SRAM-based chip like Grok or Cerebras, but worried about scalability as context lengths grew. For two years they've run skunkworks projects with memory vendors to get extremely high bandwidth to DRAM, uniting DRAM's capacity and cost with SRAM-class speed, ramping in the second half of next year.

[9:54](https://nopriors.podhood.com/3f46492d-281b-4b78-be96-6b4d5f791c45?t=594000)

### Can AI compress the chip design cycle from architect intent to GDS2?

Sarah Guo cites a top-three semiconductor CEO saying architect intent to a fully usable GDS2 file is ten years away. Walter Goodwin says divide that by four and subtract a bit, expecting end-to-end prototyping within a few years, though place-and-route tools, fab cycle times of three to five months, and three to five-year amortization windows remain bottlenecks.

[15:56](https://nopriors.podhood.com/3f46492d-281b-4b78-be96-6b4d5f791c45?t=956000)

### Why won't frontier labs just build all their own chips internally?

Walter Goodwin argues labs play an asymmetric game: if one lab goes all in on proprietary silicon and a rival discovers a computational breakthrough that only works on a different chip, it could die in the nine months before deploying enough of its own chip. So frontier players rationally deploy the same platforms and sustain third-party chip vendors.

[31:49](https://nopriors.podhood.com/3f46492d-281b-4b78-be96-6b4d5f791c45?t=1909000)

## Key moments

- **[0:00] Intro**
- **[1:20] Fractile's mission**
  - [1:24] Fractile CEO Walter Goodwin: the company builds ultra-fast inference chips for the world's largest AI models
- **[2:30] Chip landscape**
  - [2:30] Walter Goodwin founded Fractile in 2022 betting test-time compute would make inference speed the key AI bottleneck
  - [2:41] Fractile's mission: run trillion-parameter models thousands of tokens per second at extreme context lengths
  - [4:13] Hyperscaler AI chips from Google, Meta, Microsoft and OpenAI are surprisingly similar, says Walter Goodwin
- **[4:56] ASIC handoffs**
  - [4:56] Why so few chip teams go full-stack from architecture down to physical design at the foundry
  - [5:07] How a chip design goes from architect's intent to a GDS2 bitmap shipped to TSMC
  - [5:13] Broadcom builds most hyperscaler AI chips — Google TPU, Meta MTIA and Microsoft Maia all share the same HBM and TSMC packaging
- **[7:19] Full-stack team**
  - [7:37] Fractile runs a 150-person full-stack chip team to close the loop between workloads and silicon faster than rivals
- **[9:49] Technical bets**
  - [9:49] Walter Goodwin: the inference chip that recently went public was a training chip until 18 months ago
  - [10:46] Fractile abandoned its SRAM-based chip design over worries about scaling to long context lengths
  - [11:57] Fractile's new platform pairs DRAM's low cost with Grok-and-Cerebras-level speed, ramping in the second half of next year
  - [13:10] Walter Goodwin: the snappier chatbot is the faster horse — real fast-inference value is radically speeding up long-running agents
  - [13:44] Fast inference chips today can't run long context attention — customers still flip back to a GPU for that
  - [14:37] At datacenter scale, inference economics collapse to cost per gigabyte of memory, says Walter Goodwin
- **[15:05] Design cycle speed**
  - [15:56] Why every new LLM cries out for orders of magnitude more memory bandwidth, per Walter Goodwin
  - [18:05] Amdahl's law in chip design: AI can't shorten front-end timelines if downstream bottlenecks stay slow
  - [19:03] Fab cycle times are 3 to 5 months even in a super-hot lot — the innate latency of the chip industry
  - [19:49] Walter Goodwin: chips still need a 3-to-5-year amortization window, so new architectures must have a 3-year-plus lifespan
  - [20:55] Walter Goodwin: compressing the chip design cycle means more shots on goal, not shipping new chips every few weeks
  - [22:10] Walter Goodwin: a structural 3-to-6-month chip advantage wins all frontier deployments, like frontier labs' model lead
- **[23:03] Architect to GDS2**
  - [23:03] A top semiconductor CEO privately predicted 10 years until architect intent maps straight to a usable GDS2 file
  - [23:48] Walter Goodwin's heuristic: take a skeptical AI timeline and divide it by four
  - [24:28] Chip design's place-and-route tools still run for days using conventional algorithms, says Walter Goodwin
  - [25:30] AI-driven chip design hits the same wall as recursive self-improvement: the experiment becomes the bottleneck
  - [26:30] Surrogate models for finite element and thermal analysis could unlock huge returns as AI speeds up design iteration
  - [27:30] Cadence and Synopsys stay essential for final DRC and LVS sign-off even as AI transforms chip design
- **[28:17] Workload bets**
  - [28:29] Walter Goodwin: as experiments get expensive, it becomes a moral compunction to think harder before each one
  - [29:37] Barron Miller and Dwarkesh: think a hundred years of AI-equivalent thought before and after each costly experiment
  - [30:30] Fractile claims 25x more bandwidth per chip than HBM-based GPUs — and wants to pull the model landscape toward it
- **[31:17] Market structure**
  - [31:38] Sparser mixture-of-experts models save flops but hit bandwidth walls on today's HBM GPUs, says Walter Goodwin
  - [32:47] Compute flops scaled a million-fold in 20 years but memory bandwidth only 40x — Fractile bets on closing that gap
  - [33:52] Hyperscalers' internal chips exist mainly to negotiate NVIDIA prices down, says Walter Goodwin
  - [34:50] Walter Goodwin: frontier labs can't go all-in on proprietary silicon — a rival's breakthrough could kill them in nine months
- **[35:15] Outro**

## Speakers

- **Sarah Guo** (host)
- **Walter Goodwin** (guest)

## Topics

Semiconductors & Chip Design

## Mentioned

AMD (company), Broadcom (company), Cadence (company), Cerebras (company), Fractile (company), Nvidia (company), Synopsys (company), TSMC (company), TPU (product)

## Transcript

### Intro

**Walter Goodwin** [0:00]
Right now we're trying to build a single chip. If you crack open an NVIDIA system, it has anywhere between, kind of, 6 and 9 custom chips, all built by NVIDIA, to come together to build something that is really, really potent.

So there's already a gap there. You know, it'd be great to be productive enough that we could start to build our own responses in that same kind of way. This idea of being able to have, at any given moment, a kind of rolling frontier of bets that you hope are deeply aligned, in which you're ready to trigger the ramp off.

This is a huge advantage if you can kind of build that machine and build that engine. It's the equivalent of the frontier model for the chip space, is if you can just find a way to structurally carve out a 3 to 6 months advantage, you will be winning all of those deployments.

**Sarah Guo** [0:47]
Hi listeners, welcome back to No Priors. Today I'm here with Walter Goodwin, founder and CEO of full-stack AI chip company Fractile. We talk about what it means to be full-stack as a company, the technical bets they're making, why he's so focused on memory bandwidth, his predictions for frontier model architectures in the future, and the structure of the chip market at the frontier today and tomorrow, amongst NVIDIA, AMD, internal efforts, and this new class of accelerators.

Welcome, Walter. Walter, thanks so much for doing this.

**Walter Goodwin** [1:18]
Very good to be here, Sarah.

**Sarah Guo** [1:20]
You started this incredibly interesting company called Fractile. Can you give us an overview of what you guys do?

### Fractile's mission

**Walter Goodwin** [1:24]
Fractile is a chip company. We build very, very fast inference chips for the world's largest models. This sort of bet on speed is something which we've hadright from the outset. We started the company in summer of 2022, and I think at the time we were starting to see two things.

You know, one was obviously the arrival of foundation models trained on the internet and generalizing across everything. And the other was, you know, a few kind of wise people saying, you know, we need to find a way to pour more compute into these models at test time.

We need to find a way to take what we had for kind of AlphaGo, where you have a capable neural network, but then it becomes superhuman as you roll it out at scale, and bring that to kind of language modeling, for example.

And so our big pursuit over the past four years has been to find a way to build chips that will allow for kind of simultaneously taking these huge models and running them much, much, much faster than the chips of today, but do that in a way that kind of scales.

So scales to models beyond the frontier today, scales to extraordinarily long contexts. And so that's the mission that we've been on.

**Sarah Guo** [2:30]
It's become a much more interesting chip landscape in the four years since you started the company. Where would you place yourself in the overall GPU and accelerator landscape for inference?

### Chip landscape

**Walter Goodwin** [2:41]
Yeah, I'd say there's a kind of extraordinary zoo now of options available. And if you look across this kind of entire space of kind of AI ASICs, one of the things that I think is very striking is there's a lot of relatively identicate chips out there.

And this is kind of a structural, it's a kind of structural industry property where if you look especially across kind of hyperscaler efforts, and, you know, Google kind of kicked this off with the TPU more than 10 years ago, you see this paradigm where although there are a number of kind of proprietary internal chips, you know, Google with the TPU, Meta with MTIA, Microsoft with Maya, OpenAI now with Jalapeño, these are chips that are ultimately developed and delivered in partnership with kind of a relatively small number of what we might call ASIC houses, these kind of back-end delivery houses.

Broadcom is the largest, you know, $2 trillion company that help people kind of realize these designs. And so when you kind of zoom into this, this sort of seeming embarrassment of riches in the spectrum of what exists today, you see that actually there's a lot of similarity across these platforms.

You know, all of these chips, you have HBM, this type of DRAM memory that is, you know, common to NVIDIA GPUs, AMD GPUs, and all of these other ASICs. You have the same kinds of bets on Tensor cores to do your matrix multiplications, the same advanced packaging with TSMC.

And so I think one of the things that we see in the landscape is there's a relative dearth still of kind of efforts that go all the way up and down the silicon stack and try and build kind of fundamentally new capabilities.

And this is kind of a structural thing. There's just not that many teams that have decided, we're going to build from, you know, what you might call the kind of, you know, the architecture layer, the front-end design layer, which mostly looks like kind of writing code, but also all the way down to what we call kind of physical design, the process technology, the kind of foundry interactions, the stuff that essentially has you flying back and forth to Taiwan or Korea every week.

That's something which actually tends to still be left to a relatively small number of companies.

### ASIC handoffs

**Sarah Guo** [4:56]
For people who don't work in the chip industry, what is the common abstraction or handoff between what their, what, you know, architecture-focused players are doing and a player like Broadcom?

**Walter Goodwin** [5:07]
Yeah, so I think the, if you're, let's say you're Google and you're designing the TPU, you have a number of people in your team who have a really strong understanding of the workloads that they're trying to accelerate. And so, you know, going back over 10 years, this is a world that then sees you conceive of things like the Tensor core, which is a designated circuit dreamt up by an architect that is exceptionally good at matrix multiplications because you have this insight that matrix multiplications are the overwhelming majority of the kind of flop count in these models that you're running.

That's kind of the preserve of the architect, I guess, is somebody who really understands the way that chips behave, but also ideally has a very strong understanding of the kind of target workloads. What you'll then see inside those same organizations is a set of sort of, you know, very smart people that kind of lower that down into, I guess, you know, essentially a kind of circuit-level description of how this chip should behave.

And so a lot of this is what we call front-end design. You know, mechanistically, it still looks like, like writing code on a computer, and it's still code that is then handed off to a player like Broadcom. And so, you know, that kind of front-end design that describes the kind of underlying intent of the chip, the underlying logic of the chip, is then transformed ultimately into something that, you know, what gets shipped to TSMC in the end is literally a kind of bitmap.

You know, we have this file type you call GDS2, and it's literally, you know, where do I put the methyl layers? Where do I put each individual transistor? And so it's a, it's a full layout. And so you go through this process of synthesis of that RTL into a kind of set of circuits that then ultimately get laid out.

And a lot of the complexity there is around, you know, in Broadcom's case, you know, it will be ownership of analog IP that is used for chip-to-chip connectivity. And it is this kind of physical placement that is specific to a given process node at TSMC.

So, you know, we talk about like 3 nanometers, 5 nanometers, et cetera. So that piece is still today, you know, generally it is a kind of outsourced activity for all of these projects.

**Sarah Guo** [7:19]
You are describing what feels like a very large project to do end-to-end new chip creation. From a workload perspective, you're working with several frontier players to sort of characterize that and make sure that you can serve them. How did you think about, like, what does the Fractile team look like today?

### Full-stack team

**Sarah Guo** [7:35]
How can you do that as a startup?

**Walter Goodwin** [7:37]
Yeah, I think there is a, you know, it's funny, there's a lot of industries where there will be a kind of waterfall-style approach to doing things, and then suddenly, you know, along comes a new way of thinking, and it's kind of more agile to borrow from like management philosophy,right?

And I think for us as a very full-stack company, so we have a team that does span, you know, very deep workload understanding. I think we actually try to, you know, run ahead in many places and look at, you know, where would we change a model architecture to be more exceptionally aligned with our bet?

What do we think is like kind of the scaling law vector for the particular bets that we are taking? And then we can advocate to our customers and our partners on that. And it also helps us inform a set of bets for a chip that, you know, isn't going to necessarily exist in volume for, you know, one or two years.

That becomes a really important sort of part of the muscle that we have to build. But then as we look across this kind of, you know, ability to spin a very kind of agile loop, it means that inside Fractile, we do have front-end designers, we have our own physical design team, we have our own kind of back-end implementation team.

You know, we do our own work on advanced packaging, for example. And that's not an enormous headcount,right? I mean, Fractile today is about 150 people. We're sort of skinny in every single one of those sectors. But what it allows us to do is have this kind of much more agile closed loop.

And I think that's becoming increasingly essential as you look at the sort of cadence needs that are set by this industry. You know, the chasing, you know, perennially of the sort of chasing of the tail of the workloads is, you know, you have to, in the AI chip space, more than any other kind of chip play before, you have to really place your bets correctly.

So you have to have a lot of, you know, skill and a lot of luck. And you also then need to strike very fast. And so this kind of structural setup where we have the entire story in-house, it's very, very different to this kind of, I guess, prior approach where there's almost like a handoff point.

And so you get to a certain level, and then you hand off to another partner. And at some, to some extent, you're kind of at the mercy of how that other partner then behaves.

### Technical bets

**Sarah Guo** [9:49]
What do you think are the single most important technical bets that the company has made?

**Walter Goodwin** [9:54]
So for us, I think, you know, one is the direction. You know, four and a half years ago, you were obviouslyright in the heart of an inference story, but I still felt that we were educating the world about what inference meant.

For two years or so following that, some of the inference chips today, including one that recently went public, was a training chip until about 18 months ago. You know, there was a real avoidance of this idea that this sort of small, you know, final marginal cost, you know, marginal cost, I think, has two senses,right?

It either sounds like it's a small cost. What it in fact means is this is the cost that you pay every single time you deploy these models. So I think firstly, just the sheer bet that we were going to slip into a kind of deployment error and the bet on speed, that was something that kind of really informed a lot of what we then did as a kind of architectural response.

On the architectural side, you know, we've been through a journey, I think, for the first kind of two years or so of the company's life. We were working on a little like Grok or Cerebras, working on an SRAM-based chip.

So we'd sort of made the observation that SRAM is this super high bandwidth memory. It lives on the same piece of silicon as your logic. And so you have this extremely high bandwidth between your compute, where you're running the maths as you run these models, and let's say the weights of your model or the KV cache of this model as you're rolling it out.

And that's something which then can drive you to thousands of tokens per second on these language models. But I think one of the things that we started to worry about in kind of towards the end of 2023 and certainly in 2024 was the scalability of this approach.

And, you know, I think there are two things that grow with AI today. One is obviously the parameters of the model, but the other, and this was the one that kind of really got us nervous about that architectural approach, was, you know, this kind of growing context length that was becoming more and more part of the story for how we saw these models rolling out.

And it's that that has taken us over the past couple of years to a kind of what we see as a more exciting bet, I guess, in some ways, which is working more closely with memory vendors as well as our sort of logic foundry partners to find ways to get extremely high bandwidth access to higher capacity memories.

So for the past couple of years, we've been involved in these kind of skunkworks projects to move away from SRAM, you know, looking at how we can get, for instance, much, much, much higher bandwidth to DRAM memories. And this has been a very exciting bet for us because it's allowed us now to put together a platform which will be ramping in the second half of next year, which kind of unites, I guess, the scalability of these kind of higher capacity, lower cost DRAM memories that you have on a GPU, on a TPU, with all of the speed advantages that you get from a Grok chip or a Cerebras chip.

And I think where we see this as especially vital is, you know, if you look at where it is that speed really moves the dial today, there is a kind of, there's a version of this which is like a snappier chatbot.

But I think this is like, you know, when

Henry Ford sort of asked the rhetorical question of what people would have said they wanted and, you know, faster horses rather than a car. The snappier chatbot is kind of the faster horses of kind of fast inference. The place where the ability to take a multi-trillion parameter model and run it comfortably at many thousands of tokens per second, in our view, is a kind of fundamental elevator on capability for AI, is in taking these kind of very long-running agents and making them radically faster.

And so that's kind of a, it's a tantalizing mismatch today between the properties of the fast inference chips that we have, which have super high bandwidth memory, but incredibly low capacity. And if you look at the sort of technical detail of how they actually get employed and deployed today, they're not running long context attention.

You still flip back to a GPU to do that. And so there's this sort of tantalizing mismatch where precisely the area where we are almost able to go and take these things and make them much faster, we also don't have that technical capability.

And so in a sense, this is, you know, like many technical challenges, it boils down to a slightly mundane technical observation, which is we need chips that have this kind of particular ineffable property, which is incredibly high bandwidth to memory, so we can load the weights, we can load this state, you know, many thousands of times per second, but also have like economical memory.

Because when you look at what it looks like to run inference at data center scale for thousands of users, the economics actually collapse down to essentially a kind of cost per gigabyte of the memory that you're employing. And so that's been the sort of other key focus for us is unlocking this kind of fundamentally new building block, which is finding a path to get aggressively high bandwidth from this kind of, you know, from the world's lower cost memory, DRAM memory.

**Sarah Guo** [15:05]
You've said a few things that I think are not the conventional wisdom. One, in terms of making a bet yourselves on where the workloads are going. And then two, the idea that you can, you know, traditional chip people might think of delivery of chips in a generation, you know, every year or a little faster than that.

### Design cycle speed

**Sarah Guo** [15:26]
And, you know, big architectural changes come slower than that sort of tick rate. But you have aspirationally said within Fractile that you think the speed of like change in architecture and change in, you know, the technical plays you're making can be much faster than that.

Can you talk a little bit about like what would enable that to happen? Because the, you know, I think people have a better understanding of the physical limits of the supply chain today than they used to. So what is the flexibility that the industry could have?

**Walter Goodwin** [15:56]
Yeah. So I think there's a nuance on this, but what you want at any given moment in time for an AI chip is you wish that you had a chip that was precisely targeting the workload that you care about and in your hand today in volume.

But if you had that chip, you know, you're going to want something else again in six months' time. Like these workloads move on very fast.

**Sarah Guo** [16:19]
Well, today, you know, we have a model coming out about every two weeks. So it's a little tight.

**Walter Goodwin** [16:24]
That'sright. And so sanity comes from looking across those models and saying, well, okay, what's common to all of these? And like, thankfully, there's enough that is common. So, you know, it's always the case that a new LLM desperately wants to have orders of magnitude more bandwidth to memory in order to run faster.

It remains the case today that, you know, these LLMs tend to be auto-aggressive, very low batch as you generate text with this sort of fundamental trade-off again between kind of throughput efficiency and cost and the speed that you can serve these models at.

So there are certain things that these models are unified in kind of crying out for, and memory bandwidth is one of them. But then, you know, as you say, there are these evolutions. And especially if you look at the kind of frontier of open source Chinese models, you know, the exact nature of an attention mechanism will churn every couple of weeks in terms of what's kind of the frontier.

**Sarah Guo** [17:13]
Or the sparsity.

**Walter Goodwin** [17:14]
The sparsity, the sparsity level of the MOEs that you have, the sparsity on the attention itself. And, you know, we may be in for like bigger shocks as well. There may be kind of more fundamental shifts. And so I think there's this duality around sort of the limits of the physical and the kind of financial world, I guess.

And then these demands of, you know, workloads that churn at the pace of essentially software, where although there is this sort of clear desire to be able to kind of ship new platforms, you know, aggressively faster, I think there are also some sort of fundamental breaks on this.

So at Fractile, you know, we're extremely excited about being AI forward in how we do chip design. I think we've seen this ability to own the entire end-to-end gamut of the problem as really enabling a kind of fundamental rethink of how we do things.

You know, if you think about classic CS law or Zamdahl's law,right? Like everything I can paralyze becomes very, very fast, and the part that I can't paralyze doesn't. And I think it's similar when you're in a like long chain of organizations that, you know, there isn't necessarily an overwhelming need to or return on radically changing your processes as a front-end focused chip design startup, adopting AI in aggressive ways to shorten your timelines from being, you know, maybe 12 months of front-end design work to a tiny fraction of that.

If there's then going to be kind of the conventional bottlenecks and the conventional pace on the next piece. And I think that's sort of, you know, at some extent, there will always be those bottlenecks. You know, we all have the same kind of fab cycle times with our foundry partners.

So from the moment you send the chip to them to getting it back is, you know, three to five months, even in a kind of super hot lot scenario. And so these are the kind of innate latencies, I guess, in the industry.

And then when you get that chip back, I think it is vital to look at the way that these chips actually become financializable, which is that they do have a kind of payoff period. You have to have really a kind of three to five-year amortization window for that chip to make it a sensible financial decision.

And so there's these sort of two things that are intention and which I guess I believe two slightly conflicting things at the same time. One is there is a huge value in being able to take the overall chip design cycle and compress that and compress that.

There is also, I think, what we don't lose is the need to place really thoughtful and subtle bets architecturally because the chip you actually end up making does still need to have a three-year plus, you know, useful lifespan.

And so I think for me, you know, what it means to be able to over time compress that kind of chip design cycle is essentially that you get to have more shots on goal. You want to be forever ready to take a given flagship platform and say,right, that's it.

That is actually the flagship and this is the one we're going to ramp. But then it's a, you know, 12-month, 18-month volume ramp and you're expecting it to have longevity. You're expecting it to actually continue to drive value to your customers.

And so I think that the sort of subtlety here is, you know, I'm not a believer in the idea that we will ultimately get to a place where we are shipping a fundamentally new chip, you know, every few weeks just because we've shortened that down.

Because I think this is a physical world. There is a certain amount of constraint on, you know, data center power. There are latencies to installing these things and you need to finance the actual underlying silicons. It needs to have that payoff period.

But I think what you can get to is a world where the shorter you can make that latency, that gap between an observation and realizing that bet in volume, that is where there's an enormous amount of value to be captured.

And I think it's also an area where, you know, Fractile is a...

**Sarah Guo** [21:14]
It's only because the decisions that you make about what to ramp are the correct ones.

**Walter Goodwin** [21:19]
That'sright. I think, you know, you want to be, and maybe you get to have more irons in the fire as well. So, you know, one of the things that I think is so exciting about AI capability rising is as we do become more productive in doing things, you know, there's obviously these kind of two responses, you know, generally across the economy.

One is, oh no, we're not going to have as much to do. Maybe there's jobs lost. The other is we get to do more things. And I think, you know, standing here from Fractile, you know,right now we're trying to build a single chip.

We know that we're up against competitors that have, you know, if you crack open an NVIDIA system, it has anywhere between kind of six and nine custom chips all built by NVIDIA to come together to build something that is really, really potent.

So there's already a gap there. You know, it'd be great to be productive enough that we could start to build our own responses in that same kind of way. And so I think this idea of being able to have at any given moment a kind of rolling frontier of bets that you hope are deeply aligned and which you're ready to trigger the ramp of, this is a huge advantage if you can kind of build that machine and build that engine.

And the six-month gap that that might then perennially give you against your competition, that is the kind of, you know, wedge that allows you to drive in. And we see this in frontier labs,right? It's the equivalent of the frontier model for the chip space is if you can just find a way to structurally carve out a three to six months advantage, you will be winning all of those deployments.

**Sarah Guo** [22:50]
There's a meme in here somewhere of CEO saying, wow, this AI stuff is great and we structured the organization in a way that can consume it. Instead of one program, we're going to have six guys and everybody cheers.

**Walter Goodwin** [23:01]
Yeah. I think there's some truth in that.

**Sarah Guo** [23:03]
What I was talking to one of the top three semi CEOs and he wouldn't let me put this prediction, you know, actually on air, but I asked him like how soon until we can have essentially like intent from chief architect into fully usable GDS2 file and he's like 10 years.

### Architect to GDS2

**Walter Goodwin** [23:21]
Yeah.

**Sarah Guo** [23:22]
That's not going to happen like now. What is your view on that? Or where do you see the impact first in your organization?

**Walter Goodwin** [23:28]
You know, a heuristic that has been successful over the past couple of years is to just always question your sort of logical assumption and then divide it by four on timescale. So there is a world in which 10 years seems like a very sensible timeframe for that, but I would divide it by four and I probably subtract a bit from that.

You know, I think that we will have a space where there is a kind of prototyping that goes end to end, you know, in the next few years. I think one of the things about chip design that is very interesting is there are still these loops in the middle of it that are kind of conventional solutions to MP hard problems.

And so you have these sort of, you know, if you really want to output that GDS2 file, today you are still running through a bunch of layout tools that are doing place and route with conventional kind of algorithmic approaches, which will run for days on end.

You know, I think this is a really interesting thing that we see and actually to bring it back somewhat to the sort of workloads that we're excited about as well. A lot of hard problems in the world, I think, look a bit like this where there is a lot of thinking that can now be automated.

There's a lot of cerebral work. But then there is also, you know, this kind of intrinsic latency to something that you do. So I think we see this in, you know, AI for guiding AI model development today. So the RSI work, you know, yes, we can now have extreme intelligence that is guiding our experiments, but now we're just experiment bottlenecked.

You know, we're compute time bottlenecked. And so I think with chip design, it's a little bit like this. Maybe to take the kind of, you know, architect's intent through to GDS2 question, it's a little bit like the RSI question in that the things that will prevent us getting there extremely fast is the fact that actually in that flow, there are a number of things that are genuinely kind of computationally expensive in a more conventional way.

I think what we will generally see for all of these sorts of problems and.

**Sarah Guo** [25:34]
And they won't be replaced by surrogate models?

**Walter Goodwin** [25:36]
Well, I think, you know, there's definitely these sort of simulation models. I think driving approximations, you know, all of the models that we see today for kind of finite element analysis and sort of thermal estimation and so on.

I think there is a huge return to that kind of work precisely because of this kind of Amdahl's law that as we accelerate the intelligence, as we accelerate the period between these experiments, between these trials.

**Sarah Guo** [26:00]
They become the bottleneck.

**Walter Goodwin** [26:01]
It becomes much more important to accelerate those trials as well. And so, you know, we are looking internally at, you know, what is actually what are in the guts of some of those algorithms and is there a crude approximate way that you could do some of this?

You know, I think chip design is not going to change in the final sign-off for quite some time. It's incredibly valuable to have Cadence and Synopsys who've been working with TSMC and the other foundries for decades to build the sort of, you know, the final check mark that says this thing is what we call DRC and LVS clean.

It conforms to the rules that that foundry has set. That is actually where, you know, that sort of work is immensely valuable. But could you have some kind of fuzzy placement algorithm in the meantime that gets you almost all the way to your final decision?

**Sarah Guo** [26:49]
It helps you iterate faster.

**Walter Goodwin** [26:50]
Helps you iterate faster. You know, I think that is exactly where people should be doing more work because otherwise that is the bottleneck. The other part there, and I think this is true of, you know, RSI for AI experiments as well, is as we do become, you know, as our intelligence levels on the thinking that then drives that experiment gets higher and higher.

And as that experiment therefore becomes the bottleneck, I think we do more thinking. And this is just generally Fractile's workload bets. It's the reason we're so excited about hard problems and it's kind of the reason that we see Fractile as the chip that takes us toward accelerating the solution to hard problems in general is I think it becomes absolutely, it's like almost, you know, a moral compunction to think harder before every single experiment that you fire off.

I think it was maybe Baron Miller and John Dwarkesh recently who said, you know, that perhaps you should think for a hundred years, you know, before you fire off an experiment and then think for another hundred years of human equivalent with these models about the results of that AI experiment because the experiment in the middle is pretty expensive and takes a certain amount of wall clock time.

And so I think you get there also with things like chip design where there are these kind of fundamental wall clock time bottlenecks for the more conventional algorithms. We're going to be generating a lot of reasoning tokens before we go and place a given circuit.

**Sarah Guo** [28:17]
You are making workload bets and then collaborating closely with design partners on the model architectural, you know, shifts hopefully a little ahead so you can actually do some planning. What are you seeing?

### Workload bets

**Walter Goodwin** [28:31]
The biggest thing that we're doing, of course, is this kind of speed chasing,right? Maximizing memory bandwidth in order to be able to run these models faster. And one of the dualities I think you have to span as a chip company is to build something that is both exceptionally good at the kind of gamut of workloads that are coming down the pipe under this idea of the kind of hardware lottery where these will be trained on HBM-based GPUs or XPUs.

And so it's great that this is then a chip that's exceptional at taking kind of today'sultra-aggressive MOE transformers and running them much faster and more efficiently. But I think what's exciting about building chips with kind of fundamentally new capabilities is you then also get to explore whether there are kind of gravitational pulls that you yourself get to pull the new model landscape in because of those properties.

And so for Fractile, you know, this idea of having like 25 times more bandwidth per chip than an HBM-based chip, we get to explore, for instance, internally these kind of ideas of like scaling laws for bandwidth basically. You know, the scaling laws traditionally we think of them as like flop scaling laws.

You know, we get better and better performance the more flops we pour into these models at train time and test time. And if you look at just like the known landscape of ideas today, we already see some scaling laws for bandwidth.

So one is these mixture of expert models. It's pretty well known that actually ideally we would make them sparser and sparser and sparser. So for like ISO intelligence, you will save a ton of flops if you go from being like one in 16 sparse on your MOE to one in 128, one in 256.

But one of the challenges, one of the headwinds to doing that is actually that it becomes incredibly prohibitive on today's HBM-based GPUs, XPUs to serve those models efficiently. You end up often bandwidth bottlenecks. You end up serving at really, really low MFU, low flops utilization.

And so there are these areas where even as you build for today's world, you get to kind of look at where you can unlock greater capability. And, you know, with LLMs today, it's very similar actually with attention. There are forms of attention that are less bandwidth hungry.

They tend to be more flops hungry for a certain level of intelligence. And so one of the things I think we're excited about is as we elevate this other quality, this property of memory bandwidth, which hasn't really been scaled so much on chips recently.

We've scaled flops like a million fold in the last 20 years. Memory bandwidth has gone up about 40X in the same timeframe. So as you scale that frontier, you get to conserve more of the other thing. We get to reduce the number of flops that we're using on these models to get to a certain level of intelligence.

And so that's the kind of really a multiplying factor on global throughput for these models.

**Sarah Guo** [31:17]
I'll resist asking you if that changes people's ideas about how they should train for now. One last question for you just on market structure here. It's very unclear to me, you know, five years from now what a large AI player or a hyperscaler buys in terms of what chips they consume between the NVIDIA and AMDs of the world, the new accelerator class, and, you know, four, at least four of these players have their own internal efforts.

### Market structure

**Sarah Guo** [31:47]
Like how do you make sense of this?

**Walter Goodwin** [31:49]
So I think that today we see, you know, behavior where really everybody who is deploying at scale is kind of trying to deploy as many separate platforms as they can. I think one thing that will remain the case is that there will be a real need for kind of a diversity of supply as something like compute becomes so existential to these players.

There's a bit of a joke today that the sort of first-party efforts, their primary purpose is to reduce the price that people pay NVIDIA. And there may be some truth to that as well because those efforts are kind of architecturally quite similar,right?

So it's a bet that is not on enabling a fundamental capability that those other chips cannot enable. And so there I think it is this kind of gamesmanship of, you know, can we build some of this over here or can I buy something from AMD so that maybe I then get a better price from NVIDIA and so on.

And a play also for kind of aggregate capacity and a bit of control. I think as you then see this move into chips that actually give new capabilities, you know, that's a place where, you know, I suddenly see for what we're doing a real need for everybody that wants to deploy AI at the frontier.

They have to have some solution to running these models orders of magnitude faster. You know, for the frontier labs, there's this window of kind of premium capability, premium intelligence that is their entire raison d'être. You know, otherwise we'd all be using Kimi models all the time.

And so as you do have this pressure from behind from open source, you know, it does become clearly even more important for folks that want to deploy at the frontier to have every aspect of what that means. So the best model waits, but also then the fastest deployment so you can do the most reasoning in the shortest amount of time.

And so I think this kind of premium end of the chip space where you're optimizing for speed over everything else becomes a kind of very important part of this world. I think one thing that is notable about, you know, the dynamic between a chip vendor and a frontier lab, you know, if you look at Fractile, we try to be as aligned as possible with the needs of the frontier.

We try to look a bit like some of those teams as well. We have people that really think deeply about these workloads and so on. So one of the cases that I think we sometimes need to make is explaining why we believe that over, you know, multiple decades, the world sustains some kind of frontier third-party chip players.

You know, why does this not just roll up into these labs? And I think that comes down to the kind of asymmetric game that all of these labs and frontier players have to play where they're exposed to an enormous risk if they go all in on a hardware bet.

Suppose I'm, you know, I'm lab one and I've gone all in on some proprietary silicon and then lab two discovers some new computational breakthrough that delivers far better computational efficiencies for the same level of intelligence, but it only works on the chip that they've decided to deploy.

I could die in the nine months before I get to deploy enough of that chip that I've also now gained that kind of 5X in computational efficiency. And so there is a reallyurgent need for these chips to actually deploy the same, these companies to deploy the same platforms as one another.

They're playing different games,right? They're trying to now compete, you know, in the model layer. I think having deeply differentiated bets and going all in on those differentiated bets at the chip layer is sort of an irrational and a very dangerous move for anybody that is playing at the frontier.

**Sarah Guo** [35:15]
Well, I think that's a great place to end. Thanks so much, Walter.

### Outro

**Walter Goodwin** [35:17]
Thanks very much, Sarah.

**Sarah Guo** [35:21]
Find us on Twitter@NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
