OpenAI has a chip. It is called Jalapeño. Broadcom built it. That is already more than most AI hardware announcements deliver — a name, a partner, and a workload.
The workload is inference. Not training. That is a real design decision, not a hedge. Training and inference pull in opposite directions: training wants maximum throughput across a long run, inference wants low latency on a short one. Building an ASIC for one means accepting it is wrong for the other. OpenAI has chosen a side.
Broadcom is the right partner for this. They have done custom silicon for Google — the TPU family runs through Broadcom's networking and packaging lines — and they understand what it takes to get an ASIC from a block diagram to a shipping rack. OpenAI did not build this alone, and should not pretend otherwise. Broadcom built it.
What OpenAI has not released: clock speed, die size, power envelope, tokens per second per dollar, or yield numbers. Without those, Jalapeño is a name and a press release. The engineering claim is that it runs inference cheaper or faster than the Nvidia hardware it presumably replaces. That claim needs a number. No number has appeared.
The tradeoff is visible even without the spec sheet. An inference-only ASIC gives up flexibility. If OpenAI's model architecture shifts — attention head counts, context window, quantization scheme — the chip either adapts or it does not. Fixed silicon does not patch easily. That is the bet: that the inference workload is stable enough to justify freezing it in metal.
Maybe it is. The number will say.