OpenAI unveiled its first in-house AI chip, code-named “Jalapeno,” at the Hot Chips conference this week, saying benchmark tests show it delivers more AI inference work per watt than Nvidia’s current Blackwell systems.
What the Benchmarks Showed
OpenAI said Jalapeno delivered 1.5 to 1.9 times more AI work per watt at peak throughput across three tested models, with 1.7 to 3.6 times lower end-to-end latency than the best commercially available systems at the time of testing. The chip, developed with Broadcom, handles inference only, meaning it runs already-trained AI models rather than training new ones.
An Important Caveat
Analysts have noted the comparison favors Jalapeno somewhat unfairly, since it uses newer HBM4 memory that Nvidia’s currently shipping Blackwell chips don’t have. A more direct comparison is Nvidia’s upcoming Vera Rubin platform, which also uses HBM4; OpenAI says Jalapeno still produces more output tokens per megawatt than Vera Rubin, even though Nvidia’s chip uses a multi-token prediction optimization Jalapeno hasn’t adopted yet.
What Happens Next
OpenAI plans to deploy Jalapeno within its own compute infrastructure by the end of the year and says it is already developing the chip’s second and third generations. The announcement adds OpenAI to a growing list of AI companies, including Google and Amazon, building custom silicon to reduce their reliance on Nvidia.







