OpenAI presented the first real performance numbers for Jalapeño, its custom AI chip, at the Hot Chips conference on August 25, 2026. Unlike the powerful, general-purpose GPUs Nvidia sells for both training and running AI models, Jalapeño is built to do just one job: inference, the process of actually generating a response once a model has already been trained. That’s the part of AI people experience directly every time they open a chatbot and wait for a reply, and it’s also where the ongoing costs of running AI at scale really add up.

Jalapeño was co-developed with Broadcom, which handled chip and networking design, and Celestica, which managed systems integration — part of a broader deal announced in October 2025 to build out 10 gigawatts of custom AI accelerators. Design work reportedly began in mid-2024, meaning OpenAI went from assembling its hardware team to a finished chip design in roughly 16 months, an unusually fast timeline for custom silicon.
The Benchmark Numbers
OpenAI tested Jalapeño against Nvidia’s current-generation Blackwell systems using InferenceX, a public benchmark maintained by semiconductor research firm SemiAnalysis, across three open-weight AI models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The results, independently verified by SemiAnalysis engineers who visited OpenAI’s labs and ran the tests themselves, showed Jalapeño delivering 1.5 to 1.9 times more AI work per watt of power at peak throughput, along with 1.7 to 3.6 times lower end-to-end response latency compared with the Nvidia systems tested.
In one specific comparison, against an Nvidia GB200 system running the GPT-OSS 120B model, OpenAI reported roughly 1.9 times higher throughput per kilowatt and a response latency of about 1.03 seconds, compared with 1.80 seconds on the Nvidia system. SemiAnalysis CEO Dylan Patel called the results “industry leading,” noting that Jalapeño “smokes every other chip” the firm has tested on multiple top open-source models — a notable statement given that first-generation custom chips rarely perform competitively against established rivals right out of the gate.
Why OpenAI Isn’t Ditching Nvidia
Despite the strong benchmark showing, OpenAI was careful to frame Jalapeño as additive rather than a replacement for its existing hardware relationships. “Meeting growing demand for AI will require more compute from every available source,” the company wrote in its announcement. “We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads.” OpenAI plans only a very small-scale deployment of Jalapeño by the end of 2026, with meaningful production volume not arriving until 2027.
The Catch Behind the Headlines
A few caveats temper how far these results should be extrapolated. SemiAnalysis noted that a fairer comparison might actually be Nvidia’s newer Vera Rubin platform rather than Blackwell, since both use the same HBM4 memory standard — and even against Vera Rubin, Jalapeño still edged ahead on output tokens per megawatt, though the two came out roughly even on total cost of ownership per token. It’s also worth noting that Nvidia and AMD have already published benchmark results using larger, newer models that haven’t yet been tested against Jalapeño, and Nvidia’s hardware lineup won’t stand still between now and Jalapeño’s 2027 production ramp.
What This Means for the AI Chip Race
OpenAI’s benchmark results matter less as a declaration of victory over Nvidia and more as evidence that the push toward custom AI silicon among major AI companies is producing genuinely competitive results, not just cost-saving insurance policies. Google, Amazon, and Meta are all pursuing similar custom chip strategies, and OpenAI’s numbers suggest that a company without decades of chip-design experience can still produce hardware that holds its own against the market leader on specific, well-defined tasks. For everyday users of AI tools, the practical impact should eventually show up as faster responses and, potentially, lower costs for AI services, as more efficient chips work their way into the infrastructure powering them.