OpenAI’s Jalapeño Chip Shows Up to 4.1× Lower Latency in New AI Tests
OpenAI has published the first detailed performance results for Jalapeño, its custom AI inference chip, and the numbers show why the company is building its own silicon. In tests across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI says Jalapeño delivered 1.5× to 1.9× more AI work per watt at peak throughput and 1.7× to 3.6× lower end-to-end latency than the comparison systems. For one token-generation metric, OpenAI reports up to 4.1× better performance.
The announcement matters because inference hardware directly affects how quickly AI services can respond and how efficiently providers can serve millions of requests. OpenAI says it plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026.

What is OpenAI’s Jalapeño chip?
Jalapeño is OpenAI’s first custom inference accelerator, developed with Broadcom and designed specifically around the demands of large language model inference. Unlike general-purpose processors, it is built around the way modern AI systems process prompts, generate tokens and move model data through memory and networking.
OpenAI says the architecture is designed to reduce data movement and communication delays while balancing compute, memory and networking. The company originally unveiled Jalapeño in June and said it was the first step in a multigenerational hardware platform.
NewsHulk has previously covered the wider shift toward more capable AI agents and the security challenges surrounding them in AI Agent Security in 2026.
How fast is Jalapeño?
OpenAI tested Jalapeño using InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. The company compared its system with commercially available AI systems across different operating conditions, from high-throughput workloads to interactive, low-latency use.
- 1.5× to 1.9× more AI work per watt at peak throughput across the three tested models.
- 1.7× to 3.6× lower end-to-end latency than the comparison systems.
- 2.1× to 4.1× higher performance for highly interactive workloads, according to OpenAI.
On Kimi K2.5 1T, OpenAI reports approximately 1.5× higher peak performance per watt and 3.4× lower end-to-end latency than its comparison system. The company says Jalapeño remained on the Pareto frontier across the tested operating range.
Does Jalapeño beat Nvidia?
OpenAI’s results are significant, but the headline should be interpreted carefully. The company compared Jalapeño with specific commercially available systems used in the benchmark, including Nvidia-based systems. That does not mean OpenAI has made Nvidia obsolete or that Jalapeño is faster in every AI workload.
Jalapeño is primarily an inference accelerator. OpenAI says it will continue to use accelerators from Nvidia and other partners for both training and inference. The more important development is that OpenAI is adding its own silicon to the stack rather than relying entirely on external hardware.
Why OpenAI building its own chip matters
AI inference is the part of the AI stack users experience directly. Every time ChatGPT generates an answer, an agent completes a task or an API produces tokens, compute is being consumed.
More performance per watt can allow an AI provider to handle more work from the same amount of electricity and hardware. Lower latency can make conversational AI feel faster, while better efficiency can help providers manage rising demand without increasing infrastructure costs at the same rate.
This is especially important as AI agents perform longer, multi-step tasks. A small delay at every step can compound into a noticeably slower overall experience.
AI helped design the Jalapeño chip
One of the most interesting details is that OpenAI says its own AI models helped accelerate Jalapeño’s development. The company says it moved from initial design to tapeout in nine months by using AI to explore implementations, shorten verification loops and optimize parts of the chip.
OpenAI also says AI-generated implementations for selected GPT-OSS attention and mixture-of-experts blocks ran 1.5× to 1.8× faster than existing human-written implementations. Those figures apply only to selected blocks, not the complete model, so they should not be interpreted as a claim that AI-generated code makes the entire chip 1.8× faster.
When will Jalapeño be available?
OpenAI says it plans to begin deploying Jalapeño inside its compute infrastructure by the end of 2026. Initial deployment is expected to be limited, with larger-scale deployment planned as the platform matures.
The company describes Jalapeño as the first generation of a multigenerational roadmap, with a second generation already deep in development and a third generation taking shape.
What does Jalapeño mean for ChatGPT users?
ChatGPT users should not expect a new “Jalapeño mode” to appear in the app. The impact is more likely to be indirect: faster responses, better capacity during periods of heavy demand and potentially improved economics for serving increasingly capable models.
Whether users actually notice a major speed improvement will depend on how OpenAI deploys the hardware, which models run on it and how the company balances latency, throughput and cost.
Frequently asked questions
What is Jalapeño?
Jalapeño is OpenAI’s custom AI inference accelerator, developed with Broadcom for large-scale language-model workloads.
How much faster is Jalapeño?
OpenAI reports 1.7× to 3.6× lower end-to-end latency across its three public-model comparisons, with up to 4.1× higher performance on a token-generation metric for interactive workloads.
Will Jalapeño replace Nvidia?
No. OpenAI says it will continue to deploy Nvidia and other accelerators. Jalapeño expands OpenAI’s hardware options rather than replacing every external accelerator.
When will OpenAI deploy Jalapeño?
OpenAI says it plans to begin deploying the chip in its compute infrastructure by the end of 2026.
Bottom line
Jalapeño is more important than a single chip benchmark. OpenAI is trying to control more of the AI stack, from models and software to inference hardware. If the company can turn the benchmark gains into reliable large-scale deployment, the result could be faster AI responses, better power efficiency and a more diversified hardware strategy.
