Tools

Jalapeño: OpenAI Unveils Its First Custom Chip Built with Broadcom

25 June 2026 Mehdi 06:32
Jalapeño: OpenAI Unveils Its First Custom Chip Built with Broadcom

OpenAI just crossed a milestone many have been anticipating for months: the company has revealed its first custom processor, called Jalapeño, developed in collaboration with Broadcom. This is not a routine announcement. It signals that the silicon war for AI now has a new player in the ring.

OpenAI Enters the Dedicated Silicon Era with Jalapeño

Up until now, OpenAI ran heavily on Nvidia GPUs to power ChatGPT, Codex, and its other services. That model comes with massive costs and a structural dependency on a single vendor. Jalapeño is the concrete answer to that problem.

The chip was officially unveiled on June 24, 2026. It is designed exclusively for inference, meaning the execution of already-trained models in response to user requests. It is not a general-purpose GPU. It is also not a direct competitor to Nvidia’s training chips. It is a targeted accelerator, built for one specific use case at massive scale.

OpenAI notes that early internal results show significantly better performance per watt than current alternatives. Exact figures have not been made public yet, but the message is clear: energy efficiency was a core design priority.

A Deliberate Vertical Integration Strategy

OpenAI’s ambition goes well beyond simple cost reduction. The company wants to control its entire technology stack, from the model down to the silicon. That is what president Greg Brockman spelled out explicitly: “We have a deep understanding of the workloads. We look for specific workloads that are underserved, and we build something that can accelerate what’s possible.”

OpenAI put it in plain terms in its official announcement: the company is no longer just building frontier models or products. It is designing the infrastructure underneath, chip architecture, kernels, memory systems, networking, scheduling, and deployment. Every layer can be optimized around the same goal: making models faster, more reliable, and cheaper for users.

This approach mirrors what Google does with its TPUs and Amazon with its Inferentia and Trainium chips. OpenAI is catching up on that front, and doing so with a solid partner.

Broadcom as Manufacturing Partner: Why This Choice

Broadcom is no stranger to the semiconductor world. The company specializes in networking chips, connectivity, and accelerators for hyperscalers. That profile fits OpenAI’s needs perfectly.

The partnership between the two companies was officially announced in October, but discussions had started well before that. According to CNBC, early talks began roughly 18 months before the chip’s public announcement, placing them around early 2025. That timeline illustrates the complexity of such a project: designing a chip from scratch for workloads this specific is not something you improvise.

The most intensive tasks, such as model pre-training, will likely remain on Nvidia hardware. Jalapeño does not replace everything, it optimizes where the economic impact is greatest: production inference, which runs continuously and at massive scale.

AI Used to Design AI

One detail deserves particular attention: OpenAI’s own models contributed to the design of Jalapeño. It is an interesting feedback loop. AI tools accelerate hardware design, which will allow future AI tools to run even more efficiently.

This approach reflects a broader industry trend. LLMs are beginning to integrate into hardware engineering workflows, for layout optimization, RTL code generation, and formal verification. OpenAI is practicing what it preaches here.

Target Use Cases: Real-Time Inference and Agentic Systems

OpenAI explicitly highlighted the chip’s performance for real-time coding models. Codex is clearly in scope. But the chip is also designed to support future AI agent systems, which require continuous, complex, low-latency inference.

Here are the priority use cases as they emerge from the announcement:

  • Inference for ChatGPT and its API at scale
  • Running Codex models for real-time code generation
  • Supporting AI agent systems with continuous computational demands
  • Optimizing operational cost per watt in production

These use cases align with OpenAI’s product strategy. Inference is the cost line that explodes as the user base grows. Every efficiency gain at that level has a direct impact on margins.

Key Takeaways

  • Jalapeño is OpenAI’s first custom inference chip, built with Broadcom and announced on June 24, 2026.
  • It targets inference exclusively, not training, and delivers better performance per watt than current alternatives according to OpenAI.
  • The strategy is clear: full vertical integration of the stack, from model to silicon, to reduce dependency on Nvidia.
  • OpenAI’s own AI models contributed to the chip’s design, illustrating a continuous improvement loop.
  • Without public benchmarks or deployment volume details, it is still too early to measure the real impact on operational competitiveness.

If you work on AI architectures in production or are thinking about inference cost, this kind of move is worth watching closely. Feel free to share your questions or follow the blog for the next analyses.

Sources

Leave a comment

Your email address will not be published. Required fields are marked *