AI的进步,当整个系统协同提升时,其复合效应最为显著。这正是我对OpenAI计算策略的思考方式:一个整合的系统,涵盖数据中心与芯片、前沿模型、开发者平台、消费级与企业级产品,以及AI原生设备,每一层都强化着下一层。
更优秀的软件让硬件更具生产力。针对我们的工作负载设计的硬件则提升速度与效率。更强的模型解锁更好的产品,进而产生更多需求、使用和学习。这些信号回流至整个系统,帮助我们再次改进。
今天,我们分享了Jalapeño——OpenAI首款定制推理芯片——的首批实测性能结果。在InferenceX(一个使用GPT‑OSS 120B的公开基准测试)上,Jalapeño在每千瓦峰值吞吐量上高于对比中的商业系统,且token延迟更低。它在DeepSeek R1和Kimi K2上也表现出色,表明其优势跨越了不同模型系列。
Jalapeño在先前最佳TBT基础上进一步扩大领先优势
Jalapeño让我们对模型的运行方式及其服务的经济性有了更强的掌控。通过将模型、服务软件、芯片、内存和网络协同开发,我们能够将吞吐量、延迟、能效和成本作为一个整体系统来优化。它在我们使用的其他合作伙伴加速器之外,开辟了一条可靠的第一方路径,增强了我们为每项工作负载匹配最强大系统并以合理经济性运行的能力。我们现在已拥有具备实测结果的可工作的第一方硅片,未来几代产品也已在推进中。
为广度而建,以自主掌控获取杠杆
不同的工作负载对系统提出不同要求。前沿训练、高容量推理和常驻代理在芯片、软件、网络、功耗和延迟方面有着各异的诉求。
我们的目标是保持在帕累托前沿:持续为每项工作负载寻求能力、速度、可靠性、效率和成本的最优组合。不同芯片和供应商在不同维度上各有领先,而前沿也在不断移动。
我们的组合提供了满足这些需求的广度。微软的计算能力和NVIDIA的芯片一直是OpenAI成长的基石。如今,我们的组合还包括AWS、AMD、博通、Cerebras、CoreWeave、甲骨文、SB Energy和软银。每家都在云基础设施、加速计算、低延迟推理、数据中心开发和能源供应方面带来不同的优势。
我们积极管理这一组合,兼顾能力与经济性。在能力至上的地方使用高端系统,在规模与成本更关键的地方优化效率。在供应商、硬件和部署模式之间保持可信的选择,使我们能够将需求导向每美元性能最优的方向,在市场条件变化时维持定价纪律,并在更强技术出现时紧随前沿。直接掌控在更紧密的整合能改善整个系统的地方增加了杠杆。我们在生态系统能帮助我们更快推进的地方开展合作,在协同设计能创造显著优势的地方进行自建。
数据中心是另一个杠杆点。佐治亚州的山茶花项目展示了我们如何围绕客户工作负载设计设施,同时创造就业、支持本地企业、承担项目基础设施和能源成本、通过闭环系统节约用水,并接受年度独立公开审计以履行承诺。
将效率转化为经济价值
这一系统的价值以其产出衡量:每一单位计算产生更多有用的智能。
更好的模型用更少的尝试就能得出正确答案。更智能的路由和上下文管理减少了无效工作。优化的软件和专用硬件提升了速度和能效。
在Artificial Analysis编码代理指数上,GPT‑5.6 Sol在最大推理模式下达到新高,同时比另一领先模型少用了54%的输出token。对客户而言,这些改进意味着更快的结果、更可靠的产品、更少的重试、能完成更长工作流的代理,以及成功工作的更低总成本。最佳经济性来自每美元的有用智能。
随着有用智能变得更强大、更实惠,更多工作在经济上变得可行。一家公司可以为每位客户提供定制分析、审查每份合同、运行实时财务情景,并帮助工程师测试更多想法。这就是杰文斯悖论:更高的效率让更多用途变得有价值,通过完成更多工作、做出更好决策、推出更多产品、创造更多收入,扩大消费并催生新的经济活动。
复合优势
更具生产力的计算和更具竞争力的供应基础,帮助我们以更低成本服务更多客户,并将这些效率提升传递给用户。增长带来的资金持续投入研究、基础设施和安全。这就是OpenAI的复合优势:更好的技术创造更好的经济性,更好的经济性为下一波进步提供资金,而每一项成果都让整个系统更加强大。
Progress in AI compounds fastest when the entire system improves together. That is how I think about OpenAI’s compute strategy: one integrated system spanning data centers and chips, frontier models, our developer platform, consumer and enterprise products, and AI-native devices, with each layer strengthening the next.
Better software makes hardware more productive. Hardware designed for our workloads improves speed and efficiency. More capable models unlock better products, which generate more demand, usage, and learning. Those signals flow back through the system and help us improve it again.
Today, we shared the first measured performance results from Jalapeño, OpenAI’s first custom inference chip. On InferenceX, a public benchmark using GPT‑OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems in the comparison. It also performed strongly on DeepSeek R1 and Kimi K2, showing that its gains extend across model families.
Jalapeño widens the lead at previous-best TBT
Jalapeño gives us greater control over how our models run and over the economics of serving them. By developing the model, serving software, chip, memory, and network together, we can improve throughput, latency, energy efficiency, and cost as one system. It creates a credible first-party path alongside the accelerators we use from other partners, expanding our ability to match each workload to the strongest system at the right economics. We now have working first-party silicon with measured results, and future generations are already underway.
Build for breadth, own for leverage
Different workloads place different demands on the system. Frontier training, high-volume inference, and always-on agents have different requirements across chips, software, networks, power, and latency.
Our goal is to stay on the Pareto frontier: continually seeking the strongest mix of capability, speed, reliability, efficiency, and cost for each workload. Different chips and providers lead on different dimensions, and the frontier keeps moving.
Our portfolio gives us the range to meet those needs. Microsoft’s compute and NVIDIA’s chips have been foundational to OpenAI’s growth. Today, our portfolio also includes AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. Each brings different strengths across cloud infrastructure, accelerated computing, low-latency inference, data-center development, and energy delivery.
We actively manage this portfolio for both capability and economics. We use premium systems where capability matters most and optimize for efficiency where scale and cost matter more. Preserving credible choice across providers, hardware, and deployment models lets us direct demand toward the strongest performance per dollar, maintain pricing discipline as market conditions change, and move with the frontier as stronger technology emerges. Direct control adds leverage where tighter integration can improve the entire system. We partner where the ecosystem helps us move faster and build where co-design creates a meaningful advantage.
Data centers create another point of leverage. Project Camellia in Georgia shows how we can design facilities around customer workloads while creating jobs, supporting local businesses, covering project infrastructure and energy costs, conserving water through a closed-loop system, and subjecting its commitments to an annual independent public audit.
Turning efficiency into economic value
The value of this system is measured by what it produces: more useful intelligence from every unit of compute.
Better models reach the right answer with fewer attempts. Smarter routing and context management reduce wasted work. Optimized software and purpose-built hardware improve speed and energy efficiency.
On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model. For customers, improvements like these mean faster results, more dependable products, fewer retries, agents that complete longer workflows, and a lower total cost for successful work. The best economics come from useful intelligence per dollar.
As useful intelligence becomes more capable and affordable, more work becomes economically practical. A company can provide tailored analysis to every customer, review every contract, run live financial scenarios, and help engineers test more ideas. This is Jevons paradox: greater efficiency makes more uses worthwhile, expanding consumption and creating new economic activity through more work completed, better decisions, more products launched, and more revenue generated.
A compounding advantage
More productive compute and a more competitive supply base help us serve more customers at lower cost and carry those efficiency gains through to users. Growth funds continued investment in research, infrastructure, and safety. That is OpenAI’s compounding advantage: better technology creates better economics, better economics fund the next wave of progress, and every gain makes the whole system stronger.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文