Categories
Misc

NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf Debut

MLPerf Inference v5.1 is the latest version of the MLPerf Inference industry standard benchmark. With benchmark rounds held twice per year, the benchmark features many tests of AI inference performance and is regularly updated with new models and scenarios. This round, NVIDIA submitted results in the available category using the GB300 NVL72 rack-scale system, the first-ever MLPerf submissions using the Blackwell Ultra architecture.
Read Article

Categories
Misc

How to Connect Distributed Data Centers Into Large AI Factories with Scale-Across Networking

AI scaling is incredibly complex, and new techniques in training and inference are continually demanding more out of the data center. While data center…

AI scaling is incredibly complex, and new techniques in training and inference are continually demanding more out of the data center. While data center capabilities are scaling quickly, data center infrastructure is subject to fundamental physical limitations that have no impact on algorithms and models. Power availability, cooling capacity, and space constraints place limits on the physical…

Source

Categories
Misc

‘Safety First, Always,’ NVIDIA VP of Automotive Says, Unveiling the Future of AI-Defined Vehicles at IAA Mobility

At this week’s IAA Mobility conference in Munich, NVIDIA Vice President of Automotive Ali Kani outlined how cloud-to-car AI platforms are bringing new levels of safety, intelligence and trust to the road. NVIDIA and its partners didn’t just show off cars at the conference — they showed off what cars are becoming: AI-defined machines, built
Read Article

Categories
Misc

NVIDIA Rubin CPX Accelerates Inference Performance and Efficiency for 1M+ Token Context Workloads

Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent memory, and long-horizon context—enabling them to tackle complex tasks across domains such as software development, video generation, and deep research. These workloads place unprecedented demands on infrastructure, introducing new challenges in compute, memory, and networking that require a fundamental rethinking of how inference is scaled and optimized. This blog explores the next evolution in disaggregated inference infrastructure and introduces NVIDIA Rubin CPX—a purpose-built GPU designed to meet the demands of long-context AI workloads with greater efficiency and ROI.
Read Article

Categories
Misc

NVIDIA Partners With AI Infrastructure Ecosystem to Unveil Reference Design for Giga-Scale AI Factories

At this week’s AI Infrastructure Summit in Silicon Valley, NVIDIA’s VP of Accelerated Computing Ian Buck unveiled a bold new vision: the transformation of traditional data centers into fully integrated AI factories. As part of this initiative, NVIDIA is developing reference designs to be shared with partners and enterprises worldwide — offering an NVIDIA Omniverse
Read Article

Categories
Misc

NVIDIA Blackwell Ultra Sets the Bar in New MLPerf Inference Benchmark

Inference performance is critical, as it directly influences the economics of an AI factory. The higher the throughput of AI factory infrastructure, the more tokens it can produce at a high speed — increasing revenue, driving down total cost of ownership (TCO) and enhancing the system’s overall productivity. Less than half a year since its
Read Article

Categories
Misc

NVIDIA Unveils Rubin CPX: A New Class of GPU Designed for Massive-Context Inference

NVIDIA® today announced NVIDIA Rubin CPX, a new class of GPU purpose-built for massive-context processing. This enables AI systems to handle million-token software coding and generative video with groundbreaking speed and efficiency.

Categories
Misc

NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf Debut

As large language models (LLMs) grow larger, they get smarter, with open models from leading developers now featuring hundreds of billions of parameters. At the…

As large language models (LLMs) grow larger, they get smarter, with open models from leading developers now featuring hundreds of billions of parameters. At the same time, today’s leading models are also capable of reasoning, which means that they generate many intermediate reasoning tokens before delivering a final response to the user. The combination of these two trends—larger models that think…

Source

Categories
Misc

NVIDIA Rubin CPX Accelerates Inference Performance and Efficiency for 1M+ Token Context Workloads

Rendering of Rubin CPX.Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent…Rendering of Rubin CPX.

Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent memory, and long-horizon context—enabling them to tackle complex tasks across domains such as software development, video generation, and deep research. These workloads place unprecedented demands on infrastructure, introducing new challenges in…

Source

Categories
Misc

mmBERT: ModernBERT goes Multilingual