Categories
Misc

NVIDIA Blackwell Ultra Sets the Bar in New MLPerf Inference Benchmark

Inference performance is critical, as it directly influences the economics of an AI factory. The higher the throughput of AI factory infrastructure, the more tokens it can produce at a high speed — increasing revenue, driving down total cost of ownership (TCO) and enhancing the system’s overall productivity. Less than half a year since its
Read Article

Categories
Misc

NVIDIA Partners With AI Infrastructure Ecosystem to Unveil Reference Design for Giga-Scale AI Factories

At this week’s AI Infrastructure Summit in Silicon Valley, NVIDIA’s VP of Accelerated Computing Ian Buck unveiled a bold new vision: the transformation of traditional data centers into fully integrated AI factories. As part of this initiative, NVIDIA is developing reference designs to be shared with partners and enterprises worldwide — offering an NVIDIA Omniverse
Read Article

Categories
Misc

NVIDIA Rubin CPX Accelerates Inference Performance and Efficiency for 1M+ Token Context Workloads

Rendering of Rubin CPX.Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent…Rendering of Rubin CPX.

Inference has emerged as the new frontier of complexity in AI. Modern models are evolving into agentic systems capable of multi-step reasoning, persistent memory, and long-horizon context—enabling them to tackle complex tasks across domains such as software development, video generation, and deep research. These workloads place unprecedented demands on infrastructure, introducing new challenges in…

Source

Categories
Misc

NVIDIA Blackwell Ultra Sets New Inference Records in MLPerf Debut

As large language models (LLMs) grow larger, they get smarter, with open models from leading developers now featuring hundreds of billions of parameters. At the…

As large language models (LLMs) grow larger, they get smarter, with open models from leading developers now featuring hundreds of billions of parameters. At the same time, today’s leading models are also capable of reasoning, which means that they generate many intermediate reasoning tokens before delivering a final response to the user. The combination of these two trends—larger models that think…

Source

Categories
Misc

mmBERT: ModernBERT goes Multilingual

Categories
Misc

Get Started Using Generative AI for Content Creation With ComfyUI and NVIDIA RTX AI PCs

ComfyUI — an open-source, node-based graphical interface for running and building generative AI workflows for content creation — published major updates this past month, including up to 40% performance improvements for NVIDIA RTX GPUs, and support for new AI models including Wan 2.2, Qwen-Image, FLUX.1 Krea [dev] and Hunyuan3D 2.1. NVIDIA also released NVIDIA TensorRT-optimized
Read Article

Categories
Misc

How to Build AI Systems In House with Outerbounds and DGX Cloud Lepton

It’s easy to underestimate how many moving parts a real-world, production-grade AI system involves. Whether you’re building an agent that combines internal…

It’s easy to underestimate how many moving parts a real-world, production-grade AI system involves. Whether you’re building an agent that combines internal data with external LLMs or a service that generates anime on demand, the system must orchestrate multiple models and dynamic data across online and offline components. Many AI services, from LLMs to vector databases…

Source

Categories
Misc

Register for the Global Webinar: How to Prepare for NVIDIA Generative AI Certification

Join a global webinar on Oct. 7 to get everything you need to succeed on the NVIDIA generative-AI certification exams, including the new professional level…

Join a global webinar on Oct. 7 to get everything you need to succeed on the NVIDIA generative-AI certification exams, including the new professional level agentic AI and generative AI LLMs certifications. Discover exam prep strategies, see sample exam questions, and hear directly from our certification experts in a live Q&A. Register now.

Source

Categories
Misc

Just Released: NVIDIA PhysicsNeMo 25.08

NVIDIA PhysicsNeMo 25.08 is packed with powerful new workflows and recipes for CAE application developers.

NVIDIA PhysicsNeMo 25.08 is packed with powerful new workflows and recipes for CAE application developers.

Source

Categories
Misc

Accelerate Large-Scale LLM Inference and KV Cache Offload with CPU-GPU Memory Sharing

Decorative image.Large Language Models (LLMs) are at the forefront of AI innovation, but their massive size can complicate inference efficiency. Models such as Llama 3 70B and…Decorative image.

Large Language Models (LLMs) are at the forefront of AI innovation, but their massive size can complicate inference efficiency. Models such as Llama 3 70B and Llama 4 Scout 109B may require more memory than is included in the GPU, especially when including large context windows. For example, loading Llama 3 70B and Llama 4 Scout 109B models in half precision (FP16) requires approximately 140…

Source