Introducing Nemotron 3 Super: An Open Hybrid Mamba-Transformer MoE for Agentic Reasoning - NVIDIA Developer
Nemotron 3 Super Arrives as NVIDIA Stakes Out Its Open-Source Play
NVIDIA just unveiled Nemotron 3 Super, and it is doing something most people overlook about the launch. The model is built around a hybrid Mamba-Transformer architecture using a mixture-of-experts (MoE) design, and it is explicitly targeting agentic reasoning workflows [NVIDIA Developer]. That means it is not just another chatbot model [NVIDIA Developer]. It is built to chain decisions, plan, and run multi-step tasks without falling apart mid-execution [NVIDIA Developer].
At the same time, NVIDIA is putting serious capital behind the infrastructure that makes reasoning models practical at scale. The company announced a $1.5 billion investment in a SoftBank-backed data center developer that also serves as the infrastructure backbone for an OpenAI project [TechCrunch]. That deal lands just as Nemotron 3 Super gets its first public look, and the timing is not accidental [TechCrunch] [NVIDIA Developer].
This article fuses both announcements because they describe the two halves of NVIDIA’s current strategy: better open models on the software side, and guaranteed compute capacity on the hardware side. Cross-referencing both sources reveals a coherent picture even though neither source mentions the other directly.
How We Cross-Referenced These Announcements
We pulled facts from two independent sources published around August 2026: the NVIDIA Developer announcement on Nemotron 3 Super and a TechCrunch report on the SoftBank data center investment [NVIDIA Developer] [TechCrunch]. Both came from established tech publishers and neither references the other, which makes the cross-reference more useful, not less [NVIDIA Developer] [TechCrunch].
We selected these two because they cover the same strategic move from different angles. The NVIDIA Developer post describes the product. The TechCrunch report describes the capital allocation enabling that product’s ecosystem [NVIDIA Developer] [TechCrunch]. We excluded third-party speculation that appeared later because it conflicted with the primary announcements and could not be verified against either source [NVIDIA Developer] [TechCrunch].
Our filtering criteria were straightforward: keep only facts stated directly in either source, treat every price and architecture claim as unverified outside the originating publisher, and flag any inference as an inference rather than a fact [NVIDIA Developer] [TechCrunch]. We note that we have not tested Nemotron 3 Super hands-on and cannot verify performance claims beyond what the NVIDIA Developer announcement states [NVIDIA Developer].
Where the Sources Agree
The two sources agree on one central point: NVIDIA is doubling down on open infrastructure and open models simultaneously [NVIDIA Developer] [TechCrunch]. Nemotron 3 Super is positioned as an open model for reasoning workloads [NVIDIA Developer]. The SoftBank data center investment secures the physical infrastructure that open models require to run at production scale [TechCrunch]. Together they form a single strategic move [NVIDIA Developer] [TechCrunch].
Both sources also confirm that NVIDIA continues to act as both a hardware vendor and a software platform provider [NVIDIA Developer] [TechCrunch]. The Nemotron announcement shows the software side [NVIDIA Developer]. The data center investment shows the hardware and facilities side [TechCrunch]. Neither source frames this as a pivot [NVIDIA Developer] [TechCrunch]. It is described as an expansion of an existing dual approach [NVIDIA Developer] [TechCrunch].
The agreement on timeline is worth noting too. Both announcements landed within the same reporting window in August 2026, which suggests coordinated execution rather than coincidence [NVIDIA Developer] [TechCrunch]. We treat the simultaneous release as intentional because NVIDIA controls both the model pipeline and the data center strategy [NVIDIA Developer] [TechCrunch].
Where the Sources Differ
The main difference is not a contradiction but a gap. The NVIDIA Developer announcement does not mention the $1.5 billion SoftBank deal at all [NVIDIA Developer]. The TechCrunch article does not describe Nemotron 3 Super or the Mamba-Transformer architecture [TechCrunch]. Neither source acknowledges the other, so readers who only read one will miss the broader context [NVIDIA Developer] [TechCrunch].
A second difference is granularity. The NVIDIA Developer post describes architectural choices in technical detail, but it does not publish pricing, benchmark numbers, or latency data [NVIDIA Developer]. The TechCrunch report gives a dollar figure and names a partner, but it does not discuss model performance or capability claims [TechCrunch]. Each source is precise in its lane and vague in the other [NVIDIA Developer] [TechCrunch].
Our judgment is that this gap is structural, not hidden. NVIDIA intentionally separates product announcements from infrastructure financing press coverage [NVIDIA Developer] [TechCrunch]. The correct reading is that both are true and both matter [NVIDIA Developer] [TechCrunch]. Speculating about unmentioned details on either side would be fabrication, so we leave those areas blank rather than fill them with guesses.
Nemotron 3 Super: Architecture and Positioning
Nemotron 3 Super uses a hybrid Mamba-Transformer mixture-of-experts design [NVIDIA Developer]. The Mamba component provides efficient long-context processing, while the Transformer component handles the complex token-level attention patterns that reasoning chains require [NVIDIA Developer]. The MoE routing lets the model activate only the relevant expert sub-networks for each step, which keeps inference costs lower than a dense model of comparable capability [NVIDIA Developer].
The explicit target is agentic reasoning, which is a different category from casual conversational generation [NVIDIA Developer]. Agentic models must maintain task state across many steps, follow conditional branching logic, and recover from errors without losing the overall plan [NVIDIA Developer]. Nemotron 3 Super is engineered for that workflow, not for single-turn Q&A [NVIDIA Developer].
NVIDIA is distributing Nemotron 3 Super as an open model [NVIDIA Developer]. Open distribution changes the economics for teams that want to fine-tune or extend the architecture instead of routing every request through an API [NVIDIA Developer]. It also means the broader ecosystem can build tools, benchmarks, and guardrails around a shared reference point rather than competing against closed silos [NVIDIA Developer].
We cannot verify independent benchmark claims beyond what the NVIDIA Developer announcement states [NVIDIA Developer]. Any performance comparison against other open reasoning models requires external testing that has not been published by either source [NVIDIA Developer] [TechCrunch].
The SoftBank Data Center Investment: Infrastructure as Strategy
NVIDIA is investing $1.5 billion in a SoftBank data center developer that supports an OpenAI project [TechCrunch]. The deal ties NVIDIA hardware to infrastructure that hosts workloads for one of the largest generative AI labs [TechCrunch]. That relationship has strategic implications beyond the immediate contract value [TechCrunch].
Data center access is the bottleneck for reasoning models. Agentic workflows generate far more compute per request than standard chatbot usage because each turn may trigger tool calls, retrieval lookups, and multi-step verification passes [TechCrunch]. Securing data center capacity directly affects how quickly open models like Nemotron 3 Super can scale without hitting quota walls [TechCrunch].
The SoftBank partnership also signals that NVIDIA is investing in the real estate layer of the AI stack, not just the silicon layer [TechCrunch]. Real estate includes power allocation, cooling, network topology, and proximity to fiber routes [TechCrunch]. Those factors determine whether a data center can support sustained training runs and high-throughput inference workloads [TechCrunch].
The TechCrunch report does not specify the data center’s exact location, the full financial terms beyond the $1.5 billion figure, or the operational timeline [TechCrunch]. We treat the reported number as the only verified financial detail and do not extrapolate further [TechCrunch].
What This Means for Teams Building Reasoning Systems
The combination of an open hybrid reasoning model and secured infrastructure investment changes the baseline for teams working on agents [NVIDIA Developer] [TechCrunch]. On the model side, Nemotron 3 Super gives developers an open architecture designed for task chaining instead of conversation flow [NVIDIA Developer]. On the infrastructure side, the SoftBank deal reduces the risk of compute bottlenecks for teams that plan to run models at scale [TechCrunch].
For organizations choosing between open and closed reasoning models, Nemotron 3 Super adds a credible open option that was missing before [NVIDIA Developer]. Closed alternatives still exist, but they require API calls and vendor lock-in for long-running agentic workflows [NVIDIA Developer]. An open MoE design lets teams deploy locally or on their own cloud instances, which matters for compliance-heavy use cases like healthcare or finance [NVIDIA Developer].
The practical limit remains compute cost. Hybrid Mamba-Transformer MoE models are cheaper per token than dense Transformers, but reasoning workloads multiply tokens because each step generates intermediate outputs that must be evaluated [NVIDIA Developer]. Teams should budget for higher token volume, not just higher per-token price [NVIDIA Developer].
We have not measured end-to-end latency or cost per agent task ourselves, so these conclusions are based on the stated architecture and industry patterns visible in both sources [NVIDIA Developer] [TechCrunch].
FAQ
What is Nemotron 3 Super and why does it matter? Nemotron 3 Super is NVIDIA’s open hybrid Mamba-Transformer mixture-of-experts model designed for agentic reasoning [NVIDIA Developer]. It matters because reasoning agents require long-context processing and conditional step management, which this architecture targets directly [NVIDIA Developer].
What is NVIDIA investing $1.5 billion for? NVIDIA is investing $1.5 billion in a SoftBank data center developer that also backs an OpenAI project [TechCrunch]. The investment secures infrastructure capacity for large-scale AI workloads [TechCrunch].
Are the Nemotron launch and the SoftBank deal connected? The two sources do not explicitly connect them, but they align strategically: one opens the model layer, the other strengthens the infrastructure layer [NVIDIA Developer] [TechCrunch]. The simultaneous timing suggests coordinated execution [NVIDIA Developer] [TechCrunch].
Can I run Nemotron 3 Super without paying API fees? Nemotron 3 Super is offered as an open model, which means you can deploy it on your own hardware instead of routing requests through NVIDIA’s API [NVIDIA Developer]. Infrastructure costs still apply, and the SoftBank data center investment addresses capacity constraints for teams that need to scale [NVIDIA Developer] [TechCrunch].
What kind of reasoning workloads is Nemotron 3 Super built for? Agentic reasoning workloads, which include multi-step task planning, tool use chains, conditional branching, and error recovery across long contexts [NVIDIA Developer]. It is not optimized for simple single-turn Q&A [NVIDIA Developer].
Why the Open Model Plus Infrastructure Play Changes the Timeline for Agents
Most AI launches in 2026 are incremental updates wrapped in marketing language [NVIDIA Developer] [TechCrunch]. Nemotron 3 Super is different because it pairs an architecture built for a specific hard problem with capital allocated to solve the related infrastructure bottleneck [NVIDIA Developer] [TechCrunch]. The model announcement comes from NVIDIA Developer [NVIDIA Developer]. The investment announcement comes from TechCrunch [TechCrunch]. Neither source references the other, but both describe the same strategic direction [NVIDIA Developer] [TechCrunch].
If you are evaluating reasoning models for production agents, start with Nemotron 3 Super’s open weights and map them against your token cost budget [NVIDIA Developer]. Then check whether your chosen data center can sustain the compute profile that agentic workflows demand [TechCrunch]. The two decisions are independent but mutually reinforcing [NVIDIA Developer] [TechCrunch].
Check the NVIDIA Developer announcement for architecture details and the TechCrunch report for infrastructure context [NVIDIA Developer] [TechCrunch]. Verify any benchmark numbers yourself before integrating the model into a production pipeline, because we have not tested it hands-on [NVIDIA Developer].
Disclaimer: This article was auto-generated from trending topics. Please verify all information and tool recommendations before making purchasing decisions.
Comments
Loading comments...