NVIDIA: Nemotron 3 Ultra
Provider: nvidia · nvidia/nemotron-3-ultra-550b-a55b
How NVIDIA: Nemotron 3 Ultra pricing compares
At $3.6 / 1M output tokens, NVIDIA: Nemotron 3 Ultra is the 4th-cheapest of 5 nvidia models we track. It's cheaper than 38% of the 392 paid models on computetrail (245th-cheapest overall). Output costs 6.0× its input price ($0.6 / 1M in). For a typical 3:1 input-to-output workload that works out to about $1.35 / 1M blended. Reading cached input costs 67% less than fresh input, so prompt caching cuts repeat-context cost sharply. Its 512,288 tokens context window ranks 131st-largest of 392 models with a published limit.
Output price history
Across 60 daily snapshots since 2026-06-29, the output price has risen 63.6% — from $2.2 to $3.6 / 1M.
Similar-priced models
| Model | Provider | Output $/1M | Context |
|---|---|---|---|
| NVIDIA: Nemotron 3 Ultra (batch) | nvidia | $3.6 | 512,288 |
| Z.ai: GLM 5.2 | z-ai | $3.74 | 1,048,576 |
| Google: Gemini 3.6 Flash | $3.75 | 1,048,576 | |
| MoonshotAI: Kimi K2.7 Code | moonshotai | $3.4 | 262,144 |
| Qwen: Qwen3 Max | qwen | $3.9 | 262,144 |
| Z.ai: GLM 5.1 | z-ai | $3.96 | 204,800 |
Prices are recorded daily from public provider data via OpenRouter and normalized to USD per 1M tokens. See methodology.