computetrail

All models ·nvidia

NVIDIA: Nemotron 3 Ultra

Provider: nvidia · nvidia/nemotron-3-ultra-550b-a55b

Input
$0.6 /1M
Output
$3.6 /1M
Cache read
$0.2 /1M
Context
512,288 tokens

How NVIDIA: Nemotron 3 Ultra pricing compares

At $3.6 / 1M output tokens, NVIDIA: Nemotron 3 Ultra is the 4th-cheapest of 5 nvidia models we track. It's cheaper than 38% of the 392 paid models on computetrail (245th-cheapest overall). Output costs 6.0× its input price ($0.6 / 1M in). For a typical 3:1 input-to-output workload that works out to about $1.35 / 1M blended. Reading cached input costs 67% less than fresh input, so prompt caching cuts repeat-context cost sharply. Its 512,288 tokens context window ranks 131st-largest of 392 models with a published limit.

Output price history

$3.60$2.2006-2907-2908-27

Across 60 daily snapshots since 2026-06-29, the output price has risen 63.6% — from $2.2 to $3.6 / 1M.

Similar-priced models

ModelProviderOutput $/1MContext
NVIDIA: Nemotron 3 Ultra (batch)nvidia$3.6512,288
Z.ai: GLM 5.2z-ai$3.741,048,576
Google: Gemini 3.6 Flashgoogle$3.751,048,576
MoonshotAI: Kimi K2.7 Codemoonshotai$3.4262,144
Qwen: Qwen3 Maxqwen$3.9262,144
Z.ai: GLM 5.1z-ai$3.96204,800

Prices are recorded daily from public provider data via OpenRouter and normalized to USD per 1M tokens. See methodology.