computetrail

All models ·nvidia

NVIDIA: Nemotron 3 Ultra (batch)

Provider: nvidia · nvidia/nemotron-3-ultra-550b-a55b:batch

Input
$0.6 /1M
Output
$3.6 /1M
Cache read
$0.2 /1M
Context
512,288 tokens

How NVIDIA: Nemotron 3 Ultra (batch) pricing compares

At $3.6 / 1M output tokens, NVIDIA: Nemotron 3 Ultra (batch) is the 5th-cheapest of 5 nvidia models we track. It's cheaper than 37% of the 392 paid models on computetrail (246th-cheapest overall). Output costs 6.0× its input price ($0.6 / 1M in). For a typical 3:1 input-to-output workload that works out to about $1.35 / 1M blended. Reading cached input costs 67% less than fresh input, so prompt caching cuts repeat-context cost sharply. Its 512,288 tokens context window ranks 132nd-largest of 392 models with a published limit.

Output price history

$3.60$1.8008-0708-1708-27

Across 21 daily snapshots since 2026-08-07, the output price has risen 100.0% — from $1.8 to $3.6 / 1M.

Similar-priced models

ModelProviderOutput $/1MContext
NVIDIA: Nemotron 3 Ultranvidia$3.6512,288
Z.ai: GLM 5.2z-ai$3.741,048,576
Google: Gemini 3.6 Flashgoogle$3.751,048,576
MoonshotAI: Kimi K2.7 Codemoonshotai$3.4262,144
Qwen: Qwen3 Maxqwen$3.9262,144
Z.ai: GLM 5.1z-ai$3.96204,800

Prices are recorded daily from public provider data via OpenRouter and normalized to USD per 1M tokens. See methodology.