# Nemotron 3 Ultra 550B A55B

Nemotron 3 Ultra is NVIDIA's open-weight language model for reasoning, coding, and coordinating tasks through tools. Its hybrid architecture combines attention with state-space and mixture-of-experts layers, and applications can turn reasoning on or off. Its published context window is 262,144 tokens. The advertised input rate is $0.60 per million input tokens. The advertised output rate is $2.40 per million output tokens.

- Model identifier: ` nvidia/nemotron-3-ultra-550b-a55b `
- Creator: NVIDIA
- Providers: Fireworks
- Model type: Text generation
- Canonical model page: [Nemotron 3 Ultra 550B A55B](<https://impossibl.com/nvidia/nemotron-3-ultra-550b-a55b>)
- API availability: Listed as serving through Impossibl.

## Providers

Fireworks hosts Nemotron 3 Ultra 550B A55B. Impossibl routes your request intelligently and handles failover automatically when a provider is down.

| Provider | Endpoint location | Routes | Verified context window | Tool calling |
| --- | --- | --- | --- | --- |
| Fireworks | Not specified | 1 | Not verified | Not verified |

The model price applies whichever provider serves the request. Endpoint locations do not guarantee inference residency.

## Specs and Capabilities

- Context window: 262,144 tokens
- Maximum output tokens: Not published in the public catalog; this is separate from the context window.
- Inputs: text
- Outputs: text
- Tool calling: Not published

## API access

Use ` nvidia/nemotron-3-ultra-550b-a55b ` as the model identifier through the Impossibl API.

Published API endpoints:

- ` /v1/chat/completions `

[API quickstart](<https://impossibl.com/docs/agent-quickstart>) · [Models and billing documentation](<https://impossibl.com/docs/models>)

## Published API prices

All rates are in USD. An unpublished rate is unknown, not zero. Workspace-specific pricing, taxes, balance-purchase fees, billing adjustments, and tool charges can affect the amount paid.

| Usage | Published base rate | Billing unit |
| --- | --- | --- |
| Input | $0.60 | per million input tokens |
| Output | $2.40 | per million output tokens |
| Cache read | $0.12 | per million cached input tokens |
| Cache write (default retention) | Not published | per million tokens |
| Cache write (one-hour retention) | Not published | per million tokens |
| Audio input | Not published | per million audio input tokens |

## Availability and retirement

The model is listed as serving. No retirement is announced in the current public catalog.

## Frequently asked questions

### What is Nemotron 3 Ultra 550B A55B?

Nemotron 3 Ultra is NVIDIA's open-weight language model for reasoning, coding, and coordinating tasks through tools. Its hybrid architecture combines attention with state-space and mixture-of-experts layers, and applications can turn reasoning on or off.

Sources: [NVIDIA model card](<https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4>); [Fireworks model documentation](<https://fireworks.ai/models/fireworks/nemotron-3-ultra-nvfp4>)

### How much does Nemotron 3 Ultra 550B A55B cost?

Nemotron 3 Ultra 550B A55B costs $0.60 per million input tokens and $2.40 per million output tokens through the Impossibl API. Cache reads cost $0.12 per million tokens. All rates are in USD.

### What is the context length of Nemotron 3 Ultra 550B A55B?

Nemotron 3 Ultra 550B A55B has a 262,144-token context window through the Impossibl API.

### What inputs and outputs does Nemotron 3 Ultra 550B A55B support?

Nemotron 3 Ultra 550B A55B accepts text as input. Nemotron 3 Ultra 550B A55B returns text.

### How do I use Nemotron 3 Ultra 550B A55B through an API?

Nemotron 3 Ultra 550B A55B is available through the Impossibl API. Use nvidia/nemotron-3-ultra-550b-a55b as the model identifier with /v1/chat/completions. The quickstart explains how to create an API key and send your first request.

## Sources and related links

Prices, limits, capabilities, endpoints, and availability come from the public Impossibl catalog. Creator documentation supplies descriptions and release dates where available; it does not override the published API offering.

- [Public Impossibl model catalog](<https://api.impossibl.com/v1/models>)
- [Full model directory as JSON](<https://impossibl.com/models.json>)
- [Nemotron 3 Ultra 550B A55B: model details and pricing](<https://impossibl.com/nvidia/nemotron-3-ultra-550b-a55b>)
- [Model directory](<https://impossibl.com/models>)
- [NVIDIA model card](<https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4>)
- [Fireworks model documentation](<https://fireworks.ai/models/fireworks/nemotron-3-ultra-nvfp4>)
