FFree
NemotronClosed Weights

nemotron-lightning-3.5-30b-a3b

nemotron-lightning-3.5-30b-a3b

Best Input Price

$0.05 / 1M

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

Context Window

262,144 tokens

Reasoning

Supported

Tool Calling

Supported

Released

2026-08-15

Inference Providers (1)

ProviderModel IDContextInput / 1MOutput / 1MAction
Requestynemotron-lightning-3.5-30b-a3b262,144$0.05$0.20Docs ↗