Best Input Price
$0.00 / 1M
Vision language model for physical-world understanding with structured reasoning on video and images
Context Window
131,072 tokens
Reasoning
Supported
Tool Calling
Supported
Released
2025-12-01
Best Input Price
$0.00 / 1M
Vision language model for physical-world understanding with structured reasoning on video and images
131,072 tokens
Supported
Supported
2025-12-01