Best Input Price
$0.22 / 1M
Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
Context Window
128,000 tokens
Reasoning
Supported
Tool Calling
Supported
Released
2024-11-09