Image input
Models that can see
Section titled “Models that can see”| Provider | Vision |
|---|---|
| Anthropic Claude | all listed models (👁️ in the picker) |
| Google Gemini | all listed and discovered models |
| Ollama | models whose server details declare vision; older servers: recognised by family name (llava, moondream, minicpm-v, llama3.2-vision, gemma3, qwen2.5-vl, pixtral, granite-vision) |
| OpenRouter | models whose listing declares image among its input modalities |
| OpenAI, Groq, DeepSeek, Mistral, xAI, LM Studio, llama.cpp / vLLM, custom | recognised by family name: GPT-4o, GPT-4.1, GPT-5, o3, o4-mini, Pixtral, LLaVA, Qwen-VL, Gemma 3, Llama 4, Grok 4, and any model with vision in its name. The list is conservative: a vision model it misses still works, without images. |
For a model without vision Sirius strips the images and tells the model [N image(s) attached — this model cannot view images] instead of failing the request.
How images reach the model
Section titled “How images reach the model”Each provider gets images in its own format — Anthropic image blocks, Gemini inline data, OpenAI-style data URIs, Ollama’s base64 list. Images arrive three ways:
- Pasted or dropped into the chat box. The image goes to the model with your message.
- An image file attached from the workspace (PNG, JPEG, GIF or WebP, up to 5 MB) — the model receives the picture, not its path.
- Tool output. When the agent takes a screenshot with the
integrated browser’s
screenshot_page, the image goes back to a vision model as part of the tool result, and the model reasons about what it sees.
No image generation is exposed: the Gemini image models exist in the catalog but are hidden from every picker and no command uses them.