Skip to content

Image input

Provider Vision
Anthropic Claude all listed models (👁️ in the picker)
Google Gemini all listed and discovered models
Ollama models whose server details declare vision; older servers: recognised by family name (llava, moondream, minicpm-v, llama3.2-vision, gemma3, qwen2.5-vl, pixtral, granite-vision)
OpenRouter models whose listing declares image among its input modalities
OpenAI, Groq, DeepSeek, Mistral, xAI, LM Studio, llama.cpp / vLLM, custom recognised by family name: GPT-4o, GPT-4.1, GPT-5, o3, o4-mini, Pixtral, LLaVA, Qwen-VL, Gemma 3, Llama 4, Grok 4, and any model with vision in its name. The list is conservative: a vision model it misses still works, without images.

For a model without vision Sirius strips the images and tells the model [N image(s) attached — this model cannot view images] instead of failing the request.

Each provider gets images in its own format — Anthropic image blocks, Gemini inline data, OpenAI-style data URIs, Ollama’s base64 list. Images arrive three ways:

  • Pasted or dropped into the chat box. The image goes to the model with your message.
  • An image file attached from the workspace (PNG, JPEG, GIF or WebP, up to 5 MB) — the model receives the picture, not its path.
  • Tool output. When the agent takes a screenshot with the integrated browser’s screenshot_page, the image goes back to a vision model as part of the tool result, and the model reasons about what it sees.

No image generation is exposed: the Gemini image models exist in the catalog but are hidden from every picker and no command uses them.