Catalog data · Updated July 24, 2026

Vision LLM Models

Find active text-generation model identities that accept image input. Compare providers, context windows, known token prices, and related capabilities.

Objective catalog fields No overall quality score
335
Vision model identities
25
Providers shown
1
Data sections
6,154
Catalog records

Vision models with image input

Active text-generation model identities that accept image input.

Showing 50 of 335

Model Provider offers Why listed Context Input / output price Updated
OpenAI: o3
openai/o3
Github models, Nearai, Openrouter
3 offers
Image input and text output 200,000
$0.00 / $0.00
per 1M tokens, when known
April 16, 2025
OpenAI: o3 Deep Research
o3-deep-research
Openrouter
1 offer
Image input and text output 200,000
$10.00 / $40.00
per 1M tokens, when known
June 26, 2024
OpenAI: o3 Pro
o3-pro
Openai, Openrouter
2 offers
Image input and text output 200,000
$20.00 / $80.00
per 1M tokens, when known
June 10, 2025
OpenAI: o4 Mini
o4-mini
Github models, Nearai, Openrouter
3 offers
Image input and text output 200,000
$0.00 / $0.00
per 1M tokens, when known
April 16, 2025
OpenAI: o4 Mini Deep Research
o4-mini-deep-research
Openrouter
1 offer
Image input and text output 200,000
$2.00 / $8.00
per 1M tokens, when known
June 26, 2024
OpenAI: o4 Mini High
o4-mini-high
Openrouter
1 offer
Image input and text output 200,000
$1.10 / $4.40
per 1M tokens, when known
April 16, 2025
Perceptron: Perceptron Mk1
perceptron-mk1
Openrouter
1 offer
Image input and text output 32,768
$0.15 / $1.50
per 1M tokens, when known
May 12, 2026
Perplexity: Sonar
perplexity/sonar
Openrouter
1 offer
Image input and text output 127,072
$1.00 / $1.00
per 1M tokens, when known
January 27, 2025
Perplexity: Sonar Pro
sonar-pro
Openrouter, Perplexity
2 offers
Image input and text output 200,000
$3.00 / $15.00
per 1M tokens, when known
September 1, 2025
Perplexity: Sonar Pro Search
sonar-pro-search
Openrouter
1 offer
Image input and text output 200,000
$3.00 / $15.00
per 1M tokens, when known
October 30, 2025
Perplexity: Sonar Reasoning Pro
sonar-reasoning-pro
Openrouter, Perplexity
2 offers
Image input and text output 128,000
$2.00 / $8.00
per 1M tokens, when known
September 1, 2025
Phi-3.5-vision instruct (128k)
microsoft/phi-3.5-vision-instruct
Github models
1 offer
Image input and text output 128,000
$0.00 / $0.00
per 1M tokens, when known
August 20, 2024
Phi-4-multimodal-instruct
phi-4-multimodal-instruct
Github models
1 offer
Image input and text output 128,000
$0.00 / $0.00
per 1M tokens, when known
December 11, 2024
Pixtral 12B
pixtral-12b
Mistral
1 offer
Image input and text output 128,000
$0.15 / $0.15
per 1M tokens, when known
September 1, 2024
Pixtral Large (latest)
pixtral-large-latest
Mistral
1 offer
Image input and text output 128,000
$2.00 / $6.00
per 1M tokens, when known
November 4, 2024
QVQ Max
qvq-max
Alibaba
1 offer
Image input and text output 131,072
$1.20 / $4.80
per 1M tokens, when known
March 25, 2025
Qwen 3.5 35B A3B
qwen3-5-35b-a3b
Venice
1 offer
Image input and text output 256,000
$0.31 / $1.25
per 1M tokens, when known
June 11, 2026
Qwen 3.5 397B
qwen3-5-397b-a17b
Venice
1 offer
Image input and text output 128,000
$0.75 / $4.50
per 1M tokens, when known
June 11, 2026
Qwen 3.5 9B
qwen3-5-9b
Venice
1 offer
Image input and text output 256,000
$0.10 / $0.15
per 1M tokens, when known
June 11, 2026
Qwen 3.6 27B
qwen3-6-27b
Venice
1 offer
Image input and text output 256,000
$0.33 / $3.25
per 1M tokens, when known
June 11, 2026
Qwen 3.6 35B A3B
qwen3-6-35b-a3b
Venice
1 offer
Image input and text output 256,000
$0.15 / $1.00
per 1M tokens, when known
July 22, 2026
Qwen 3.6 Plus Uncensored
qwen-3-6-plus
Venice
1 offer
Image input and text output 1,000,000
$0.63 / $3.75
per 1M tokens, when known
June 11, 2026
Qwen 3.7 Max
qwen-3-7-max
Venice
1 offer
Image input and text output 1,000,000
$2.70 / $8.05
per 1M tokens, when known
June 11, 2026
Qwen 3.7 Plus
accounts/fireworks/models/qwen3p7-plus
Fireworks ai
1 offer
Image input and text output 262,144
$0.40 / $1.60
per 1M tokens, when known
June 12, 2026
Qwen 3.7 Plus
qwen-3-7-plus
Venice
1 offer
Image input and text output 1,000,000
$0.50 / $2.00
per 1M tokens, when known
June 11, 2026
Qwen-Omni Turbo
qwen-omni-turbo
Alibaba
1 offer
Image input and text output 32,768
$0.07 / $0.27
per 1M tokens, when known
March 26, 2025
Qwen-VL Max
qwen-vl-max
Alibaba
1 offer
Image input and text output 131,072
$0.80 / $3.20
per 1M tokens, when known
August 13, 2025
Qwen-VL OCR
qwen-vl-ocr
Alibaba
1 offer
Image input and text output 34,096
$0.72 / $0.72
per 1M tokens, when known
April 13, 2025
Qwen-VL Plus
qwen-vl-plus
Alibaba
1 offer
Image input and text output 131,072
$0.21 / $0.63
per 1M tokens, when known
August 15, 2025
qwen/qwen3.5-flash
qwen3.5-flash
Zenmux
1 offer
Image input and text output 1,020,000
$0.10 / $0.40
per 1M tokens, when known
March 20, 2026
Qwen2.5-Omni 7B
qwen2-5-omni-7b
Alibaba
1 offer
Image input and text output 32,768
$0.10 / $0.40
per 1M tokens, when known
unknown
Qwen2.5-VL 72B Instruct
qwen2-5-vl-72b-instruct
Alibaba
1 offer
Image input and text output 131,072
$2.80 / $8.40
per 1M tokens, when known
unknown
Qwen2.5-VL 7B Instruct
qwen2-5-vl-7b-instruct
Alibaba
1 offer
Image input and text output 131,072
$0.35 / $1.05
per 1M tokens, when known
unknown
Qwen3-Omni Flash
qwen3-omni-flash
Alibaba
1 offer
Image input and text output 65,536
$0.43 / $1.66
per 1M tokens, when known
September 15, 2025
Qwen3-VL 235B-A22B
qwen3-vl-235b-a22b
Alibaba, Venice
2 offers
Image input and text output 131,072
$0.21 / $1.90
per 1M tokens, when known
June 11, 2026
Qwen3-VL 30B-A3B
qwen3-vl-30b-a3b
Alibaba
1 offer
Image input and text output 131,072
$0.20 / $0.80
per 1M tokens, when known
unknown
Qwen3-VL 30B-A3B Instruct
Qwen3-VL-30B-A3B-Instruct
Nearai
1 offer
Image input and text output 256,000
$0.15 / $0.55
per 1M tokens, when known
September 23, 2025
Qwen3-VL Plus
qwen3-vl-plus
Alibaba
1 offer
Image input and text output 262,144
$0.20 / $1.60
per 1M tokens, when known
September 23, 2025
Qwen3.5 397B A17B
Qwen3.5-397B-A17B
Togetherai
1 offer
Image input and text output 262,144
$0.60 / $3.60
per 1M tokens, when known
June 15, 2026
Qwen3.5 9B
Qwen3.5-9B
Togetherai
1 offer
Image input and text output 262,144
$0.17 / $0.25
per 1M tokens, when known
March 3, 2026
Qwen3.5 Plus
qwen3.5-plus
Alibaba, Opencode, Zenmux
3 offers
Image input and text output 1,000,000
$0.20 / $1.20
per 1M tokens, when known
March 20, 2026
qwen3.5:397b
qwen3.5:397b
Ollama cloud
1 offer
Image input and text output 262,144
N/A / N/A
per 1M tokens, when known
February 17, 2026
Qwen3.6 Plus
qwen3.6-plus
Alibaba, Opencode, Openrouter
3 offers
Image input and text output 1,000,000
$0.33 / $1.95
per 1M tokens, when known
April 2, 2026
Qwen: Qwen2.5 VL 72B Instruct
qwen2.5-vl-72b-instruct
Openrouter
1 offer
Image input and text output 128,000
$0.80 / $1.00
per 1M tokens, when known
February 1, 2025
Qwen: Qwen3 VL 235B A22B Instruct
qwen3-vl-235b-a22b-instruct
Openrouter
1 offer
Image input and text output 262,144
$0.21 / $1.90
per 1M tokens, when known
September 23, 2025
Qwen: Qwen3 VL 235B A22B Thinking
qwen3-vl-235b-a22b-thinking
Openrouter
1 offer
Image input and text output 131,072
$0.26 / $2.60
per 1M tokens, when known
September 23, 2025
Qwen: Qwen3 VL 30B A3B Instruct
qwen3-vl-30b-a3b-instruct
Openrouter
1 offer
Image input and text output 262,144
$0.15 / $0.60
per 1M tokens, when known
October 6, 2025
Qwen: Qwen3 VL 30B A3B Thinking
qwen3-vl-30b-a3b-thinking
Openrouter
1 offer
Image input and text output 262,144
$0.13 / $1.56
per 1M tokens, when known
October 6, 2025
Qwen: Qwen3 VL 32B Instruct
qwen3-vl-32b-instruct
Openrouter
1 offer
Image input and text output 131,072
$0.10 / $0.42
per 1M tokens, when known
October 23, 2025
Qwen: Qwen3 VL 8B Instruct
qwen3-vl-8b-instruct
Openrouter
1 offer
Image input and text output 262,144
$0.12 / $0.46
per 1M tokens, when known
October 14, 2025

How the vision list works

The page groups active text-generation offers that list image input and text output into conservative model identities.

Inclusion criteria

  • The provider offer has explicit typed text-generation support.
  • The input modality set includes image and the output modality set includes text.

Exclusions

  • The list excludes image-generation-only records that do not return text.
  • The list excludes catalog-only, deprecated, retired, and disallowed offers.

Data rules

  • Vision meaning: Vision means that the catalog lists image as an input modality. It does not mean image generation.
  • Grouping: Provider offers are grouped by the conservative model identity rule used by the LLM models directory.

Limits of the data

  • Image input metadata does not measure visual accuracy, supported image size, or document quality.
  • A missing modality can mean that the source does not have a confirmed value.

What vision means here

This page uses the catalog modality fields. A model is present when an active text offer accepts images and returns text. The list does not claim that every image model has the same level of visual reasoning.

Sources

  • llm_db — Source for typed execution, modality, provider, context, and price data. Checked 2026-07-30.