Catalog data · Updated July 24, 2026

Vision LLM Models

Find active text-generation model identities that accept image input. Compare providers, context windows, known token prices, and related capabilities.

Objective catalog fields No overall quality score
335
Vision model identities
25
Providers shown
1
Data sections
6,154
Catalog records

Vision models with image input

Active text-generation model identities that accept image input.

Showing 35 of 335

Model Provider offers Why listed Context Input / output price Updated
Qwen: Qwen3 VL 8B Thinking
qwen3-vl-8b-thinking
Openrouter
1 offer
Image input and text output 131,072
$0.12 / $1.37
per 1M tokens, when known
October 14, 2025
Qwen: Qwen3.5 397B A17B
qwen3.5-397b-a17b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.39 / $2.34
per 1M tokens, when known
February 15, 2026
Qwen: Qwen3.5 Plus 2026-02-15
qwen/qwen3.5-plus-02-15
Openrouter
1 offer
Image input and text output 1,000,000
$0.26 / $1.56
per 1M tokens, when known
February 16, 2026
Qwen: Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420
Openrouter
1 offer
Image input and text output 1,000,000
$0.30 / $1.80
per 1M tokens, when known
April 27, 2026
Qwen: Qwen3.5-122B-A10B
qwen3.5-122b-a10b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.26 / $2.08
per 1M tokens, when known
February 23, 2026
Qwen: Qwen3.5-27B
qwen3.5-27b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.20 / $1.56
per 1M tokens, when known
February 23, 2026
Qwen: Qwen3.5-35B-A3B
qwen3.5-35b-a3b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.14 / $1.00
per 1M tokens, when known
February 23, 2026
Qwen: Qwen3.5-9B
qwen3.5-9b
Openrouter
1 offer
Image input and text output 262,144
$0.10 / $0.15
per 1M tokens, when known
February 23, 2026
Qwen: Qwen3.5-Flash
qwen/qwen3.5-flash-02-23
Openrouter
1 offer
Image input and text output 1,000,000
$0.07 / $0.26
per 1M tokens, when known
February 25, 2026
Qwen: Qwen3.6 27B
qwen3.6-27b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.29 / $2.40
per 1M tokens, when known
April 22, 2026
Qwen: Qwen3.6 35B A3B
qwen3.6-35b-a3b
Alibaba, Openrouter
2 offers
Image input and text output 262,144
$0.14 / $1.00
per 1M tokens, when known
April 17, 2026
Qwen: Qwen3.6 Flash
qwen3.6-flash
Alibaba, Openrouter
2 offers
Image input and text output 1,000,000
$0.19 / $1.13
per 1M tokens, when known
April 27, 2026
Qwen: Qwen3.7 Plus
qwen3.7-plus
Alibaba, Openrouter, Zenmux
3 offers
Image input and text output 1,000,000
$0.32 / $1.28
per 1M tokens, when known
June 4, 2026
Reka Edge
reka-edge
Openrouter
1 offer
Image input and text output 16,384
$0.10 / $0.10
per 1M tokens, when known
March 20, 2026
Sakana: Fugu Ultra
fugu-ultra
Openrouter
1 offer
Image input and text output 1,000,000
$5.00 / $30.00
per 1M tokens, when known
June 24, 2026
Seed 2.1 Turbo
seed-2-1-turbo
Venice
1 offer
Image input and text output 256,000
$0.63 / $3.13
per 1M tokens, when known
July 24, 2026
Step 3.7 Flash (Free)
step-3.7-flash-free
Zenmux
1 offer
Image input and text output 256,000
$0.00 / $0.00
per 1M tokens, when known
May 29, 2026
Step-3
step-3
Zenmux
1 offer
Image input and text output 65,536
$0.21 / $0.57
per 1M tokens, when known
July 31, 2025
StepFun: Step 3.7 Flash
step-3.7-flash
Openrouter, Zenmux
2 offers
Image input and text output 262,144
$0.20 / $1.15
per 1M tokens, when known
May 29, 2026
Thinking Machines: Inkling
thinkingmachines/inkling
Openrouter
1 offer
Image input and text output 1,048,576
$1.00 / $4.05
per 1M tokens, when known
July 15, 2026
Venice Role Play Uncensored
venice-uncensored-role-play
Venice
1 offer
Image input and text output 128,000
$0.50 / $2.00
per 1M tokens, when known
June 11, 2026
Venice Uncensored 1.2
venice-uncensored-1-2
Venice
1 offer
Image input and text output 128,000
$0.20 / $0.90
per 1M tokens, when known
June 11, 2026
x-ai/grok-4.2-fast
grok-4.2-fast
Zenmux
1 offer
Image input and text output 2,000,000
$3.00 / $9.00
per 1M tokens, when known
March 20, 2026
x-ai/grok-4.2-fast-non-reasoning
grok-4.2-fast-non-reasoning
Zenmux
1 offer
Image input and text output 2,000,000
$3.00 / $9.00
per 1M tokens, when known
March 20, 2026
xAI: Grok 4.20
grok-4.20
Openrouter
1 offer
Image input and text output 2,000,000
$1.25 / $2.50
per 1M tokens, when known
March 31, 2026
xAI: Grok 4.20 Multi-Agent
grok-4.20-multi-agent
Openrouter
1 offer
Image input and text output 2,000,000
$1.25 / $2.50
per 1M tokens, when known
March 31, 2026
xAI: Grok 4.3
grok-4.3
Openrouter, Xai, Zenmux
3 offers
Image input and text output 1,000,000
$1.25 / $2.50
per 1M tokens, when known
May 7, 2026
xAI: Grok 4.5
grok-4.5
Openrouter, Xai, Zenmux
3 offers
Image input and text output 500,000
$2.00 / $6.00
per 1M tokens, when known
July 8, 2026
xAI: Grok Latest
grok-latest
Openrouter
1 offer
Image input and text output 500,000
$2.00 / $6.00
per 1M tokens, when known
July 8, 2026
Xiaomi: MiMo-V2.5
mimo-v2.5
Openrouter, Zenmux
2 offers
Image input and text output 1,050,000
$0.14 / $0.28
per 1M tokens, when known
April 22, 2026
z-ai/glm-4.6v-flash
glm-4.6v-flash
Zenmux
1 offer
Image input and text output 200,000
$0.02 / $0.21
per 1M tokens, when known
December 8, 2025
z-ai/glm-4.6v-flash-free
glm-4.6v-flash-free
Zenmux
1 offer
Image input and text output 200,000
$0.00 / $0.00
per 1M tokens, when known
December 8, 2025
Z.ai: GLM 4.5V
glm-4.5v
Openrouter, Zai
2 offers
Image input and text output 65,536
$0.60 / $1.80
per 1M tokens, when known
August 11, 2025
Z.ai: GLM 4.6V
glm-4.6v
Openrouter, Zai, Zenmux
3 offers
Image input and text output 200,000
$0.14 / $0.42
per 1M tokens, when known
December 8, 2025
Z.ai: GLM 5V Turbo
glm-5v-turbo
Openrouter, Zai, Zenmux
3 offers
Image input and text output 202,752
$0.73 / $3.19
per 1M tokens, when known
April 1, 2026

How the vision list works

The page groups active text-generation offers that list image input and text output into conservative model identities.

Inclusion criteria

  • The provider offer has explicit typed text-generation support.
  • The input modality set includes image and the output modality set includes text.

Exclusions

  • The list excludes image-generation-only records that do not return text.
  • The list excludes catalog-only, deprecated, retired, and disallowed offers.

Data rules

  • Vision meaning: Vision means that the catalog lists image as an input modality. It does not mean image generation.
  • Grouping: Provider offers are grouped by the conservative model identity rule used by the LLM models directory.

Limits of the data

  • Image input metadata does not measure visual accuracy, supported image size, or document quality.
  • A missing modality can mean that the source does not have a confirmed value.

What vision means here

This page uses the catalog modality fields. A model is present when an active text offer accepts images and returns text. The list does not claim that every image model has the same level of visual reasoning.

Sources

  • llm_db — Source for typed execution, modality, provider, context, and price data. Checked 2026-07-30.