Step 3.7 Flash

LLM
Stepfun

Step 3.7 Flash is a sparse MoE vision-language model built for agentic workflows combining perception, search, and reasoning. It supports a 256k context window and offers three selectable reasoning levels to balance speed and

Context tokens

262,144

Output tokens

8,192

Released

Jun 12, 2026

Schema

Supports three selectable reasoning levels (low, medium, and high) that let you balance speed and depth.

Schema documentation

Capabilities

Tool use
Structured output
Thinking
Multimodal

Supported input media

image
text
video

Supported tools

No tools enabled.

Pricing

TypeCreditsUnits
InputCredits per 1k tokens
OutputCredits per 1k tokens

Variants

No variants available for this model.

Try This Model

Write a prompt and experiment with Step 3.7 Flash in the model experiments page. You can compare it with other models side by side.