Step 3.7 Flash (free)
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning. Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.
Offers
Same model, different listings. First-party, router, and aggregator rows stay separate so you can see the spread instead of a single blended number. Free/$0 listings are hidden by default (3 omitted). Show free/$0 offers.
No priced offers in the current snapshot.
Capabilities
- Reasoning
- Tool calling
- Attachments
- Open weights
- In: text
- In: image
- In: video
- Out: text
Scores
Sparse coverage on purpose. Missing a score means we have not linked one yet, not that the model is unranked. Full boards on /benchmarks.
- AA intelligence30.3
- AA coding39.6
- AA agentic21.5
Sources: Arena AI, Artificial Analysis. See /benchmarks.