Step 5 Preview API
Coming SoonStepFun's flagship 2026 multimodal MoE model for agentic work, coding, long-context reasoning, document analysis, and video understanding.
Step 5 Preview API Background
Overview
Step 5 Preview is StepFun's flagship 2026 model for agentic work, exposed as the Step 5 Preview API under the model ID step-5-preview. It is designed for demanding software engineering, long-context reasoning, multimodal analysis, and professional knowledge tasks, with notable strength in finance-oriented workflows. The model accepts text, image, and video inputs and returns text output, making it suitable for API-driven assistants that must read large document sets, inspect visual materials, call tools, and maintain progress across complex multi-step tasks.
Development History
Released in 2026, Step 5 Preview represents StepFun's flagship preview model for advanced agentic applications and production-oriented knowledge work. The Step 5 Preview API was positioned around frontier-level performance in software engineering and professional reasoning, especially where long context and tool use matter. Its launch emphasized a combination of sparse Mixture-of-Experts scaling, a 1,048,576-token context window, multimodal input support, and structured API features such as streaming, JSON outputs, JSON Schema, prompt caching, and configurable reasoning effort for complex workflows.
Key Innovations
- Sparse Mixture-of-Experts architecture with 600 billion total parameters and only 27 billion activated per token, improving scale-efficiency for advanced reasoning and coding tasks
- Native multimodal ingestion across text, images, and video within the Step 5 Preview API, enabling unified handling of documents, screenshots, charts, and short video content
- A 1 million token context window paired with tool calling, structured output, and reasoning effort controls for long-running agent workflows and large knowledge synthesis
Step 5 Preview API Technical Specifications
Architecture
The model uses a sparse Mixture-of-Experts architecture optimized for large-scale agentic workloads. Step 5 Preview has 600 billion total parameters, with 27 billion parameters activated per token, balancing high-capacity reasoning with more efficient per-token computation than a dense model of similar total size. The Step 5 Preview API supports text, image, and video input natively and produces text output, while also exposing streaming, tool calling, JSON Mode, JSON Schema output, prompt caching, and selectable reasoning effort levels of low, medium, and high.
Parameters
Step 5 Preview operates at very large scale with 600 billion total parameters and 27 billion active parameters per token. It supports a context window of 1,048,576 tokens, which is substantial enough for multi-document analysis, large codebases, extensive tool traces, and continuous agent execution. For multimodal requests, the Step 5 Preview API supports up to 60 images per request in JPG, JPEG, PNG, WebP, and static GIF formats, plus video input in MP4, QuickTime, and Matroska, with individual MP4 files under 128 MB and under 5 minutes recommended.
Capabilities
- Long-context understanding and reasoning across very large documents, multiple sources, tool outputs, and ongoing task state
- Programming and software engineering assistance across multiple languages, including bug localization, code modification, and test recommendation
- Multi-step agent tasks using tools to retrieve information, process documents, and generate structured outputs through the Step 5 Preview API
- Multimodal understanding for screenshots, charts, image-based question answering, and short video summarization or interpretation
Limitations
- Output is text-only, so the Step 5 Preview API can analyze images and video but does not generate image, audio, or video responses
- Video handling is best suited to relatively short files, with MP4 files under 128 MB and under 5 minutes recommended, which constrains long-form video workflows
Step 5 Preview API Performance
Strengths
- Frontier-level performance in software engineering and professional knowledge work, especially for long-context and tool-augmented tasks
- Particularly strong results in finance-related reasoning and analysis, making the Step 5 Preview API well suited to document-heavy professional domains
- Strong multimodal comprehension that combines textual reasoning with chart reading, screenshot interpretation, and short video understanding
- Reliable support for structured and agentic workloads through streaming, JSON outputs, reasoning controls, and prompt caching
Real-world Effectiveness
In practical use, Step 5 Preview is most effective when an application needs to combine large context ingestion, multi-step reasoning, and API-native workflow control. The Step 5 Preview API is well matched to assistants that must read lengthy reports, traverse code repositories, inspect screenshots, summarize short videos, and call tools without losing task continuity. Its long context window reduces fragmentation across requests, while structured output and reasoning effort controls improve integration into enterprise systems that require predictable formatting, orchestration, and auditable outputs.
Step 5 Preview API When to Use
Scenarios
- You have a software engineering team that needs help navigating a large repository, identifying regressions, proposing code changes, and recommending targeted tests. The Step 5 Preview API is ideal because it combines strong coding performance with a 1 million token context window, so it can reason over extensive code, issue descriptions, and tool outputs in one workflow. This can reduce manual triage time, improve developer productivity, and support more consistent debugging and remediation across multi-file changes.
- You have a research, legal, or financial operations workflow that requires reading many long documents, comparing sources, extracting conclusions, and returning structured results. The Step 5 Preview API fits this scenario because it is optimized for long-context reasoning and professional knowledge work, with particular strength in finance. It can process large bodies of material in fewer handoffs, preserve cross-document relationships, and deliver JSON-formatted outputs that are easier to validate, automate, and integrate into downstream business systems.
- You have a multimodal support or analytics task involving screenshots, charts, product images, and short videos alongside written instructions or questions. The Step 5 Preview API is a strong choice because it natively accepts text, image, and video input in one model, allowing a single workflow to interpret visual evidence and explain findings in text. This improves operational efficiency for use cases such as dashboard review, chart analysis, screenshot troubleshooting, and concise video summarization without building separate modality-specific pipelines.
Best Practices
- Use the Step 5 Preview API with structured output modes such as JSON Mode or JSON Schema when integrating with business workflows, analytics pipelines, or tool-based agents that require predictable machine-readable responses
- Select reasoning effort levels based on task complexity and combine prompt caching, tool calling, and long-context inputs to improve consistency and efficiency in repeated or multi-step workflows
- Provide clear task framing, relevant documents, and modality-specific inputs together so the Step 5 Preview API can reason across code, text, images, and short videos in a single coherent context