StepFun, an AI company based in Shanghai, China, develops multimodal AI models that process text, image, video, and audio data. Its offerings focus on providing advanced large language models (LLMs) and multimodal AI capabilities primarily through an API, catering to high-frequency AI developers and professionals. The platform’s designed for those building AI agents and integrating with coding tools.
StepFun: Chinese AI Platform with Multimodal Models
StepFun’s core functionality revolves around its API-accessible models. These include reasoning models capable of complex problem-solving, mathematical reasoning, and code generation. It also provides audio models for text-to-speech conversion and voice cloning. Beyond foundational models, StepFun offers specialized tools:
- Stepfun Diligence Check: An AI-powered search tool that includes agent-verified citations.
- Deep Research Agent: An autonomous agent designed to conduct multi-step research, encompassing web browsing, data analysis, and report generation.
These tools are meant to support professionals — they’re built for in fields such as finance, consulting, healthcare, and research who require in-depth reports and insights. Industries like content creation, manufacturing, intelligent cars, game entertainment, and government can also utilize StepFun for various applications, from smart editing to legal consulting.
Step-1V Vision-Language Model and Step-2 Reasoning
StepFun’s models, such as Step 3.5 Flash, use a Sparse Mixture-of-Experts (MoE) architecture. This design activates only a subset of parameters per token, so you’re getting efficient processing. Key technical details include:
- Inference Speed: Step 3.5 Flash can achieve speeds of up to 350 tokens per second.
- Context Window: Step 3.5 Flash supports a substantial 262,144 token context window, so it can handle extensive inputs.
- Architecture: Models are built for tool calling and agentic workflows.
- API Access: Rate limiting applies to API usage, based on requests per minute (RPM), tokens per minute (TPM), and concurrency.
Free Tier Available, Pay-per-Token API
Access to StepFun’s models is primarily via API, with pricing based on token usage. Multimodal models also incur costs for image processing.
| Plan | Input Tokens (per 1M) | Output Tokens (per 1M) | Key Details |
|---|---|---|---|
| Step 1 (32K) | $2.05 | $9.59 | |
| Step 3.5 Flash | $0.09 | $0.30 | Images billed at 400 tokens per image |
StepFun also offers a "Step Plan" subscription service tailored for high-frequency AI developers. These plans provide varying prompt limits and access to flagship models. Examples include Flash Mini at $6.99 and Flash Max at $99.
Strong Chinese-Language Performance and Government Partnerships
StepFun’s Step 3.5 Flash model is optimized for AI agent workflows, offering a combination of efficiency, strong reasoning capabilities, and low-latency responses. That makes it a solid pick for complex, multi-step tasks. The "Deep Research" agent aims to automate thorough research, from web searching and data analysis to report generation, which can save knowledge workers a lot of time — you won’t need to do it all manually. The "Step Plan" subscriptions are marketed as providing "2x the usage of comparable competitor tiers," suggesting a cost-effective option for intensive API usage.
Chinese Regulatory Compliance, Limited English Documentation
A notable limitation is the upcoming cessation of operations for the "StepFun AI (International version)" on May 27, 2026. This will result in the deletion of chat history for users of that specific version. Some users have also expressed concerns regarding the output quality of certain StepFun tools; for instance, one reviewer found the "Stepfun Diligence Check" output to be comparable to other services. Additionally, the Step 3.5 Flash model has been observed to sometimes "overthink" or be "a bit too passive/narrating too much" in interactive scenarios. There have also been discussions questioning the accuracy of self-reported benchmarks for some models, so you’d be wise to verify performance against their specific needs.


