Meituan’s LongCat project introduces a powerful suite of AI models, offering advanced capabilities in language processing, video generation, and image creation. This collection of open-source tools, including LongCat-Flash-Chat, LongCat-Video, and LongCat-Image, caters primarily to developers and AI researchers seeking reliable, deployable AI solutions.
Flash-Chat (560B MoE), Video (5-Min), and Image in One Ecosystem
LongCat isn’t just one tool; it’s an ecosystem of specialized AI models:
- LongCat-Flash-Chat: This large language model (LLM) is built on a high-speed Mixture of Experts (MoE) architecture. It’s designed for rapid, multi-modal interactions, excelling in reasoning, coding, instruction following, and complex agentic tasks. It handles long documents and conversations with a 128K token context window.
- LongCat-Video: This model generates videos up to 5 minutes long from text prompts, images, or existing footage. It maintains visual consistency across frames and supports text-to-video, image-to-video, and video continuation within a unified architecture.
- LongCat-Image: Offering AI image generation and editing, this tool supports text-to-image, various editing tasks like object manipulation and style transfer, and multi-round editing. It’s particularly noted for its superior rendering of Chinese text.
Developers, Content Creators, and Chinese Text Rendering
LongCat’s open-source nature (MIT license) makes it a valuable resource for specific audiences:
- Developers and AI Researchers: The primary target, benefiting from the models’ deployability and API access. LongCat-Flash-Chat serves as an advanced "conversational pair programmer" or a tool for studying MoE models.
- Content Creators: LongCat-Video is ideal for social media, educational materials, marketing videos, and indie game development, providing consistent long-form video generation.
- Data Analysts: Can use LongCat-Flash-Chat for automating complex data interpretation.
- Production Teams: Use LongCat-Video for concept previews and animatics.
- Users needing precise image control: LongCat-Image offers high-quality generation and editing, especially for those working with Chinese text.
100+ Tokens/Second on H800 and 13.6B DiT Video Architecture
LongCat’s models are built on advanced architectures:
- LongCat-Flash-Chat: Employs a Mixture of Experts (MoE) architecture with 560 billion total parameters, dynamically activating about 27 billion per token. This allows for high-throughput inference, exceeding 100 tokens/second on H800 GPUs.
- LongCat-Video: Uses a 13.6 billion dense parameter Diffusion Transformer (DiT) architecture. It generates 720p videos at 30 frames per second, utilizing a coarse-to-fine generation strategy and Block Sparse Attention for efficiency and temporal consistency.
- LongCat-Image: Supports "One-Step Multi-Edit" instructions and a "Visual Guidance" workflow for precise control.
MIT License, $0.7/Million Output Tokens, and Credit Packs
LongCat’s core models are largely accessible without direct licensing fees:
- Open-Source Models: LongCat-Flash-Chat and LongCat-Video are open-source under the MIT license, allowing free commercial use, modification, and deployment.
- Hosted Services:
- LongCat-Flash-Chat API: Costs around $0.7 per 1 million output tokens when using H800 GPUs.
- LongCat Video: Offers credit packs for its hosted service, with one-time purchases such as $9.9 for 90 credits or $29.9 for 400 credits for video and avatar generation.
5-Minute Video Consistency and Unified Video Workflow
LongCat distinguishes itself in several ways:
- Long-Form Video Consistency: LongCat-Video excels at generating extended videos (up to 5 minutes) while maintaining high temporal consistency, visual coherence, and subject identity. This addresses a common weakness in many open-source video models.
- Efficient Agentic AI: LongCat-Flash-Chat’s MoE architecture provides efficient agentic AI capabilities with high processing speed and reduced operational costs, performing comparably to leading LLMs.
- Unified Video Workflow: LongCat-Video’s single architecture for text-to-video, image-to-video, and video continuation streamlines workflows, eliminating the need for multiple specialized tools.
- Open-Source Freedom: The MIT license for Flash-Chat and Video offers developers complete control, freedom from vendor lock-in, and the ability to integrate and modify models for commercial applications without licensing fees.
NVIDIA-Only, 25-35 Min Generation Time, and Evolving Docs
LongCat models carry specific requirements and limitations for deployment:
- GPU Incompatibility: LongCat-Video specifically requires NVIDIA GPUs with CUDA support and isn’t compatible with Apple Silicon (M1/M2/M3) chips.
- Video Generation Speed: A 4-minute video can take 25-35 minutes to generate on a single GPU, which may not suit applications requiring instant results.
- Prompt Sensitivity: LongCat-Video can be sensitive to prompt complexity, requiring detailed and structured prompts for optimal output.
- Quality Nuances: Motion realism in LongCat-Video may not match top proprietary systems, and fine details like hands can sometimes appear blurred or distorted. It’s also sensitive to unsupported resolutions.
- Image Artifacts: LongCat-Image may occasionally produce images with a reddish or yellowish tint.
- Deployment Complexity: The absence of official web demos for some models means users must self-deploy, which can be a barrier for non-technical individuals. Community support and documentation are still evolving due to the models’ recent release.


