
Vidu
A Chinese text-to-video model chasing Sora and Kling on quality, at a lower price.
What is Vidu?
Vidu is a text-to-video and image-to-video generator built by Shengshu AI, a Chinese startup with roots in Tsinghua University research. It turns short prompts or reference images into a few seconds of video with reasonably consistent characters and motion, and it's positioned as one of China's clearest answers to Sora, Kling, and Veo.
Vidu entered the text-to-video race early and has kept iterating quickly, adding longer clips, better character consistency, and reference-image controls with each new version. It sits in a crowded field of Chinese video models racing Western labs on the same underlying problem — turning a sentence into believable motion — and it gets compared most often to Kling, since the two share a similar audience and price point.
Key features
- Text-to-video and image-to-video generation
- Character and scene consistency across shots
- Reference-image control for style and subject
- Multiple aspect ratios for social formats
- Camera-motion controls for more directed shots
- Fast generation times compared to some Western rivals
Pros
- Output quality genuinely competes with bigger-name rivals
- Ships new versions and improvements quickly
- Cheaper than most Western text-to-video tools
Cons
- English-language documentation and support lag behind the product
- Clip lengths are still short compared to what most marketers want
- Motion and physics still glitch on complex or crowded prompts

