UniVideo represents a significant advancement in AI video generation, merging video understanding, generation, and editing into a unified workflow. By utilizing a dual-stream architecture, it combines the reasoning capabilities of Multimodal Large Language Models (MLLM) with the generative power of Multimodal Diffusion Transformers (MMDiT). This innovative approach allows for deep semantic understanding of user instructions, enabling complex tasks such as object replacement, style transfer, and consistent character editing.