一种结合文本图像模型的AI视频输出系统及方法
By combining text and image models, the AI video output system solves the problems of style drift and audio-visual asynchrony in multi-camera and multi-view video scenes, realizes the dissemination and maintenance of artistic style across multiple perspectives, ensures the consistency of structure and style of generated content, and improves the automation of video generation and the efficiency of audience comprehension.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SHUYOU WENLV TECH CO LTD
- Filing Date
- 2025-11-03
- Publication Date
- 2026-07-17
AI Technical Summary
Existing illustration and AI video generation methods suffer from problems such as style drift, structural inconsistency, mismatch between script semantics and visuals, and audio-visual asynchrony in multi-camera and multi-perspective video scenes, especially when it is difficult to balance the need to preserve the hand-drawn art style and the structural accuracy of physical references.
An AI video output system combining text and image models is adopted. Through techniques such as perspective-aware style propagation, hand-drawn-object dual encoding, cross-modal alignment, perspective transformation and multi-perspective generation, script analysis and storyboard planning, audio synthesis and lip-syncing, the system can achieve the propagation and maintenance of artistic style across multiple perspectives, ensure the consistency of the structure and style of the generated content, and achieve audio-visual synchronization.
It significantly reduces style drift and distortion inconsistencies under multiple perspectives or lenses, improves content consistency and visual professionalism, ensures that generated content follows the geometric structure of real objects and hand-drawn style, and achieves automated script-to-screen correspondence and audio-visual synchronization, making it suitable for product display, character design and other scenarios.
Smart Images

Figure CN121442166B_ABST