一种结合文本图像模型的AI视频输出系统及方法

By combining text and image models, the AI ​​video output system solves the problems of style drift and audio-visual asynchrony in multi-camera and multi-view video scenes, realizes the dissemination and maintenance of artistic style across multiple perspectives, ensures the consistency of structure and style of generated content, and improves the automation of video generation and the efficiency of audience comprehension.

CN121442166BActive Publication Date: 2026-07-17BEIJING SHUYOU WENLV TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SHUYOU WENLV TECH CO LTD
Filing Date
2025-11-03
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing illustration and AI video generation methods suffer from problems such as style drift, structural inconsistency, mismatch between script semantics and visuals, and audio-visual asynchrony in multi-camera and multi-perspective video scenes, especially when it is difficult to balance the need to preserve the hand-drawn art style and the structural accuracy of physical references.

Method used

An AI video output system combining text and image models is adopted. Through techniques such as perspective-aware style propagation, hand-drawn-object dual encoding, cross-modal alignment, perspective transformation and multi-perspective generation, script analysis and storyboard planning, audio synthesis and lip-syncing, the system can achieve the propagation and maintenance of artistic style across multiple perspectives, ensure the consistency of the structure and style of the generated content, and achieve audio-visual synchronization.

Benefits of technology

It significantly reduces style drift and distortion inconsistencies under multiple perspectives or lenses, improves content consistency and visual professionalism, ensures that generated content follows the geometric structure of real objects and hand-drawn style, and achieves automated script-to-screen correspondence and audio-visual synchronization, making it suitable for product display, character design and other scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121442166B_ABST
    Figure CN121442166B_ABST
Patent Text Reader

Abstract

本发明公开了一种结合文本图像模型的AI视频输出系统及方法,涉及AI视频输出技术领域,所述系统包括:视角变换与多视角生成模块:包括视角编码器和基于语义‑风格隐表示及视角向量的多视角生成器,用于生成针对不同摄像机参数的多角度插画序列,并在生成过程中通过时域一致性约束保持不同视角之间的结构与风格一致性。本发明通过视角感知风格传播,实现艺术风格按照三维视角变换规则在多帧、多视角间传播与保持,且生成器在采样时受视角条件约束输出一致的多角度插画序列,克服多视角或多镜头下的风格漂移与形变不一致问题,避免因视角变化导致笔触、纹理与主体结构不连贯,显著减少后期人工校正,提升内容一致性与观感专业度。
Need to check novelty before this filing date? Find Prior Art