AI video output system and method combined with text image model
By combining text and image models, the AI video output system solves the problems of style drift and audio-visual asynchrony in multi-camera and multi-view video scenes, realizes the dissemination and preservation of artistic style across multiple perspectives, and improves the structural consistency of video generation and the efficiency of audience comprehension.
Patent Information
- Application Number
- CN202511589099.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-11-03
AI Technical Summary
Existing illustration and AI video generation methods suffer from problems such as style drift, structural inconsistency, mismatch between script semantics and visuals, and audio-visual asynchrony in multi-camera and multi-perspective video scenes, especially when it is difficult to balance the need to preserve the hand-drawn art style and the structural accuracy of physical references.
An AI video output system combining text and image models is adopted. Through techniques such as perspective-aware style propagation, hand-drawn-object dual encoding, cross-modal alignment, perspective transformation and multi-perspective generation, script analysis and storyboard planning, audio synthesis and lip-syncing, the system can achieve the propagation and maintenance of artistic style across multiple perspectives, ensure the consistency of the structure and style of the generated content, and achieve audio-visual synchronization.
It significantly reduces style drift and structural inconsistency issues under multiple perspectives or lenses, improves content consistency and visual professionalism, simplifies the video production process, and enhances the controllability of audio-visual synchronization and script-to-screen automation, making it suitable for scenarios such as product demonstrations and character design.
Smart Images

Figure CN121442166A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI video output technology, specifically to an AI video output system and method that combines text and image models. Background Technology
[0002] Existing illustration and AI video generation methods often rely on a single modality or prioritize style and perspective control, leading to issues such as style drift, structural inconsistency, mismatch between script semantics and visuals, and audio-visual asynchrony in multi-camera and multi-perspective video scenes. In particular, when creators want to preserve the hand-drawn art style while also needing to combine the structural accuracy of real-world references, traditional single-channel generation methods struggle to balance structural fidelity and artistic consistency. In the video compositing stage, automatically generated subtitles, narration, and visuals are often semantically disconnected and difficult to synchronize with lip movements.
[0003] Patent CN119046481B discloses a multimedia video stream management system and method based on artificial intelligence. The patent enables adaptive monitoring of the performance of the AI module in the field of multimedia video stream tag processing.
[0004] The aforementioned patents effectively improve the discoverability of AI in multimedia video content and user experience. However, existing illustration and AI video generation rely heavily on a single modality or prioritize style and perspective control, leading to problems such as style drift, structural inconsistency, script semantics mismatch with visuals, and audio-visual asynchrony in multi-camera and multi-perspective video scenarios.
[0005] To this end, this application proposes an AI video output system and method that combines a text image model to achieve the propagation and preservation of artistic style across multiple frames and perspectives according to the three-dimensional perspective transformation rules, and to output a consistent multi-angle illustration sequence by the generator under the constraint of perspective conditions during sampling. Summary of the Invention
[0006] The purpose of this invention is to provide an AI video output system and method that combines text and image models to solve the problems mentioned in the background art, such as the difficulty in balancing structural fidelity and artistic consistency, and the semantic disconnect and difficulty in lip-syncing between automatically generated subtitles, narration and images in the video synthesis process.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an AI video output system and method combining text and image models, the system comprising:
[0008] Reference input module: used to receive physical reference images, hand-drawn reference images, hand-drawn style keywords, and descriptive text;
[0009] Dual encoder module: includes a physical encoder and a hand-drawn encoder;
[0010] Cross-modal alignment module: used to map geometric priors and style tokens from dual encoders into fused semantic-style latent representations through a multimodal alignment strategy;
[0011] The perspective transformation and multi-view generation module includes a perspective encoder and a multi-view generator based on semantic-style implicit representation and perspective vectors, which is used to generate multi-angle illustration sequences for different camera parameters, and maintains structural and stylistic consistency between different perspectives through temporal consistency constraints during the generation process.
[0012] Script parsing and storyboard planning module: used to parse the input introductory text or video script into shot units, and map the shot units into view vectors, generation parameters and timeline instructions;
[0013] Question and answer caption generation module: used to automatically generate question and answer data, captions or explanatory text corresponding to the semantics of each shot based on semantic-style implicit representation and shot unit.
[0014] Preferably, the system further includes:
[0015] Audio synthesis and lip-sync module: used to generate emotional synthesized speech based on subtitles or explanatory text and align the phoneme timeline of the generated speech with the target lip-sync template for time-speech-lip-sync.
[0016] Timeline Composition Module: Used to combine multi-angle illustration sequences, Q&A subtitles, and audio tracks into a video output file in chronological order according to the timeline instructions of the camera unit;
[0017] Evaluation and Optimization Module: This module evaluates the generated video for aspects such as viewpoint consistency, style consistency, semantic alignment, and audio-visual synchronization, and outputs optimization suggestions.
[0018] Preferably, the physical encoder includes sub-modules for semantic segmentation, dense depth estimation, and key point detection, and the hand-drawn encoder includes sub-modules for line drawing, brushstroke feature extraction, and sketch vectorization. The depth and key points output by the physical encoder are used to guide 3D layout and occlusion processing in the viewpoint transformation and multi-view generation modules.
[0019] Preferably, the cross-modal alignment module adopts a contrastive learning and cross-attention mechanism to map text style keywords, hand-drawn style tokens and visual structure priors to a unified semantic-style embedding space, and introduces a few-sample style fine-tuning strategy during training to support style transfer under a small number of hand-drawn samples.
[0020] The cross-modal alignment module works in collaboration with the perspective transformation and multi-view generation module to achieve the perceptual propagation of style across multiple perspectives and frames based on the temporal consistency constraints of depth and optical flow and the cross-frame attention mechanism, thereby ensuring the consistency of the structure and artistic style of the output video in shot transitions and shot movements.
[0021] Preferably, the viewpoint transformation and multi-viewpoint generation module adopts a conditional diffusion generation network or a hybrid NeRF and diffusion generation architecture. The viewpoint decoder can map the camera parameters parsed from the script into latent space transformation vectors, and introduces temporal loss based on depth consistency and optical flow constraints during the generation process to reduce cross-frame texture drift.
[0022] Preferably, the script parsing and storyboard planning module parses the introductory text into a set of shot descriptions based on a natural language processor, and automatically generates a scene map for each shot. The scene map includes the subject, background, props, actions, and emotion tags. The scene map is used to determine the composition, perspective, and shot motion parameters of the multi-view generator.
[0023] Preferably, the question-and-answer subtitle generation module includes a question-and-answer generator based on a generative large model and a subtitle style generator. When generating question-and-answer pairs, the question-and-answer generator uses semantic similarity filtering and clustering steps to ensure that the generated questions and answers are relevant to the prominent objects in the shot. The subtitle CCTV generator outputs subtitle files of different granularities and formats according to the target platform.
[0024] Preferably, the audio synthesis and lip-sync module includes: an emotion control TTS submodule, a phoneme timeline generator, and a lip-sync synthesizer based on keyframes or 2D skeletons, which supports controllable adjustment of synthesized speech through emotion parameters, speech rate parameters, and timbre parameters, and aligns the phoneme timeline with the mouth keyframes of the target character to achieve visual lip-sync.
[0025] The timeline compositing module supports incremental re-rendering strategies, optical flow-driven inter-frame interpolation, color matching and lens transition interpolation, and also supports exporting editable project files and differential caches to improve iteration efficiency.
[0026] Preferably, the method includes the following steps:
[0027] S1. Receive at least one physical reference image, at least one hand-drawn reference image, keywords for the hand-drawn style, and descriptive text;
[0028] S2. Encode the physical reference image and the hand-drawn reference image respectively to obtain geometric priors and style tokens;
[0029] S3. Obtain a fused semantic-style implicit representation by aligning geometric priors and style tokens across modalities;
[0030] S4. Parse the introductory text or script into a list of shot units, and generate camera parameters and timeline instructions for each shot;
[0031] S5. For each shot, based on the semantic-style implicit representation and camera parameters in S3, a view-aware multi-view generator is used to generate the corresponding illustration frame sequence, while applying view consistency and temporal consistency constraints during the generation process.
[0032] S6. Based on the content of S3 and S4, automatically generate question-and-answer pairs and subtitle texts corresponding to the semantics of each shot, and synthesize emotional language based on the subtitle texts;
[0033] S7. Combine the illustration sequence generated in S5, the subtitles generated in S6, and the audio into the final video output file according to the timeline instructions in S4.
[0034] S8. Evaluate the video for consistency of viewpoint, style, voice alignment, and audio-visual synchronization, and perform local re-rendering or parameter fine-tuning based on the evaluation results.
[0035] Preferably, the training and generation of the multi-view generator in S5 simultaneously employ the following loss or constraint:
[0036] Content-aware loss, style loss, perspective consistency loss, adversarial detail enhancement loss, and cross-modal semantic consistency loss are included. Furthermore, cross-frame attention and latent space smoothing strategies are employed during generation to maintain structural and stylistic coherence between different perspectives and consecutive frames. In the fine-tuning stage, a few hand-drawn samples are used for fine-tuning to adapt to the client's specified style.
[0037] Compared with the prior art, the beneficial effects of the present invention are:
[0038] 1. This invention achieves the propagation and maintenance of artistic style across multiple frames and perspectives according to the three-dimensional perspective transformation rules through perspective perception style propagation. Moreover, the generator outputs a consistent multi-angle illustration sequence under the constraint of perspective conditions during sampling, which overcomes the problems of style drift and deformation inconsistency under multiple perspectives or lenses, avoids the inconsistency of brushstrokes, textures and main structure due to perspective changes, significantly reduces post-processing manual correction, and improves content consistency and visual professionalism.
[0039] 2. This invention achieves the goal of preserving the characteristics of hand-drawn style while ensuring that the generated content follows the geometric structure and appearance constraints of the real object. It resolves the contradiction between style-first approach leading to structural distortion or structure-first approach leading to style loss. It can faithfully reflect the form and spatial relationship of the real object and present the artistic style specified by the user. It is suitable for scenarios such as product display and character setting that require accurate structure and stylized expression.
[0040] 3. This invention achieves automatic generation of corresponding camera parameters, composition constraints, and timeline instructions through script-scene graph mapping and automatic storyboarding, bridging the expression gap between script and screen parameters, avoiding the inefficient and error-prone operation of manually converting text intent into shot parameters, and ensuring a direct correspondence between storyboard semantics and generated screens. This significantly improves the automation and controllability of script-to-screen transition, enabling non-professional users to quickly generate storyboard-style video drafts from text, accelerating the content production process and facilitating subsequent manual fine-tuning.
[0041] 4. This invention achieves audio-visual lip-sync through multimodal question-and-answer, subtitle linkage, and emotional TTS lip-sync, solving the problems of semantic disconnect between automatically generated text and visuals, as well as the visual disjointedness caused by asynchrony between speech and lip movements. It reduces the workload of manually writing subtitles and manually adjusting lip movements, improves audience comprehension efficiency and immersive experience, and facilitates the generation of multi-purpose content for teaching, explanation, marketing, etc. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of the AI video output system framework and process of the present invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Please see Figure 1 This invention provides an embodiment of an AI video output system and method combining text and image models. Input types and formats include: physical reference images: JPEG / PNG, with a recommended resolution ≥1024×1024; optional depth maps or photos of the same object from multiple perspectives; hand-drawn reference images: scanned or photographed line drawings or sketches, supporting vector (SVG) and bitmap (PNG / TIFF); hand-drawn style keywords: natural language descriptions, such as "watercolor elegance," "pencil sketch," "high-saturation oil painting brushstrokes," "cyberpunk neon lines," etc.; introductory text / script: natural language paragraphs, possibly containing scene cues (such as close-ups, slow motion, character narration), duration indicators, or emotional tags.
[0045] Sketch processing: binarization or multi-level grayscale quantization; vectorization of hand-drawn lines, extraction of line width and line density statistics, and calculation of stroke direction field;
[0046] Real-world image processing: semantic segmentation (foreground, background, subject, props), dense depth estimation (monocular depth prediction or multi-view SFM results), key point detection (face, human body, object key points);
[0047] Standardization: color space unification (sRGB), resolution adjustment, edge cleanup, noise suppression;
[0048] Metadata generation: standard input source, collection time, user priority flags (e.g., style priority or structure priority);
[0049] Real-world encoder: Extracts geometric and appearance priors; Output: Dense depth map D(x,y), semantic segmentation label S(x,y), keypoints K={k1,...,kn}, color and texture statistical vector T; The Transformer-based visual encoder employs a multi-task learning strategy to share features and improve consistency; If multiple real-world images are provided, more accurate depth and surface normals can be obtained through SFM or multi-view consistency algorithms;
[0050] Hand-drawn encoder: Extracts stroke shape, line structure, and style tokens; Output: Contour image C(x,y), stroke vector set P (including line width, curvature, and direction), and style token set ST; Uses a lightweight network trained on line art, such as a line art feature extractor + image vectorization module; If the hand-drawn image is in vector format, path information can be directly parsed and pen pressure and pen width can be calculated; Tolerates incomplete errors in the sketch, and generates structural priors that can be used by subsequent fusion modules;
[0051] Style word embedding: User-provided natural language style keywords are encoded into style vectors and merged with hand-drawn style token ST to form a unified style representation, supporting few-shot fine-tuning to map customer styles to the token space.
[0052] Please see Figure 1 The present invention provides an embodiment of an AI video output system and method that combines text image models. The cross-modal alignment module aligns the text representations from the physical encoder, the hand-drawn encoder and style keywords to a unified semantic-style latent space L, so that the generator can faithfully render the style specified by the user while maintaining structural accuracy.
[0053] Embedding Alignment: Employing contrastive learning or CLIP-style multimodal alignment strategies, images, sketches, and text encodings are aligned. Training data can include triples (object images, sketches, style descriptions); Cross-Attention Fusion: The fusion layer conditionally injects structural information (depth, keypoints) and style tokens into the fusion layer, outputting a semantic-style code for each object or region; Few-Shot Adaptation: Using adapter technology, styles are quickly mapped to the latent space after the client provides a small number of samples, supporting few-shot style transfer; Output: Semantic-style latent vector Li for each shot or object;
[0054] Optional mechanisms: Style intensity control: An adjustable weight αstyl is reserved in Li, which is used to adjust the style intensity (from 0 to 1) by pressing a slider at runtime; Local style map: In addition to the global token, the system can generate a local style map to support different styles for different objects or regions;
[0055] Script analysis and storyboard planning:
[0056] The NLP parser is used to break down the introductory text or script into shot units. Each shot includes: shot description, suggested duration, subject, props, action tags, emotion tags, and priority controls (such as keeping the head in close-up). If the script contains explicit duration and frame count, they are directly mapped; otherwise, a default duration template or an estimation based on shot type is used.
[0057] Scene graph construction: Construct scene-graph nodes for each shot: subject, background, props, actions, and emotions; each node is associated with semantic-style latent Li as well as priority and constraints;
[0058] Camera parameter generation: Generates a camera parameter set C={Pitch, Yaw, Distance R, Focal Length F, Motion trajectory M(t)} for the lens. These parameters can be generated directly by the script parser or selected by the user template. Supported lens motion types: stationary, translation, push-pull, track, rotation. The motion trajectory can be represented as the camera's path in 3D space or a 2D plane interpolated path.
[0059] Output: Returns a Shot List: Each shot contains a shot ID, duration, camera parameter C, scene-graph, and corresponding Li, and is written to the timeline.
[0060] Please see Figure 1The present invention provides an embodiment of an AI video output system and method that combines a text image model. The multi-view generator generates several frames of illustrations corresponding to the lens direction based on the semantic-style implicit representation Li and the camera parameters C. The core objective is to maintain structural and stylistic consistency between different viewpoints and consecutive frames, while ensuring high quality and strong controllability of single frames.
[0061] Option A (Diffusion Condition Generator): Employs a conditional diffusion model, injecting Li and viewpoint encoding into the conditional network during the generation process, and controlling the generation style intensity and level of detail through sampling;
[0062] Option B (NeRF + Diffusion Hybrid): First, use NeRF / 2.5D scene to infer a rough geometry rendering map consistent with the viewpoint, and then use the diffusion model to render details and style. This option is more robust in terms of viewpoint consistency and is suitable for scenes with multiple viewpoint inputs or requiring high 3D coherence.
[0063] Option C (Implicit Representation Based on Differentiable Rendering): Construct an implicit function representation of view conditions and obtain a 2D image through a renderer, then adjust the brush strokes and textures using a stylization module;
[0064] Depth Consistency: The generator introduces a depth preservation mechanism during the training or inference phase to match the depth map of the generated image with the depth prior provided by the physical encoder. L1 / L2 depth loss or perceptual depth loss can be used.
[0065] Optical flow consistency: For adjacent frames, optical flow estimation or optical flow loss is introduced to ensure that texture points follow a reasonable trajectory during movement, preventing texture jitter and drift;
[0066] Latent space tracking: Maintain a unique identifier (ID) for each salient object in the latent space and perform cross-frame latent smoothing to ensure that the object style and geometry do not change drastically with each frame;
[0067] Cross-frame attention: The generator performs attention queries on features from previous or keyframes when generating the current frame in order to copy and preserve detailed structure;
[0068] Frame materials: Supports RGBA PNG sequences, as well as sliced textures and multi-layer images; Metadata: Each frame includes a depth map, semantic segmentation map, key point location, and object ID mapping table, which facilitates post-production compositing and lip-syncing.
[0069] Title, Q&A, and subtitle generation:
[0070] Automatically generate question-and-answer pairs, subtitle text, and social media copy that correspond to the semantics of each shot, while maintaining semantic relevance;
[0071] Prompt Design: Combine the Li and the shot description into a prompt for generative large models (such as language models based on Transformer) to generate questions and answers or captions: Example: "Shot description: {shot text}; main subject: {subject tag}; style: {style keywords}; please generate short captions and three relevant questions and answers for teaching scenarios";
[0072] Filtering and clustering: Generate several candidate texts, cluster them using semantic similarity, and select one or more texts that best match the semantics of the image;
[0073] Formatting; outputting in formats such as SRT / ASS / JSON, and generating a plain text version that can be directly consumed by TTS;
[0074] Optional: Multi-granularity and multi-purpose output: Platform adaptation: short videos (short subtitles, emphasis hooks), instructional videos (explanatory, step-by-step subtitles), product promotion (highlights + CTA); Question and answer types: factual (descriptive), interactive (questioning to guide users), quiz (assessment questions); Dataset output: question and answer pairs can be packaged into JSON / CSV for downstream training or retrieval.
[0075] Please see Figure 1 The present invention provides an embodiment of an AI video output system and method that combines a text image model;
[0076] Audio synthesis and lip-sync:
[0077] TTS Synthesis: TTS with Emotion Control Support: Receives parameters such as emotion tags, speech rate, pause markings, and timbre selection, and outputs a timestamped phoneme sequence and a complete audio stream; Output content: audio file, phoneme time alignment table, and sentence emotion annotation;
[0078] Lip-sync: Mapping factor sequences to viseme sequences, commonly using mapping tables or language-specific mappings; Lip-sync keyframe generation: Obtaining the lip-sync category or mouth morphology parameters (e.g., mouth opening degree, lip protrusion degree, etc.) for each video frame based on factor timeline interpolation; For 2D / 2.5D tasks, skeletal-based mouth deformation or shape-interpolation-based lip synthesis can be used; For real-life face replacement scenes, face reenactment technology can be used; Lip-sync calibration: Phoneme-viseme mapping calibration can be performed using a small number of samples to adapt to different character styles;
[0079] Use timeline commands to align audio to each shot. If the shot duration does not match the audio clip, you can stretch or cut the clip. If the dialogue is too early or too late, the system can provide fine-tuning suggestions (adjusting shot duration or re-synthesizing audio).
[0080] Timeline compositing and exporting:
[0081] Composite multi-layered materials (foreground, subject, background, effects, subtitles, QA pop-up, audio track) along the timeline, supporting layer blending, masking, motion interpolation, and special effects filters; inter-shot transition plugins: dissolve, erase, push-pull, rotation transitions, and custom stylized transitions;
[0082] When a user fine-tunes a shot parameter or keyframe, only the affected frame or adjacent frame is re-rendered, while the differential cache of the unchanged frame is kept to save computing resources. The differential cache records: frame ID, source latency, view parameters, and rendering version number. When the system attempts to reuse the cache, it first compares the latency with the parameter hash. If they match, the cache is reused directly.
[0083] Video output: MP4, MOV, supporting different resolutions and bitrate configurations; Project export: AE / PR XML, JSONproject includes timeline, layer paths and metadata, facilitating subsequent manual retouching; Material package: frame-by-frame PNG sequence, depth map, segmentation map, lip-sync keyframe table, subtitle file and QA data.
[0084] Working principle: The system first receives mixed reference information - physical reference image, hand-drawn reference image, hand-drawn style keywords and introductory text. The physical image and the hand-drawn image are respectively fed into the physical encoder and the hand-drawn encoder, which extract structural and geometric information such as density depth, semantic segmentation, and key points, as well as style information such as lines, strokes, and style tokens. The style keywords are converted into adjustable style vectors by the style embedder. In this stage, the structural prior and style prior are encoded in parallel, laying the foundation for subsequent cross-modal fusion.
[0085] After encoding, the structural prior, style vector and text semantics are mapped to a unified semantic-style latent space through the cross-modal alignment module. Based on this latent representation, the system introduces a view decoder, a view transformation module and a view-aware multi-view generator to generate multi-angle illustration sequences under specified or script-generated camera parameters. During the generation process, depth, optical flow temporal consistency, cross-frame attention and latent space smoothing are applied in parallel to ensure the consistency of structure and artistic style between different viewpoints and frames.
[0086] The system parses the introductory text or script into storyboards and scene graphs, and maps these into viewpoint vectors, generation parameters, and timeline instructions. Based on the semantics of the shots, it automatically generates question-and-answer and subtitle text, uses emotional TTS to synthesize speech, and maps the phoneme timeline to lip-sync keyframes. Finally, the timeline synthesis module synthesizes the illustrations, subtitles, QA, and audio tracks into a complete video according to the storyboard. The evaluation module provides quality feedback and fine-tuning for viewpoint consistency, style consistency, and audio-visual synchronization.
[0087] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An AI video output system incorporating a text image model, characterized by: The system comprises: a reference input module: to receive real object reference images, hand-drawn reference images, hand-drawn style keywords and introduction texts; a dual encoder module: including a real object encoder and a hand-drawn encoder; a cross-modal alignment module: for mapping the geometric prior and style token from the dual encoder into a fused semantic-style latent representation through a multi-modal alignment strategy; a view transformation and multi-view generation module: including a view encoder and a multi-view generator based on semantic-style latent representation and view vectors, for generating a multi-angle illustration sequence for different camera parameters, and maintaining the structural and style consistency between different views through temporal consistency constraints during the generation process; a script analysis and storyboard planning module: for analyzing the input introduction text or video script into shot units, and mapping the shot units into view vectors, generation parameters and time axis instructions; a question and answer subtitle generation module: for automatically generating question and answer data, subtitles or explanatory texts corresponding to each shot semantic based on semantic-style latent representation and shot units.
2. The AI video output system combined with a text image model according to claim 1, characterized in that: The system further comprises: an audio synthesis and lip synchronization module: for generating emotional synthesized speech based on the subtitles or explanatory texts, and aligning the phoneme time axis of the generated speech with the target lip template to achieve temporal speech-lip synchronization; a timeline synthesis module: for combining the multi-angle illustration sequence, question and answer subtitles and audio tracks into a video output file in chronological order according to the time axis instructions of the shot units; an evaluation and optimization module: for evaluating the generated video in terms of view consistency, style consistency, semantic alignment and audio-visual synchronization, and outputting optimization suggestions.
3. The AI video output system combined with a text image model according to claim 1, characterized in that: The real object encoder includes sub-modules for semantic segmentation, dense depth estimation and key point detection, and the hand-drawn encoder includes sub-modules for line drawing, brush stroke feature extraction and sketch vectorization, and the depth and key points output by the real object encoder are used to guide three-dimensional layout and occlusion processing in the view transformation and multi-view generation module.
4. The AI video output system combined with a text image model according to claim 1, characterized in that: The cross-modal alignment module uses contrastive learning and cross-attention mechanisms to map text style keywords, hand-drawn style tokens and visual structure prior into a unified semantic-style embedding space, and introduces a small sample style fine-tuning strategy during training to support style transfer with a small number of hand-drawn samples; The cross-modal alignment module and the view transformation and multi-view generation module work together to realize the perceptual propagation of style across multiple views and multiple frames based on depth, optical flow temporal consistency constraints and cross-frame attention mechanisms, thereby ensuring the consistency of the structure and artistic style of the output video in shot switching and shot motion.
5. The AI video output system in combination with a text image model according to claim 1, characterized in that: The view transformation and multi-view generation module uses a conditional diffusion generation network or a NeRF and diffusion hybrid generation architecture, the view decoder can map the camera parameters analyzed from the script into a latent space transformation vector, and introduce a temporal loss based on depth consistency and optical flow constraints during the generation process to reduce cross-frame texture drift.
6. The AI video output system in combination with a text image model according to claim 1, wherein: The script analysis and shot planning module parses the introduction text into a set of shot descriptions based on a natural language processor, and automatically generates a scene graph for each shot, which includes subject, background, props, action, and emotion tags. The scene graph is used to determine the composition, perspective, and shot motion parameters for the multi-perspective generator.
7. The AI video output system in combination with a text image model according to claim 1, wherein: The question and answer subtitle generation module includes a generative large model-based question and answer generator and a subtitle style generator. The question and answer generator uses semantic similarity filtering and clustering steps to ensure that the generated questions and answers are directed towards the prominent objects in the shot. The subtitle generator outputs subtitle files of different granularities and formats based on the target platform.
8. The AI video output system combined with a text image model according to claim 2, characterized in that: The audio synthesis and lip-sync module includes an emotion control TTS sub-module, a phoneme timeline generator, and a keyframe-based or 2D skeleton-based lip synthesis module. It supports controllable adjustment of synthesized speech through emotion parameters, speech rate parameters, and timbre parameters, and aligns the phoneme timeline with the target character's mouth keyframes to achieve visual lip-sync. The timeline synthesis module supports incremental re-rendering strategy, optical flow-driven inter-frame interpolation, color matching, and shot transition interpolation. It also supports exporting editable project files and differential cache to improve iteration efficiency. 9.A method for AI video output combined with a text image model, applied to the AI video output system combined with the text image model according to any one of claims 1-8, characterized in that: The method comprises the following steps: S1, receiving at least one real object reference image, at least one hand-drawn reference image, hand-drawn style keywords, and introduction text; S2, encoding the real object reference image and the hand-drawn reference image respectively to obtain geometric priors and style tokens; S3, aligning the geometric priors and style tokens through cross-modal alignment to obtain a fused semantic-style hidden representation; S4, parsing the introduction text or script into a list of shot units, and generating camera parameters and timeline instructions for each shot; S5, for each shot, based on the semantic-style hidden representation of S3 and the camera parameters, using a perspective-aware multi-perspective generator to generate a corresponding sequence of illustration frames, while imposing perspective consistency and temporal consistency constraints during the generation process; S6, based on the content of S3 and S4, automatically generating question and answer pairs corresponding to each shot, and synthesizing emotional language based on the subtitle text; S7, synthesizing the illustration sequence generated in S5, the subtitle and voice generated in S6 into a final video output file according to the timeline instructions in S4; S8, evaluating the video for perspective consistency, style consistency, voice alignment, and audio-visual synchronization, and performing local re-rendering or parameter fine-tuning based on the evaluation results.
10. The AI video output method in combination with a text image model according to claim 9, characterized in that: The multi-perspective generator in S5 uses the following losses or constraints during training and generation: content-aware loss, style loss, perspective consistency loss, adversarial detail enhancement loss, and cross-modal semantic consistency loss. During generation, cross-frame attention and hidden space smoothing strategies are used to maintain structural and stylistic coherence between different perspectives and consecutive frames. In the fine-tuning stage, few-shot fine-tuning is used with a small number of hand-drawn samples to adapt to customer-specified styles.
Citation Information
Patent Citations
A multimedia video stream management system and method based on artificial intelligence
CN119046481B
Text-guided portrait nerve drawing method based on implicit coding
CN119809919A
Video portrait consistency editing method based on optical flow guiding and text driving
CN119887502A
Video generation control method and computer readable storage medium
CN120017931A
Face consistency multi-angle lens video generation method and device
CN120321472A
Cited By
AIGC video background generation control method based on freehand background structure prior
CN121883657A
Intelligent video propaganda product design method based on large model and knowledge base
CN122064843A
Creative arrangement method and compilation management system based on hexagonal box
CN122372814A