A movie-level animation video generation system and method based on multi-modal fusion

The multimodal fusion system solves the problems of character drift, insufficient 3D depth perception, inaccurate audio-visual synchronization, low image quality, and inability to uniformly model multimodal information in animation video generation, and realizes the generation of high-quality, commercially viable industrial-grade animation videos.

CN122340329APending Publication Date: 2026-07-03黄承斌

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
黄承斌
Filing Date
2026-03-25
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing AI video generation technologies suffer from problems in film-level and industrial-level animation video production, such as character drift, lack of 3D depth perception, inaccurate audio-visual synchronization, abrupt frame transitions, low image quality, and the inability to uniformly model multimodal information. These issues result in uncontrollable generation effects and make large-scale deployment difficult.

Method used

The system adopts a multimodal fusion architecture, which achieves character consistency, 3D depth perception, audio-visual synchronization, image quality enhancement and unified modeling of multimodal information through character ID embedding and persistence, depth estimation and attention fusion, audio-visual cross-attention, professional camera movement coding and high-definition rendering.

Benefits of technology

It achieves unified modeling of permanent character locking, realistic 3D depth perception, high-definition image quality, professional film and television effects, and multimodal information, supports multi-GPU training and large-scale deployment, and generates animation videos that meet industrial-grade commercial standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a kind of movie-level animation video generation system and method based on multimodal fusion, it is related to artificial intelligence and digital content creation technical field.The system includes six big core modules of multimodal input, multimodal feature coding, audio feature processing, DiT main network, high-definition rendering, video synthesis and output;Method flow is to collect text, character ID, shot, style, audio and other multimodal input information, after feature coding, audio processing and feature fusion, time series modeling is carried out by DiT main network, video features are generated by combining depth perception, character consistency maintenance, film panning control technology, and then high-definition rendering, film and picture synchronization are completed to output movie-level animation video.The present application effectively solves the problems of character drift, lack of 3D space feeling, poor audio-visual synchronization, stiff panning and low image quality in existing AI video generation, realizes permanent locking of characters, real 3D depth perception, accurate audio-visual synchronization and 4K high-definition image output, and the overall system can be deployed at industrial level, suitable for commercial scenarios such as animation production and film creation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, computer vision and deep learning, specifically to a cinematic animation video generation system and method based on multimodal fusion, belonging to the category of generative AI and digital content creation technology. Background Technology

[0002] While current AI video generation technology is developing rapidly, it still has many core shortcomings in film-level and industrial-grade animation video production scenarios: First, the appearance and posture of characters are prone to drift during multi-frame generation, making it impossible to permanently lock the character; second, it is based solely on 2D image generation and lacks... The film suffers from several shortcomings: First, it lacks true 3D depth perception, resulting in insufficient spatial depth; second, the audio-visual synchronization is low, failing to accurately match audio beats with camera movement and lip movements; third, the frame transitions are abrupt, lacking professional cinematic camerawork logic, leading to poor visual presentation; and fourth, the image resolution is low. The existing technologies fail to meet 4K cinematic commercial standards; sixth, they cannot achieve unified modeling of multi-dimensional information such as text, characters, audio, shots, and style, resulting in poor controllability of the generated effects. None of these existing technologies have solved all the above problems, making it impossible to create scalable, industrial-grade animation. Video generation solutions. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a cinematic animation video generation system and method based on multimodal fusion. Addressing the shortcomings of existing AI video generation methods in areas such as character consistency, 3D depth perception, audio-visual synchronization, temporal coherence, high-definition image quality, and multimodal fusion. Overcoming technical shortcomings, we can achieve industrial-grade video generation capabilities that are trainable, commercially viable, and scalable.

[0004] The system of this invention integrates various user input information through a multimodal input module, and then converts it into unified feature vectors through a multimodal feature encoding module. The audio feature processing module performs deep audio processing and audio-video feature fusion, while the DiT backbone network module implements multi-feature fusion and temporal processing. The modeling and high-definition rendering modules complete image decoding and quality optimization, while the video compositing and output module finally generates and exports the video. The system achieves permanent character locking through character ID embedding and feature persistence, and constructs realistic 3D spatial relationships through depth estimation and attention fusion. The system achieves precise audio-visual synchronization through a cross-attention mechanism, enables controllable generation of multiple styles through style mixing branches, and utilizes professional camera work. Encodes to achieve cinematic-quality camera effects. Attached Figure Description

[0005] Figure 1 The system's overall architecture flowchart illustrates the connection relationships and data transmission flow of each core module, clarifying the overall operating logic of the system.

[0006] Figure 2 This diagram illustrates the process of permanently locking a character, showing the complete process of character ID creation, from input, feature generation, enhancement, to persistent storage and multi-frame reuse. Implementation process.

[0007] Figure 3 This is a schematic diagram of the 3D depth perception process, showing the extraction of depth information after an image frame is input, feature projection onto the image, and integration of the image with the main visual features. Processing steps.

[0008] Figure 4 This is a schematic diagram of the audio-visual synchronization process, illustrating audio input, feature extraction, beat and lip-sync generation, audio-video feature fusion, and beat-driven processing. The complete logic of camera movement.

[0009] Figure 5 This is a flowchart illustrating the Style Hybrid Expert (MoE) process, showing style parameter processing, multi-branch weight allocation, and multi-style feature fusion output. The implementation process.

[0010] Figure 6 This is a schematic diagram of the cinematic camera movement control process, illustrating the flow from shot type input, encoding, and feature projection to camera movement enhancement and professional camera movement implementation. Process steps.

[0011] Figure 7 This is a schematic diagram of the overall video generation method, showing the complete steps from user information input to the final video synthesis output. Detailed Implementation Users can input video description text, specify character ID, select shot type, set style parameters, and upload audio through the multimodal input module. The system receives the complete input information, determines the target frame number, and then initiates the generation process. The multimodal feature encoding module separately encodes text, angles, and other data. Color ID, lens type, and style parameters are encoded, and inter-frame temporal relationship modeling is completed through the temporal coding submodule to generate a unified... Multimodal feature vectors.

[0012] The audio feature processing module synchronously encodes the input audio, extracts core audio features, and captures the audio beat through the beat extraction submodule. The signal, the lip-sync submodule generates lip features that match the human voice, and then the audio-video cross-attention submodule combines the audio features with multiple signals. Modal features are deeply fused, and the beat signal is associated with the camera movement through the beat-driven mirror module to achieve audio-visual coordination.

[0013] The fused features are input into the DiT backbone network module, the depth perception attention submodule incorporates 3D depth features, and the role consistency preservation submodule... Reusing fixed character traits to avoid drift; the film camera movement attention submodule controls inter-frame spatial changes based on shot encoding; style modulation before... The feed module maintains a consistent overall style, while the temporal DiT block module completes inter-frame temporal coherence modeling and generates high-quality video features.

[0014] The high-definition rendering module converts video features into image frames through the VAE decoding submodule, while the depth estimation submodule optimizes the spatial stereoscopic effect of the image. The 4K supramolecular module upscales image resolution to 4K, while the light and shadow enhancement submodule optimizes lighting effects to achieve cinematic image quality. Finally, the video synthesis and output module combines multiple frames into a video sequence, and the audio-visual synchronization submodule completes the precise audio and video synchronization. The alignment and character saving submodules persistently store character features for easy reuse later. The file export submodule exports the final video in a commercially usable format. Export the formula to complete the entire generation process.

[0015] This system supports multi-GPU distributed training and mixed-precision training, can be visualized through a WebUI, and can be scaled up using Docker. The system can be deployed in a customized manner, flexibly generating cinematic-quality animation videos with different styles, camera angles, and image qualities based on different user input parameters. It is suitable for various commercial applications such as animation production, film and television post-production, advertising creation, and game development.

[0016] 6. Beneficial effects Compared with existing AI video generation technologies, this invention has the following significant advantages: 1. Achieve permanent character locking by embedding character IDs, enhancing features, and persistently storing them to ensure that character features remain unchanged during multi-frame generation. The movement and consistency are maximized, completely resolving the character deformation issue; 2. Possesses true 3D depth perception capabilities, simulating real spatial distances and lighting changes through depth estimation and attention fusion, significantly improving... The visual depth and realism; 3. Achieve high-precision audio-visual synchronization through audio-video cross-attention fusion and beat-driven camera movement, matching camera movement with audio beats and lip movements. The synchronized visuals and voices significantly enhance the immersive experience of the video. 4. Supports professional cinematic camera movement control, enabling various professional shot effects such as close-ups, wide shots, panning, and rotation, with smooth frame transitions. Equipped with professional-quality film and television visuals; 5. Capable of outputting 4K cinematic high-definition image quality, meeting various commercial display and dissemination needs through super-resolution and lighting optimization; 6. Multimodal information is fused and modeled in a unified manner, with strong controllability of input parameters, and the generated results are highly matched with user needs; 7. The system has industrial-grade deployment capabilities, supports multi-GPU training, visualization operation, and large-scale application, and has outstanding commercial value and practicality.

[0017] Figure 1 Multimodal input module → Multimodal feature encoding module → Audio feature processing module → DiT backbone network module → High-definition rendering module → Video synthesis and output module Figure 2 Character ID input → Character embedding vector generation → Position encoding enhancement → Feature normalization → MLP feature enhancement → Persistent storage of character features → Reusing the same feature across multiple frames → Achieving permanent character locking Figure 3 Input image frame → Depth estimation network → Depth map generation → Depth feature projection → Depth-aware attention fusion → Integration of visual features from the DiT backbone → Realization of 3D spatial perception modeling Figure 4 Audio input → Audio encoding feature extraction → Beat signal extraction → Lip shape feature generation → Audio-video cross-attention fusion → Visual feature synchronization with audio signal → Beat-driven camera movement → Output audio-visual synchronization features Figure 5 Style parameter input → Gated network weight allocation → Anime style expert branch → Film style expert branch → Realistic style expert branch → Weighted fusion of multiple branches → Output unified style features Figure 6 Lens type input → Lens type encoding → Camera movement feature projection → Integration of visual features → Enhanced camera movement attention → Control of inter-frame spatial changes → Realization of professional camera movements such as push-pull, pan, and tilt. Figure 7 The process involves: receiving input information → multimodal feature encoding → audio feature processing and fusion → temporal modeling and video feature generation → high-definition rendering → video synthesis and output.

Claims

1. A cinematic animation video generation system based on multimodal fusion, characterized in that, Including multimodal input modules connected in sequence, Multimodal feature encoding module, audio feature processing module, DiT backbone network module, high-definition rendering module, and video synthesis and output module; The multimodal input module is used to receive text, character ID, shot type, style parameters, and audio information; The multimodal feature encoding module is used to transform various types of input information into a unified feature vector; The audio feature processing module is used for audio feature extraction, beat recognition, lip shape generation, and audio-video feature fusion. The DiT backbone network module is used for temporal modeling and video feature generation of multimodal fusion features; The high-definition rendering module is used to decode, optimize, and enhance the generated features. The video synthesis and output module is used for frame sequence synthesis, audio-visual synchronization, character information saving, and file export.

2. The system according to claim 1, characterized in that, The multimodal input module includes a text input submodule, a character ID input submodule, a shot type input submodule, a style parameter input submodule, and an audio input submodule.

3. The system according to claim 1, characterized in that, The multimodal feature encoding module includes a text encoding submodule, a character ID embedding submodule, a shot type encoding submodule, a style mixing encoding submodule, and a temporal encoding submodule.

4. The system according to claim 1, characterized in that, The audio feature processing module includes an audio encoding submodule, a beat extraction submodule, a lip-sync submodule, an audio-video cross-attention submodule, and a beat-driven motion mirroring module.

5. The system according to claim 1, characterized in that, The DiT backbone network module includes a depth perception attention submodule, a role consistency maintenance submodule, a cinematic camera movement attention submodule, a style modulation feedforward submodule, and a temporal DiT block module.

6. The system according to claim 1, characterized in that, The high-definition rendering module includes a VAE decoding submodule, a depth estimation submodule, a 4K supramolecular module, and a lighting enhancement submodule.

7. The system according to claim 1, characterized in that, The video synthesis and output module includes a frame sequence synthesis submodule, an audio-visual synchronization submodule, a character saving submodule, and a file export submodule.

8. A method for generating cinematic-quality animated videos based on multimodal fusion, characterized in that, Includes the following steps: Step 1: Obtain text prompts, character ID, camera type, style parameters, audio, and target frame number through the multimodal input module; Step 2: Encode the above information using the multimodal feature encoding module to generate a multimodal feature vector; Step 3: Process the audio using the audio feature processing module to achieve audio feature extraction and audio-video feature fusion; Step 4: Input the fused features into the DiT backbone network module for temporal modeling to generate video features; Step 5: Decode and render the video features using the high-definition rendering module to obtain high-definition image frames; Step 6: Complete frame synthesis, audio-visual synchronization, character saving, and video output through the video synthesis and output module.

9. The method according to claim 8, characterized in that, The audio and video processing flow in step 3 is as follows: audio input → audio encoding → beat extraction → lip shape feature generation → audio and video cross-attention fusion → beat-driven camera movement → output synchronization features.

10. The method according to claim 8, characterized in that, In step 4, the DiT backbone network synchronously integrates deep perception features, character consistency features, cinematic camera movement features, and style features to generate video features that are sequential, stable in character, and smooth in camera movement.