Automatic video generation system and method based on AI Agent multi-mode cooperative control

The AIAgentic multimodal collaborative control automated video generation system solves the problem of insufficient multimodal fusion and dynamic adaptation capabilities in existing technologies, and achieves spatiotemporal consistency in high-quality video generation and improves user experience.

CN121126084AActive Publication Date: 2025-12-12BEIJING SOIN TECH CORP LTD

Patent Information

Application Number
CN202511213638.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-12
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing automated video generation methods have shortcomings in multimodal fusion, dynamic adaptation capabilities, and physical-semantic processing, resulting in poor spatiotemporal consistency and user experience of the generated videos.

Method used

The automated video generation system employing AIAgentic multimodal collaborative control achieves intelligent control throughout the entire process from multimodal input to high-quality video output through hierarchical intent parsing and dynamic graph construction, multi-scale spatiotemporal capsule scene generation, physical-semantic joint constraint verification, and spatiotemporal consistency optimization.

Benefits of technology

It improves the accuracy, coherence, and user experience of generated videos, and achieves flexibility in intent parsing, physical-semantic collaborative verification, and strong guarantees of spatiotemporal consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126084A_ABST
    Figure CN121126084A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic video generation system and method based on AI Agent multi-mode cooperative control. The system comprises a demand acquisition module used for acquiring multi-modal generation demand information; the hierarchical intention analysis and dynamic graph construction module is used for generating an intention graph with priorities and constraint conditions; the multi-scale space-time capsule scene generation module is used for generating multi-scale space-time capsule modeling information; the physical-semantic joint constraint verification module is used for acquiring the verified modeling information; the space-time consistency optimization and content fusion module is used for acquiring a video frame sequence after space-time consistency optimization; and the dynamic code rate video synthesis and output module is used for generating a video file. Through dynamic adaptive architecture and global joint optimization, the problems of stiffness and separation of a traditional video generation technology are solved, and full-process intelligent control from multi-mode input to high-quality video output is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video generation, in particular to an AIAgentic multi-modal collaborative control automatic video generation system and an AIAgentic multi-modal collaborative control automatic video generation method. BACKGROUND

[0002] The existing automatic video generation method adopts a traditional pipeline architecture, which is divided into four stages of multi-modal input analysis, static intent mapping, scene step-by-step generation, and physical post-verification. Multi-modal processing relies on fixed rules (such as keyword matching and template library), and the intent analysis is rigid and cannot understand emerging combinations (such as the implicit contradiction of "dragon spewing fire + clear sky"). The intent mapping decomposes sub-intents through hard-coded rules, and the graph is statically fixed and lacks dynamic adjustment capability. Scene generation is divided into macro, meso, and micro steps for splicing, action matching is preset in the library, and details rely on texture mapping, which can easily lead to time-space inconsistency. Physical verification and semantic checking are processed in series, only the error is marked for manual parameter adjustment, and the physical law and user intent cannot be optimized jointly. The core defects are weak dynamic adaptive capability, simple multi-modal fusion, and separation of physical-semantic processing, which makes it difficult to balance the quality and feasibility. SUMMARY

[0003] The purpose of the present application is to provide an AIAgentic multi-modal collaborative control automatic video generation system to at least solve one of the above technical problems.

[0004] In one aspect of the present application, an AIAgentic multi-modal collaborative control automatic video generation system is provided, which comprises:

[0005] A requirement acquisition module for acquiring multi-modal generation requirement information provided by a user;

[0006] A hierarchical intent analysis and dynamic graph construction module for analyzing the multi-modal generation requirement information to generate an intent graph with priority and constraint conditions;

[0007] A multi-scale spatiotemporal capsule scene generation module for generating multi-scale spatiotemporal capsule modeling information according to the dynamically evolving intent graph;

[0008] A physical-semantic joint constraint verification module for verifying the multi-scale spatiotemporal capsule modeling information to obtain verified modeling information;

[0009] The spatio-temporal consistency optimization and content fusion module is configured to perform spatio-temporal consistency optimization and content fusion on the verified modeling information, thereby obtaining a sequence of video frames after spatio-temporal consistency optimization.

[0010] The dynamic code rate video synthesis and output module is configured to generate a video file according to the sequence of video frames after spatio-temporal consistency optimization.

[0011] Optionally, the requirement obtaining module comprises:

[0012] The text information obtaining module is configured to obtain text information provided by a user.

[0013] The visual information obtaining module is configured to obtain visual information provided by a user.

[0014] The audio information obtaining module is configured to obtain audio information provided by a user.

[0015] The text processing module is configured to extract text features in the text information and expand the text features, thereby obtaining text modality features, wherein the text modality features comprise a text semantic vector, an expanded keyword list, and corresponding semantic weights.

[0016] The visual processing module is configured to extract features of the visual information, thereby obtaining visual modality features, wherein the visual modality features comprise an object-level visual feature vector and a visual dynamic trajectory feature vector.

[0017] The audio processing module is configured to extract the audio information, thereby obtaining audio modality features, wherein the audio modality features comprise an acoustic feature vector and an emotion label.

[0018] The dynamic modality weight distribution module is configured to respectively distribute weights to the text modality features, the visual modality features, and the audio modality features, thereby obtaining a multi-modal feature vector after dynamic weight distribution.

[0019] The cross-modal residual learning and alignment module is configured to perform residual adjustment on the multi-modal feature vector after dynamic weight distribution, thereby obtaining a multi-modal feature vector after residual adjustment, which is taken as the requirement information.

[0020] Optionally, the hierarchical intention analysis and dynamic graph construction module comprises:

[0021] an intent classification module configured to generate a core intent set according to the multi-modal generation requirement information, the core intent set including at least one core intent information;

[0022] a confidence score module configured to score the confidence of each core intent information;

[0023] a sub-intent recursive decomposition module configured to decompose each core intent information respectively to obtain executable sub-intents, one core intent information including at least one sub-intent;

[0024] a hierarchical graph construction module configured to generate a hierarchical intent graph according to the core intent information and the sub-intents; wherein the nodes of the hierarchical intent graph include the core intent information and the sub-intents, and the edges are dependency relationships;

[0025] a dynamic graph evolution and adjustment module configured to generate a dynamically adjusted intent graph according to the hierarchical intent graph;

[0026] an intent priority ranking and constraint generation module configured to adjust the priority and constraint conditions of each node in the dynamically adjusted intent graph to generate an intent graph with priority and constraint conditions.

[0027] Optionally, the multi-scale spatio-temporal capsule scenario generation module includes:

[0028] a macro-scale scenario layout generation module configured to generate a macro scenario capsule according to the intent graph with priority and constraint conditions;

[0029] a meso-scale role action and special effect generation module configured to generate a meso action and special effect capsule according to the macro scenario capsule and the requirement information;

[0030] a micro-scale detail texture and light generation module configured to generate a micro detail and light capsule according to the macro scenario capsule, the meso action and special effect capsule, and the requirement information;

[0031] a spatio-temporal capsule routing and cross-scale connection module configured to fuse the macro scenario capsule, the meso action and special effect capsule, and the micro detail and light capsule into a spatio-temporal consistent capsule set;

[0032] a multi-scale content fusion and preliminary rendering module, configured to perform capsule fusion on the set of spatiotemporal consistency capsules and preliminary rendering, so as to obtain a sequence of preliminary rendered video frames as multi-scale spatiotemporal capsule modeling information.

[0033] Optionally, the physical-semantic joint constraint verification module comprises:

[0034] a joint constraint field construction module, configured to generate a joint constraint field model according to the multi-scale spatiotemporal capsule modeling information and the intention graph with priority and constraint conditions;

[0035] a simulated annealing optimization module, configured to perform optimization on the joint constraint field model by a simulated annealing algorithm, so as to obtain optimized generation parameters;

[0036] a real-time physical engine verification module, configured to verify the optimized generation parameters, so as to generate physical violation labels and correction suggestions;

[0037] a real-time semantic discriminator verification module, configured to perform semantic scoring and evaluation on the optimized generation parameters, so as to obtain semantic violation labels and correction suggestions;

[0038] a generation parameter adjustment module, configured to adjust the optimized generation parameters according to the physical violation labels and correction suggestions and the semantic violation labels and correction suggestions, so as to obtain adjusted generation parameters;

[0039] a re-rendering module, configured to adjust the multi-scale spatiotemporal capsule modeling information according to the adjusted generation parameters, so as to obtain verified modeling information.

[0040] Optionally, the spatiotemporal consistency optimization and content fusion module comprises:

[0041] a spatiotemporal consistency detection and labeling module, configured to perform spatiotemporal consistency detection on the verified modeling information, so as to obtain spatiotemporal inconsistency labels, the spatiotemporal inconsistency labels comprising frame numbers and region coordinates that need to be optimized;

[0042] an optical flow prediction and motion trajectory smoothing module, configured to process according to the spatiotemporal inconsistency labels, so as to form a smoothed motion vector field, the smoothed motion vector field comprising corrected inter-frame motion parameters;

[0043] The temporal discriminator verification and coherence optimization module is configured to generate a sequence of video frames that pass the temporal coherence verification according to the smoothed motion vector field;

[0044] The multi-scale content fusion and final rendering is configured to generate a sequence of spatio-temporally optimized video frames according to the sequence of video frames that pass the temporal coherence verification, the macroscopic scene capsules, the mesoscopic action and special effect capsules, and the microscopic detail and lighting capsules.

[0045] Optionally, the dynamic bit rate video synthesis and output module comprises:

[0046] The content importance analysis and region division module is configured to identify the sequence of spatio-temporally optimized video frames to obtain a content importance atlas, the content importance atlas comprising importance scores of each region of each frame;

[0047] The dynamic bit rate allocation strategy formulation module is configured to perform dynamic bit rate allocation according to the content importance atlas to generate a dynamic bit rate allocation table, the dynamic bit rate allocation table comprising bit rate values of each region of each frame;

[0048] The video encoding and dynamic bit rate compression is configured to encode and compress the sequence of spatio-temporally optimized video frames according to the dynamic bit rate allocation table to obtain a dynamic bit rate compressed video file.

[0049] Optionally, the hierarchical intent analysis and dynamic atlas construction module further comprises:

[0050] The intent contradiction detection and self-repair module is configured to perform contradiction detection and repair on the intent atlas adjusted dynamically to obtain a repaired intent atlas; and the intent priority ordering and constraint generation module is configured to process the repaired intent atlas.

[0051] Optionally, the intent contradiction detection and self-repair module comprises:

[0052] The intent contradiction detection module is configured to perform intent contradiction detection by using an ICOF formula to obtain a contradiction mark;

[0053] The intent contradiction repair module is configured to perform intent contradiction repair according to the contradiction mark to obtain a repaired intent atlas.

[0054] The application also provides an AI Agentic multi-modal collaborative control automated video generation method, characterized in that the AI Agentic multi-modal collaborative control automated video generation method comprises:

[0055] obtaining multi-modal generation requirement information provided by a user;

[0056] analyzing the multi-modal generation requirement information, thereby generating an intention graph with priority and constraint conditions;

[0057] generating multi-scale spatiotemporal capsule modeling information according to the dynamically evolved intention graph;

[0058] verifying the multi-scale spatiotemporal capsule modeling information, thereby obtaining verified modeling information;

[0059] performing spatiotemporal consistency optimization and content fusion on the verified modeling information, thereby obtaining a video frame sequence after spatiotemporal consistency optimization;

[0060] generating a video file according to the video frame sequence after spatiotemporal consistency optimization.

[0061] The AI Agentic multi-modal collaborative control automated video generation method of the application breaks through the rigidity and separation problem of traditional video generation technology through dynamic adaptive architecture and global joint optimization, realizes intelligent control of the whole process from multi-modal input to high-quality video output. The core advantages are flexibility of intention analysis, collaborative verification of physics-semantic, strong guarantee of spatiotemporal consistency, and precision of output adaptation, which significantly improve the accuracy, continuity and user experience of the generated video. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is a system schematic diagram of an AI Agentic multi-modal collaborative control automated video generation system of an embodiment of the application. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical scheme and advantages of the application clearer, the technical scheme of the embodiments of the application will be described in more detail below in combination with the drawings of the embodiments of the application. In the drawings, the same or similar notations represent the same or similar elements or elements with the same or similar functions throughout. The described embodiments are some embodiments of the application, not all embodiments. The embodiments described below by reference to the drawings are exemplary and are intended to explain the application, and cannot be understood as limiting the application. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the application. The embodiments of the application will be described in detail below in combination with the drawings.

[0064] The AIAgentic multi-modal collaborative control automated video generation system as shown in Figure 1 includes a requirement acquisition module, a hierarchical intention analysis and dynamic graph construction module, a multi-scale spatio-temporal capsule scene generation module, a physical-semantic joint constraint verification module, a spatio-temporal consistency optimization and content fusion module, and a dynamic code rate video synthesis and output module, wherein

[0065] The requirement acquisition module is configured to acquire multi-modal generation requirement information provided by a user.

[0066] The hierarchical intention analysis and dynamic graph construction module is configured to analyze the multi-modal generation requirement information to generate an intention graph with priority and constraint conditions.

[0067] The multi-scale spatio-temporal capsule scene generation module is configured to generate multi-scale spatio-temporal capsule modeling information according to the dynamically evolving intention graph.

[0068] The physical-semantic joint constraint verification module is configured to verify the multi-scale spatio-temporal capsule modeling information to obtain verified modeling information.

[0069] The spatio-temporal consistency optimization and content fusion module is configured to perform spatio-temporal consistency optimization and content fusion on the verified modeling information to obtain a sequence of video frames after spatio-temporal consistency optimization.

[0070] The dynamic code rate video synthesis and output module is configured to generate a video file according to the sequence of video frames after spatio-temporal consistency optimization.

[0071] The AIAgentic multi-modal collaborative control automated video generation method of the present application breaks through the rigidity and separation of traditional video generation technology through dynamic adaptive architecture and global joint optimization, and realizes intelligent control of the whole process from multi-modal input to high-quality video output. The core advantages are flexibility of intention analysis, collaborative verification of physical-semantic, strong guarantee of spatio-temporal consistency, and precision of output adaptation, which significantly improves the accuracy, coherence and user experience of the generated video.

[0072] In the present embodiment, the requirement acquisition module includes:

[0073] A text information acquisition module configured to acquire text information provided by a user.

[0074] A visual information acquisition module configured to acquire visual information provided by a user.

[0075] An audio information acquisition module configured to acquire audio information provided by a user.

[0076] a text processing module, configured to extract text features in the text information and expand the text features, so as to obtain text modality features, the text modality features including a text semantic vector, an expanded keyword list and a corresponding semantic weight;

[0077] a visual processing module, configured to extract features of the visual information, so as to obtain visual modality features, the visual modality features including an object-level visual feature vector and a visual dynamic trajectory feature vector;

[0078] an audio processing module, configured to extract the audio information, so as to obtain audio modality features, the audio modality features including an acoustic feature vector and an emotion label;

[0079] a dynamic modality weight distribution module, configured to respectively distribute weights for the text modality features, the visual modality features and the audio modality features, so as to obtain a multi-modal feature vector after dynamic weight distribution;

[0080] a cross-modal residual learning and alignment module, configured to perform residual adjustment on the multi-modal feature vector after dynamic weight distribution, so as to obtain a residual-adjusted multi-modal feature vector, the residual-adjusted multi-modal feature vector serving as the requirement information.

[0081] For example, the text information acquisition module can perform the following operation: a user inputs by typing: “want a video of a red dragon fighting on a snow mountain, with lightning and roar sound”, corrects “roar” to “dragon roar”, and finally obtains the cleaned text: “want a video of a red dragon fighting on a snow mountain, with lightning and dragon roar sound”.

[0082] For example, the visual information acquisition module can perform the following operation: a user uploads a picture of “snow mountain background” and a 10-second short video of “flying dragon”, the module first scales the picture to 512x512, extracts one key frame every two frames of the video (a total of 5 frames), and removes the watermark in the corners of the video, and finally obtains 6 visual reference materials.

[0083] For example, the audio information acquisition module can perform the following operation: a user uploads a mixed audio containing “thunder” and “wind”, the module first performs noise reduction processing, and then divides the audio into “thunder (3 seconds)” and “wind (5 seconds)” two independent segments through mute detection, as the input for subsequent audio feature extraction.

[0084] For example, the text processing module can perform the following operation: the text “red dragon fighting on a snow mountain” is processed as follows:

[0085] Semantic Vector: CLIP-Text encoded 768-dimensional vector containing the combined semantics of "red + dragon + snow mountain + battle";

[0086] Extended Keywords: [red dragon, dragon, snow mountain, battle, fight, snow, cliff, flame];

[0087] Semantic Weights: "Dragon" (0.9), "Battle" (0.8), "Red" (0.7), "Snow Mountain" (0.6), and the rest of the associated words (0.3-0.5).

[0088] For example, the visual processing module can perform the following operations:

[0089] Object-level visual feature vector:

[0090] Identify core objects in the visual material (such as "dragon", "snow mountain", "sky") using the object detection model YOLOv8;

[0091] For each object, extract a 2048-dimensional feature vector (containing shape, color, texture information) using the CNN model ResNet50.

[0092] Visual dynamic trajectory feature vector (for video / sequence frames):

[0093] Calculate the motion vector of adjacent frames using the optical flow model FlowNet to capture the object motion trajectory;

[0094] Time-series encode the trajectory (such as LSTM) to generate a 128-dimensional dynamic feature vector (containing speed, direction, acceleration information).

[0095] For example, the audio processing module can perform the following operations:

[0096] Acoustic feature vector:

[0097] Extract acoustic features such as MFCC (Mel Frequency Cepstral Coefficient, 39-dimensional), spectral entropy (1-dimensional), and zero-crossing rate (1-dimensional);

[0098] Fuse the features into a 256-dimensional vector (containing pitch, loudness, rhythm information) using a CNN model.

[0099] Emotion label generation:

[0100] Classify the audio using a pre-trained audio emotion classification model (such as CNN-LSTM) to output emotion labels (such as "majestic", "tense", "calm"); for audio without explicit emotions (such as pure wind), mark it as "neutral".

[0101] For example, the dynamic modality weight allocation module can perform the following operations:

[0102] Dynamically calculate weights with attention mechanism (such as multi-head attention of Transformer), ensure that the sum of weights is 1, and specifically, the weight calculation is based on the following:

[0103] Modal quality: text clarity (unambiguous -> high weight), visual resolution (high definition -> high weight), audio signal-to-noise ratio (no noise -> high weight);

[0104] User input intensity: if the user provides detailed text + fuzzy picture -> high text weight; if clear picture + brief text is provided -> high visual weight;

[0105] Cross-modal consistency: if the text "red dragon" is consistent with the color of the dragon in the visual picture, the weights of both are increased; if there is a conflict, the weight of the conflicting modality is reduced.

[0106] In this embodiment, the cross-modal residual learning and alignment module can perform the following operations:

[0107] Use the contrastive learning mechanism of CLIP to calculate the distance of text, visual, and audio features in the semantic space (the greater the distance, the greater the difference);

[0108] Through the residual network (ResNet), the modality features with large differences are corrected (such as adjusting the color dimension of the visual feature if the visual feature has a large deviation from the text "red");

[0109] Through cross-modal similarity calculation (such as cosine similarity), ensure that the similarity of each modality feature after adjustment is ≥0.8 (threshold can be dynamically adjusted).

[0110] In this embodiment, the hierarchical intent analysis and dynamic graph construction module includes:

[0111] An intent classification module, configured to generate a core intent set according to the multi-modal generation requirement information, the core intent set including at least one core intent information;

[0112] A confidence score module, configured to score the confidence of each core intent information;

[0113] A sub-intent recursive decomposition module, configured to decompose each core intent information respectively to obtain executable sub-intents, and one core intent information includes at least one sub-intent;

[0114] A hierarchical graph construction module, configured to generate and construct a hierarchical intent graph according to the core intent information and the sub-intents; wherein the nodes of the constructed hierarchical intent graph include the core intent information and the sub-intents, and the edges are dependency relationships;

[0115] a dynamic graph evolution and adjustment module configured to generate a dynamically adjusted intention graph according to the hierarchical intention graph;

[0116] an intention priority ranking and constraint generation module configured to adjust priorities and constraint conditions of each node in the dynamically adjusted intention graph, thereby generating an intention graph with priorities and constraint conditions.

[0117] In this embodiment, the intention classification module can perform the following operations:

[0118] The cross-modal aligned feature vector (from the demand acquisition module) is input into a "CLIP+BiLSTM" hybrid model, and classification is performed through a pre-trained intention category library (such as "scene category", "subject category", "action category", "special effect category", "style category").

[0119] In this embodiment, the confidence score module can evaluate through modal consistency, feature clarity, and model output probability. Specifically, the modal consistency is that if the text, vision, and audio all refer to the same intention (such as "red dragon" appearing in both text and vision), the score is improved; the feature clarity is that the score of an intention with unambiguous text and high vision resolution is high; and the model output probability is that the prediction probability of the classification model for the intention (such as the classification probability of CLIP) directly affects the score.

[0120] For example, the confidence score is calculated by multi-dimensional weighting, considering the cross-modal consistency of the intention, the quality of the associated material, and the certainty of the model classification. The specific rules are as follows:

[0121] Weight distribution: cross-modal feature alignment score accounts for 40% of the weight, material quality score accounts for 30% of the weight, and intention classification probability accounts for 30% of the weight (this weight ratio is verified through a large number of experiments, which can balance the multi-modal consistency and material reliability, and avoid single-dimensional deviation).

[0122] Score calculation: first, convert the 0-1 original score of each dimension to a standard score of 0-100 (such as alignment score 0.9→90 points, material quality score 0.8→80 points, classification probability 0.85→85 points), then calculate the weighted sum according to the weight, and finally the final result is kept to one decimal place.

[0123] The formula is as follows: confidence score = (cross-modal alignment standard score × 40%) + (material quality standard score × 30%) + (classification probability standard score × 30%).

[0124] In this embodiment, the sub-intention recursive decomposition module can perform the following operations:

[0125] Step 1: Core intent matching with decomposition knowledge base:

[0126] For each high-confidence core intent of the confidence scoring module, according to its category and description text, match the corresponding standard decomposition rules in the domain intent decomposition knowledge base. For example, the core intent "Dragon Fighting" (action category), matches the decomposition rule "fighting action -> attack action / defense action / move action" in the knowledge base, and determines the initial decomposition direction of the core intent as three types of action branches.

[0127] If a core intent has no complete matching rule in the knowledge base (such as the emerging style class intent "Cyberpunk Snow Mountain"), call the similar intent decomposition rule for adaptation, and mark it as "to be verified decomposition direction", and later combine with manual review to supplement and improve.

[0128] Step 2: First-level sub-intent generation:

[0129] Based on the matched decomposition rules, generate the first-level sub-intents of the core intent. For example, "Dragon Fighting", according to the decomposition direction "attack action / defense action / move action", generates three first-level sub-intents "Dragon Attack", "Dragon Defense", and "Dragon Battlefield Move". Each first-level sub-intent needs to clearly describe the core action direction, such as "Dragon Attack" needs to specify the active attack behavior against the enemy target, and assign a unique identifier to each first-level sub-intent, and associate the corresponding core intent ID.

[0130] Step 3: Second-level sub-intent decomposition:

[0131] For each first-level sub-intent, further decompose it into more detailed second-level sub-intents. For example, "Dragon Attack", combined with the "attack action" sub-rule in the domain knowledge base, decomposes into "fire spewing attack", "claw swinging attack", and "tail sweeping attack". In the decomposition process, it needs to judge whether the second-level sub-intent is close to the atomic operation, if a second-level sub-intent can initially extract some parameters (such as "fire spewing attack" can initially determine the spewing range), then focus on marking the parameter extraction direction, and prepare for the next step of atomic-level decomposition.

[0132] Step 4: Atomic-level sub-intent verification and determination:

[0133] For each secondary sub-intention, check against the atomic operation library to determine if it can be directly mapped to an atomic operation. If the secondary sub-intention "flame jet attack" can find the "special effect generation-flame jet" operation in the atomic operation library, and the jet duration (such as 2-3 seconds by default according to "battle scene") and flame color (such as orange-red by default according to "dragon red appearance") can be extracted from the intention description, then the secondary sub-intention is converted into an atomic sub-intention, marked as a leaf node; if the secondary sub-intention "claw swing attack" cannot directly extract complete parameters (such as swing amplitude is not clear), it is further decomposed into three atomic sub-intentions "claw lifting", "claw swinging forward", "claw retracting", each corresponding to the "limb movement-claw lifting", "limb movement-claw swinging", "limb movement-claw retracting" operation in the atomic operation library, and the execution time of each action is specified (such as lifting for 0.5 seconds, swinging for 1 second, and retracting for 0.5 seconds).

[0134] Fifth step: decomposition tree integration

[0135] Integrate the core intention, primary sub-intention, secondary sub-intention, and atomic sub-intention into a decomposition tree according to the hierarchical relationship, supplement the execution dependency conditions of each node (such as "claw swing" depends on the completion of "claw lifting"), and ensure that the decomposition tree structure is clear and logically coherent. At the same time, perform integrity verification on the decomposition tree to check if there are any core intentions that have not been decomposed or sub-intentions that have not been associated with atomic operations, if so, return to the previous step to re-decompose and adjust.

[0136] In this embodiment, the hierarchical graph construction module can perform the following operations:

[0137] First step: node information extraction and initialization:

[0138] From the hierarchical sub-intention decomposition tree, extract the unique identifier, description text, belonging level, and associated core intention ID of each intention one by one, and match the scores in the intention confidence list to integrate into the basic information of the graph node. The initial state of each node is set to "to be activated", and it is classified according to the belonging level (such as core intention is classified as top-level node, atomic operation is classified as bottom-level node), to ensure that there is no missing node information (such as nodes without associated core intention ID, need to backtrack the decomposition tree for supplement).

[0139] Second step: dependency relationship sorting and classification:

[0140] First, extract explicit dependency relationships from the "execution dependency conditions" of the decomposition tree (such as "claw waving depends on claw lifting"), and then combine the domain intent association knowledge base to mine implicit dependency relationships (such as "all battle-related sub-intents depend on the core intent 'dragon battle'" and "special effect sub-intents in the snow mountain scene need to depend on the'snow mountain environment' scene intent"). For each group of dependency relationships, classify them by "premise / timing / space" type, such as "dragon exists → fire spewing" as premise dependency, "claw lifting → claw waving" as timing dependency, and "dragon stepping on → snow flying" as spatial dependency.

[0141] Step 3: Dependency strength weight calculation:

[0142] For each group of classified dependency relationships, calculate the dependency strength weight: first take the average of the confidence scores of the two associated intents (such as A confidence 90 points, B confidence 85 points, average 87.5 points, converted to 0.875), then adjust according to the association coefficient of this type of dependency in the domain knowledge base (if it is a strong association, add 0.05, if it is a weak association, subtract 0.05, if there is no clear association, do not change), finally constrain the result between 0-1 (such as 1.02, take 1.0, 0.03, take 0.03). For example, "dragon exists → fire spewing", average 0.875, knowledge base for strong association plus 0.05, maximum weight 0.925; "snow flying → rock collapse", average 0.75, no clear association, weight 0.75.

[0143] Step 4: Graph structure building and visual mapping:

[0144] Build the graph according to the layout rules of "hierarchy from top to bottom, dependency from left to right": the top core intent node is in the center, the first level sub-intent nodes are distributed around the core node, the second level sub-intent nodes are under the corresponding first level sub-intent nodes, and the fourth level atomic operation nodes are under the corresponding second level sub-intent nodes. Connect the corresponding nodes with "edges" according to the dependency relationships, and adjust the thickness of the edges according to the dependency strength weight (the higher the weight, the thicker the edge), and label the dependency type and description beside the edge (such as "premise dependency: fire spewing needs dragon existence"). At the same time, add color to the nodes to distinguish the types (core intent red, first level sub-intent blue, atomic operation green), and ensure that the graph is intuitive and distinguishable.

[0145] Step 5: Graph information verification and supplement:

[0146] Check the integrity of the graph: confirm that all intentions extracted from the decomposition tree have been converted into nodes without missing; all dependencies teased out have been converted into edges without missing. Check the logic for reasonableness: ensure that there is no circular dependency (e.g., A depends on B, B depends on A), no abnormal dependency across levels (e.g., atomic operation node depends on core intent node, need to confirm whether it meets the knowledge base rules), if there is a problem, backtrack to the second step to re-tease the dependency relationship. Finally, add a unique identifier to each node and edge to complete the graph structure integration.

[0147] In this embodiment, the dynamic graph evolution and adjustment module can perform the following operations:

[0148] First step: feedback information analysis and classification:

[0149] For the input real-time feedback information, extract the "associated intent ID, feedback type, adjustment suggestion" one by one, and classify and archive according to "user feedback / module verification feedback / resource feedback". For example, user feedback "add night scene" (associated ID: none, type: add, suggestion: add "night scene" core intent), physical engine feedback "dragon mass exceeds standard" (associated ID: INT-Subject-001, type: conflict, suggestion: reduce dragon mass parameter), resource feedback "particle special effect insufficient" (associated ID: INT-Special Effect-001, type: resource limitation, suggestion: simplify special effect). At the same time, check the completeness of the feedback information, if the associated ID or adjustment suggestion is missing, return to the previous step to supplement (such as "semantic conflict" feedback without associated intent, need to prompt the relevant module to supplement).

[0150] Second step: matching evolution rules and generating preliminary adjustment scheme:

[0151] For each type of feedback, match the corresponding adjustment strategy in the graph evolution rule library and generate a preliminary scheme. For example, user adds an intent feedback, matches the "add node rule", determines the node level ("night scene" is a core intent, belongs to the top level), the initial confidence (user's clear demand, set to 90 points), and the initial state (to be activated); physical conflict feedback, match "conflict handling rules", reduce the confidence of the associated node "dragon" (from 90 points to 80 points), and mark the node state as "conflict to be repaired"; resource limitation feedback, match "resource compromise rules", change the "particle special effect" sub-intention to "simplified flame effect", adjust the node description and reduce the confidence (from 85 points to 70 points).

[0152] Third step: scheme feasibility check and conflict investigation:

[0153] For the preliminary adjustment scheme, check from two aspects of "logical feasibility" and "resource feasibility": logical feasibility needs to check whether the newly added dependency meets the hierarchical rules (such as the top-level node cannot depend on the bottom-level node), and whether the deletion of the node affects the execution of the core intent (such as deleting "dragon exists" will cause all battle sub-intents to fail, so the deletion scheme needs to be vetoed); resource feasibility needs to combine the current system resources (such as computing power, storage) to judge whether the adjusted intent can be executed (such as whether the computing power demand of "simplify flame effect" is within the range of graphics card bearing). If a conflict is found (such as the newly added "night scene" conflicts with the existing "snow mountain daytime" scene), conflict negotiation is started, and the "night scene" is preferentially retained, and the "snow mountain daytime" node and related edges are deleted.

[0154] Fourth step: actual adjustment of graph nodes and edges

[0155] According to the scheme passed by the check, the specific adjustment operation is performed:

[0156] Node adjustment: add node (insert into graph according to hierarchy, such as "night scene" added to top-level core node), delete node (remove node and all references corresponding to associated ID), state change (such as "conflict to be repaired" -> "repaired"), confidence modification (update score in node information);

[0157] Edge adjustment: add edge (connect new node and associated node, such as "night scene" -> "moonlight" prerequisite dependency, weight 0.9), delete edge (remove all edges related to deleted node), weight recalculation (such as "dragon" confidence reduced to 80 points, associated edge "dragon -> flame jet" weight recalculated: average value (80+85) / 2 = 82.5 -> 0.825, strong association + 0.05 -> 0.875). During the adjustment process, real-time operation logs (including time, operation content, feedback ID) are recorded to ensure traceability.

[0158] Fifth step: verification and confirmation of the adjusted graph

[0159] Check the integrity of the adjusted graph: whether the newly added node has added all necessary information (ID, type, confidence, etc.), whether the deleted node and edge have been completely removed, and whether the weight adjustment is within the range of 0-1. Check logical consistency: no new circular dependency (such as "night scene -> moonlight" and "moonlight -> night scene"), node state and feedback match (such as conflict repaired node state updated). Finally, the adjusted graph is associated with the feedback information and archived to generate a "graph evolution report" (including before and after comparison, key change explanation), which is submitted to the next module.

[0160] In this embodiment, the intent priority sorting and constraint generation module can perform the following operations:

[0161] Step 1: Input data collation and preprocessing:

[0162] Extract the ID, type, description, confidence, and state of all nodes, as well as the dependency type and weight of the edges from the dynamically adjusted intent graph. Integrate user demand preferences (e.g., "core is dragon fighting") and system resource configurations (e.g., 16GB GPU, 32GB memory). Retrieve the weight of each dimension from the priority evaluation index library (demand core degree 40%, confidence 30%, etc.) to ensure that all input data is complete. If a node is missing confidence (e.g., a newly added node is not labeled), set it to 80 points if the user explicitly demands it, or 60 points if the demand is not explicit, to avoid evaluation bias.

[0163] Step 2: Intent node priority calculation and sorting:

[0164] Multi-dimensional score calculation: Calculate the four-dimensional score for each node one by one. Take "dragon fighting" (core intent, confidence 90 points, user explicitly core demand, 10 strong dependency nodes with dependency weight above 0.9, resource consumption 40%) as an example:

[0165] Demand core degree: User explicitly core, 100 points (40% weight → 40 points);

[0166] Confidence: 90 points (30% weight → 27 points);

[0167] Dependency tightness: 10 strong dependency nodes, score increased by 20 points (20% weight → (90+20) × 20% = 22 points);

[0168] Resource consumption rationality: Consumes 40% of the computing power, 60 points (10% weight → 6 points);

[0169] Total score: 40 + 27 + 22 + 6 = 95 points;

[0170] Priority level mapping: 95 points correspond to priority level 10, labeled "priority 10, sorting reason: user core demand + high confidence + multiple nodes strong dependency, resource consumption is reasonable";

[0171] Priority correction: For nodes with similar scores (e.g., "night scene" 92 points, "simplified flame effect" 88 points), set "night scene" priority to 10 and "simplified flame effect" to 9 to ensure that the dependency relationship and priority are matched;

[0172] Global sorting: Sort by priority level from high to low (10 level → 9 level → … → 1 level), and arrange the same level in descending order of score. Generate "priority sorting table" and label the ID, level, score, and reason of each node.

[0173] Step 3: Multi-type constraint condition generation:

[0174] Temporal constraint generation:

[0175] Extract temporal dependency edges in the graph (e.g., "night scene → moonlight", "dragon exists → claw raises");

[0176] Combine video length 15 seconds, frame rate 30fps, set specific time difference, "night scene" (priority level 10, activation time 0 seconds) → "moonlight" (priority level 9, activation time 0-2 seconds, corresponding to 0-60 frames); "dragon exists" (priority level 10, activation time 3 seconds) → "claw raises" (priority level 8, activation time 4 seconds, corresponding to 91-120 frames);

[0177] Assign an ID to each temporal constraint (e.g., "temporal constraint-night scene-moonlight-001"), label associated node ID, time range, and violation consequences (e.g., "more than 2 seconds without activating moonlight, automatically trigger scene light compensation mechanism");

[0178] Resource constraint generation:

[0179] Statistical resource consumption of each node (e.g., "dragon action calculation" 40% computing power, "simplified flame effect" 30% computing power, "moonlight" 5% computing power);

[0180] Combine system computing power upper limit 100%, set constraints: "dragon action calculation" computing power occupancy ≤45%, "simplified flame effect" ≤35%, "moonlight" ≤8%, reserve 12% dynamic space;

[0181] Label resource monitoring threshold (e.g., "simplified flame effect" occupancy exceeds 35% to automatically reduce particle number), generate "resource constraint table";

[0182] Style constraint generation:

[0183] Extract style parameters from "night snow mountain fantasy scene" node (main color dark purple + silver, saturation 75%, light intensity 200 lux);

[0184] Convert to constraints: "all scene elements main color RGB30-80, saturation 60%-80%, light intensity 150-250 lux; prohibit using RGB200 above bright color, prohibit realistic style texture";

[0185] Associate all visual nodes (e.g., "snow mountain", "dragon", "moon");

[0186] Step 4: Integration of priority and constraint graph:

[0187] The level, score, and reason in the "priority ranking table" are supplemented into the corresponding nodes of the dynamic intention graph, and each node adds a "priority information" field; the constraint conditions in the "timing / resource / style constraint table" are associated with the graph nodes, and a "constraint ID" field is added in the edge set (such as associating the corresponding timing constraint ID with the timing dependent edge), forming an "intention graph with priority and constraint conditions". At the same time, "priority and constraint description documents" are added to the graph to explain the sorting logic and constraint basis (such as "the 'night scene' priority is 10 levels because it is a multi-node premise" and "the reference GPU 16GB configuration is used for power constraint"), which is convenient for subsequent modules to understand and execute.

[0188] Fifth step: verification and confirmation of priority and constraint

[0189] Priority verification

[0190] Check whether the priority of the core node (such as "dragon fight" and "night scene") is 10 levels, whether the priority of the auxiliary node (such as "moonlight") is 9 levels or above, and whether the priority of the edge node (such as "distant flying bird") is 5 levels or below, which meets the "core → auxiliary → edge" level logic;

[0191] Verify the consistency of the dependency relationship and the priority (such as the priority of the premise node being higher than that of the dependent node), and there is no "dependent node priority higher than premise node" exception;

[0192] Constraint verification

[0193] Check whether the timing constraint time range is reasonable (such as "moonlight" activated within 0-2 seconds meeting the scene introduction logic, and there is no "time range exceeding video duration" error);

[0194] Check whether the total resource constraint is ≤ system upper limit (45% + 35% + 8% +… = 88% ≤ 100%), and whether the reserved space is sufficient;

[0195] Check whether the style constraint covers all visual nodes, and whether the parameters are quantifiable (such as "RGB 30-80" is clear, and there is no "style beautiful" or other ambiguous expression);

[0196] Final confirmation: generate a "priority and constraint verification report" to record the verification results (such as "priority ranking meets the logic, and constraint conditions are executable"), and submit the intention graph with priority and constraint to the multi-scale spatiotemporal capsule scene generation module.

[0197] In this embodiment, the multi-scale spatiotemporal capsule scene generation module comprises:

[0198] A macro-scale scene layout generation module, which is configured to generate a macro scene capsule according to the intention graph with priority and constraint conditions;

[0199] Mesoscale role action and special effect generation module, the mesoscale role action and special effect generation module is used for generating mesoscale action and special effect capsules according to the macroscopic scene capsule and demand information;

[0200] Microscale detail texture and light and shadow generation module, the microscale detail texture and light and shadow generation module is used for generating microscale detail and light and shadow capsules according to the macroscopic scene capsule, mesoscale action and special effect capsule and demand information;

[0201] Spacetime capsule routing and cross-scale connection module, the spacetime capsule routing and cross-scale connection module is used for fusing the macroscopic scene capsule, mesoscale action and special effect capsule and microscale detail and light and shadow capsule into a spacetime consistency capsule set;

[0202] Multi-scale content fusion and preliminary rendering module, the multi-scale content fusion and preliminary rendering module is used for capsule fusion and preliminary rendering to the spacetime consistency capsule set, so as to obtain a preliminary rendered video frame sequence, and the preliminary rendered video frame sequence is used as multi-scale spacetime capsule modeling information.

[0203] In the embodiment, the macroscopic scale scene layout generation module can perform the following operations:

[0204] First step: scene related intention and constraint extraction:

[0205] From the intention atlas with priority and constraint conditions, the intention (high priority ensures that the core element is not missed) of which the type is scene class / subject class and the priority is greater than or equal to 8 is filtered out, and the description text, constraint condition and space dependency of each intention are extracted. For example, the “night snow mountain scene” (scene class, priority 10, constraint: “contains main peak, valley, sky moon, style fantasy dark tone”) and “red dragon” (subject class, priority 9, constraint: “moves around the main peak of the snow mountain, large body size”) are extracted, and the space dependency of “moonlight” (special effect class, dependent on the sky area of “night snow mountain”) is recorded. If a constraint condition is ambiguous (such as “large body size” without specifying the size), the default value is matched from the scene basic parameter library (such as “the default size of the large body is 1 / 3-1 / 2 of the height of the core element of the scene”).

[0206] Second step: space layout construction:

[0207] Determine the overall spatial range: Based on the default parameters of the scene class core intent, combined with constraint adjustments, such as the "snow mountain scene" default 1000m x 800m x 500m (length x width x height), the width is expanded to 900m due to the constraint requirement "contains valley", ensuring that the valley area has enough space; At the same time, define a three-dimensional coordinate system (such as taking the lower left corner of the scene as the origin (0, 0, 0), x as length, y as width, and z as height), which is convenient for subsequent labeling of position coordinates.

[0208] Divide the core area: According to "core function area → interaction area → background area", such as the (400-700, 300-600, 200-500) coordinate range of "night snow mountain" is defined as the main peak combat core area (main action core area) (700-900, 300-600, 100-300) is defined as the valley interaction area (main body and scene interaction area), and the remaining area is the background transition area (such as the outline of the snow mountain in the distance, the edge of the sky); Each area is labeled with function positioning and spatial constraints (such as "no background elements blocking the main body" in the core area, "need to reserve snow splashing space" in the interaction area).

[0209] Third step: Time axis layout construction:

[0210] According to the video length and frame rate requirements, divide the scene stages and define the state of each stage. For example, the video length is 15 seconds, the frame rate is 30fps (a total of 450 frames):

[0211] Opening stage (0-3 seconds, 0-90 frames): Scene introduction, state definition as "from long shot to main peak panorama, moon slowly rises from the right side of the sky, snow mountain outline gradually appears with moonlight, no main body element";

[0212] Combat stage (3-12 seconds, 91-360 frames): Core interaction, state definition as "main body dragon flies into the core area from the left side of the main peak, scene dynamically responds with dragon action (such as wing flapping brings airflow, mountain body slightly shakes), moon maintains stable lighting";

[0213] End stage (12-15 seconds, 361-450 frames): Transition freeze frame, state definition as "dragon stays on the top of the main peak, wings slowly fold, moon halo enhances, scene long shot gradually darkens, finally freezes in the dragon and main peak panorama";

[0214] Each stage is labeled with key frame nodes (such as 90 frames "moon rises to the center of the sky", 360 frames "dragon reaches the top of the main peak"), ensuring that the time axis is synchronized with subsequent action generation.

[0215] Fourth step: Core element layout and style adaptation:

[0216] Element Position and Size Determination: According to the core element layout rules, mark the coordinates and sizes, for example, the center point coordinates of the main peak of the snow mountain (550, 450, 350), height 300m; the initial position of the dragon (300, 450, 250) (left side of the main peak), body length 80m (about 1 / 3 of the height of the main peak); the position of the moon (550, 100, 450) (sky area, aligned with the center of the main peak), diameter 50m; all coordinates need to ensure that the elements do not exceed the scene space range, and the size ratio is coordinated.

[0217] Style Adaptation Adjustment: According to the style constraints in the intention graph (such as "fantasy dark tone"), add style processing to the core elements, add light purple ice crystal texture to the surface of the snow mountain, cover the top of the main peak with silver-white snow; add a light blue halo to the edge of the moon, with a range of 2 times the diameter of the moon; add dark gold reflection to the dragon scales, echoing the scene tone; at the same time, determine the scene tone parameters (main tone RGB(30, 40, 80), saturation 75%, brightness 30%), ensure the style is unified.

[0218] Step 5: Global Constraint Standard Formulation and Capsule Integration:

[0219] Formulate Global Constraints: Integrate spatial constraints, style constraints, and lighting constraints, for example, spatial constraints: "the activity range of the dragon is limited to (300-800, 300-600, 200-400) coordinates"; style constraints: "all scene elements cannot use bright colors above RGB(200, 200, 200), special effect colors need to be derived from the main tone"; lighting constraints: "the moon is the only main light source, with a light intensity of 200 lux, the ground object shadow direction is (-1, -1, 0) (corresponding to the left-up to right-down direction), and the shadow length is 2 times the height of the object".

[0220] Macroscopic Scene Capsule Integration: Integrate spatial layout, time axis layout, core element layout, and global constraint standards into a structured "macroscopic scene capsule", mark the unique identification of the capsule (such as "macroscopic capsule-night snow mountain-001"), generate time and associate intention graph ID, submit to the mesoscale role action and special effect generation module, and at the same time, keep the capsule generation log (including parameter adjustment basis, constraint source) for subsequent tracing.

[0221] In this embodiment, the mesoscale role action and special effect generation module operates as follows:

[0222] First Step: Input Data Analysis and Requirement Mapping:

[0223] Extract the spatial range from the macro-scene capsule (such as the core area (400-700), the time axis stage (3-12 seconds of the battle stage), the initial position of the character (300, 450, 250), and the style constraint (fantasy dark tone); From the user's multi-modal requirements, sort out the character action requirements (such as "the dragon battle needs to include wing flapping, claw waving, and flame spewing"); From the action and effect base library, call the default action template of "dragon" (such as wing flapping frequency 2 times per second) and the "fantasy dark tone" effect template (such as flame color RGB (160, 70, 40)), preliminarily associate the requirements with the templates, and clearly generate the direction (such as "wing flapping" reference template parameters, and "flame spewing" adjust the color according to user requirements).

[0224] Second step: Character action sequence generation:

[0225] Take the red dragon as an example, generate actions according to the time axis stage:

[0226] Pre-battle action (3-4 seconds, 91-120 frames): Corresponding to the macro-scene "dragon flies into the core area", you can refer to the "dragon flight" template in the action base library (which can be preset in advance), combine the initial position (300, 450, 250) and the core area entrance (400, 450, 250), and design the "wing flapping flight" action: flapping frequency 2 times per second, each flapping angle from -35° to +35°, flight speed 10 m / s (1 second from the initial position to the core area entrance), action parameter label "action ID: ACT-dragon-001, time 91-120 frames, wing flapping angle -35°~+35°, flight speed 10 m / s";

[0227] In-battle action (4-10 seconds, 121-300 frames): Corresponding to "battle attack", according to user requirements "wave claws, flame spewing", design action sequence, for example, 4-4.5 seconds (121-135 frames): "claw lifting" action, claw from below the body to 60° position in front, lifting speed 0.5 m / s, transition from "wing flapping" action, transition time 0.2 seconds;

[0228] 4.5-5 seconds (136-150 frames): "claw waving attack" action, claw from 60° position to 30° position below, waving speed 2 m / s, accompanied by "roar" action (mouth opening amplitude 50 cm);

[0229] 5-7 seconds (151-210 frames): "flame spewing" action, mouth keeps open state, spewing lasts for 2 seconds, action parameters are associated with subsequent effect trigger conditions;

[0230] Each action is annotated with physical constraints (e.g., claw waving speed should not exceed 2.5 m / s to avoid exceeding the character's muscle load);

[0231] Step 3: Scene special effect sequence generation:

[0232] According to the character's actions and macro scene constraints, generate matching special effects, for example:

[0233] "Air flow special effect": trigger condition is "wing flapping" action, time interval is synchronized with wing flapping (3-12 seconds), reference special effect base library "fantasy air flow" template, adjust style parameters: air flow color is light purple (RGB(80,80,150), matching dark tone), particle density is positively correlated with flapping intensity (flapping frequency 2 times per second, density 300 per frame, 0.5 times per second, density 100 per frame), spatial range is around the wings 5-8m area (not exceeding the size of the character), special effect ID is labeled "EFF-air flow-001", associated action ID is "ACT-dragon-001";

[0234] Step 4: Spatio-temporal adaptation of action and special effect, for example:

[0235] Spatial adaptation: check whether the spatio-temporal position of each action and special effect is within the macro scene range, such as "flame spewing" length 30m, starting point (350,450,270), end point (380,450,270), all within the core area (400-700,300-600,200-500), conforming to the spatial constraints; if the "snow splashing" range exceeds the interaction area, adjust the splashing range to 8m to ensure that it does not exceed the macro scene partition;

[0236] Temporal adaptation: ensure that the time interval of action and special effect is synchronized with the macro time axis, such as "dragon flying into core area" action 3-4 seconds, corresponding to the beginning of the macro battle stage, "stopping" action 12 seconds (360 frames) completely synchronized with the macro key frame; if the "flame spewing" special effect lasts for 2 seconds (5-7 seconds), it does not exceed the battle stage (3-12 seconds), the timing is reasonable;

[0237] Dynamic adaptation: adjust the rhythm of action and special effect to ensure coherence and naturalness, such as "claw lifting" (4-4.5 seconds) and "claw waving" (4.5-5 seconds) without pause, transition time 0.2 seconds, smooth action connection; "flame spewing" special effect triggers 0.1 seconds after action starts, simulating the natural delay of "spewing fire after opening mouth", avoiding disconnection between special effect and action.

[0238] Step 5: Integration of meso-action and special effect capsules:

[0239] Integrate the action sequence of the character with the scene special effect sequence, label the unique identification of the capsule (such as "mesoscopic capsule-dragon fight-001"), and clearly indicate the association between each action and special effect (such as "ACT-dragon-003 (flame spitting action) is associated with EFF-flame-001 (flame spitting special effect)"), and supplement "adaptation instructions" (such as "the action fan frequency is referenced to the size of the dragon, and the special effect color is matched to the macroscopic dark tone"). After integration, submit to the micro-scale detail texture and light generation module, while retaining the generation log (including action parameter adjustment basis, special effect template modification record) for subsequent tracing and adjustment.

[0240] In this embodiment, the micro-scale detail texture and light generation module operates as follows:

[0241] First step: input data analysis and detail requirement mapping:

[0242] Extract the global illumination (moon direction from top left to bottom right, intensity 200 lux), style constraints (fantasy dark tone), and element size (dragon length 80m, snow mountain main peak 300m) from the macroscopic scene capsule; extract the character action parameters and special effect parameters from the mesoscopic action and special effect capsule; retrieve matching templates (dragon scale metal texture template, snow mountain ice crystal semi-transparent texture template, etc.) from the micro-scale detail material library; extract the detail requirements (for example, "scale dark gold reflection") from user preferences, and integrate these information to clearly indicate the detail generation direction of each element (such as adding dark gold reflection to dragon scales, and ice crystal containing light purple lines).

[0243] Second step: detail texture generation and dynamic adaptation:

[0244] Generate textures in the order of "character elements -> scene elements".

[0245] Third step: dynamic light generation and timing synchronization:

[0246] Generate in the order of "global light -> local light -> dynamic adjustment", and ensure synchronization with macroscopic and mesoscopic.

[0247] In this embodiment, the spatiotemporal capsule routing and cross-scale connection module operates as follows:

[0248] First step: capsule information analysis and index establishment:

[0249] Extract the core parameters of each scale: extract the global coordinate system, scene partition coordinates, time axis stage, and key frame from the macroscopic capsule; extract the ID, local coordinates, time interval, and size parameters of the action / special effect from the mesoscopic capsule; extract the ID, associated elements, dynamic rules, and spatiotemporal parameters of the texture / light from the micro-scale capsule;

[0250] Establish a cross-scale index table: according to the hierarchy of "macro-scene partitioning → meso-action / special effect → micro-texture / lighting", assign a unique associated ID to each element, and ensure that each meso and micro element can be traced back to the corresponding macro partition.

[0251] Second step: space coordinate unification and range verification:

[0252] Meso-coordinate mapping: convert the local coordinates of meso-action / special effect to global coordinates, such as the local coordinates of meso "dragon wing flapping" (-50~+50, 0, 0) (with the dragon's head as the origin), and the dragon's position in the macro at 3-3.5 seconds is (350, 450, 260), then the global coordinates = local coordinates + macro position, i.e. the flapping range of the wing is (300-400, 450, 260);

[0253] Macro range verification: check if the mapped meso coordinates fall within the corresponding macro partition, such as the flapping range of the wing (300-400, 450, 260) corresponding to the macro core area (400-700, 300-600, 200-500), the left boundary 300-400 exceeds the left boundary 400 of the core area, the deviation is 0-100m, the dragon's macro position needs to be adjusted to (450, 450, 260), so that the flapping range of the wing becomes (400-500, 450, 260), completely falling within the core area;

[0254] Micro-coordinate binding: bind the coordinates of micro-texture / lighting to the macro coordinates of meso-action / special effect, such as micro "scale texture" associated with meso wing, wing macro range (400-500, 450, 260), then the scale texture is only generated within this coordinate range.

[0255] Third step: time axis synchronization and trigger linkage:

[0256] Frame-level time conversion: convert all time intervals of different scales to frame units (30fps);

[0257] Time sequence deviation verification: check the synchronization of time intervals between different scales, such as micro "snow splashing" special effect original time 4.6-5.1 seconds = 139-154 frames, meso "claw waving" action 4.5-5 seconds = 136-150 frames, the deviation is 3 frames, the special effect time needs to be adjusted to 4.5-5 seconds = 136-150 frames, completely synchronized with the action;

[0258] Trigger link establishment: Establish time trigger link in the order of "action-> special effect-> details-> light and shadow", such as the action of the macro "dragon mouth opening" at 4.8 seconds = 144 frames, delay 3 frames (0.1 seconds) to trigger the macro "flame jet" at 147 frames, delay 3 frames (0.1 seconds) to trigger the micro "flame lighting" at 150 frames, and the micro "snow mountain texture brightness change" in the lighting range is started at 150 frames, forming a complete time linkage chain, and recording the frame number and delay time of each trigger node.

[0259] Fourth step: dynamic dependent routing and consistency adjustment:

[0260] Top-down constraint transmission: transmit macro constraints to meso and micro, such as macro "moonlight direction left up to right down", transmit to meso "dragon shadow direction right down", and meso shadow parameters (length 160m, blur radius 2m) are transmitted to micro "shadow texture details" (add 1m wide gradient blur to shadow edge, match blur radius);

[0261] Bottom-up problem feedback: feedback micro resource problems to meso adjustment, such as micro "4K scale texture total size 120MB, over memory limit 100MB", feedback to meso, reduce scale texture resolution to 2K in non-core area (wing edge), total size to 80MB, ensure resource adaptation;

[0262] Peer collaborative adjustment: synchronize dynamic parameters of meso and micro, such as meso "wing fan angle from -35°->+35°" (91-105 frames), synchronize micro "scale stretch ratio from 10%->15%", each frame fan angle change 1°, corresponding scale stretch ratio change 0.25%, ensure action and detail dynamic rhythm consistent.

[0263] Fifth step: capsule integration:

[0264] Capsule set integration: integrate "macro partition-meso action / special effect-micro texture / light and shadow" unit into spatiotemporal consistency capsule set, each unit contains association index, unified spatiotemporal parameter, dynamic dependent link and verification mark; For example, "core area unit" contains: association index (LINK-core area-ACT001-TEX001), unified spatiotemporal (400-700, 300-600, 200-500, 91-360 frames), dependent link (macro core area->meso wing fan->micro scale stretch), and verification mark (pass).

[0265] In this embodiment, the multi-scale content fusion and preliminary rendering module specifically operates as follows:

[0266] First step: spatiotemporal consistency capsule set preprocessing and hierarchical loading:

[0267] After receiving the spatiotemporal consistency capsule set, the data is classified according to the "macroscopic scene layer → mesoscopic dynamic layer → microscopic detail layer", for example, the macroscopic layer extracts the spatial range of the core area (400-700, 300-600, 200-500), the global illumination (moonlight 200 lux, direction left up to right down), the correlation rules and parameters of the time axis key frame (90 / 360 frames), and ensures that there is no omission of data in each layer.

[0268] Layered loading and resource allocation: adopt the "priority loading" strategy, first load the macroscopic scene layer data (such as the main peak of the snow mountain and the basic model of the sky background), allocate 20% of the memory to store the scene framework; then load the mesoscopic dynamic layer (dragon skeleton model, flame particle special effect template), allocate 40% of the memory to carry dynamic data; finally load the microscopic detail layer (scale texture map, light and shadow calculation shader), allocate 30% of the memory to store detail resources, and reserve 10% of the memory for real-time calculation, to avoid resource conflicts caused by chaotic loading order (such as loading microscopic textures without mesoscopic models to carry).

[0269] Second step: multi-scale content space fusion and dynamic binding:

[0270] Macroscopic-mesoscopic space binding: anchor the mesoscopic dynamic elements to the corresponding positions of the macroscopic scene, for example, according to the unified parameter "dragon position 450, 450, 260" in the capsule set, load the mesoscopic dragon skeleton model to the (450, 450, 260) coordinate point of the macroscopic coordinate system, to ensure that the model center point and the parameter are completely coincident.

[0271] Mesoscopic-microscopic detail attachment: bind the microscopic texture to the specific parts of the mesoscopic dynamic elements, according to the rule "scale texture associated with wings", accurately paste the 4K scale texture map to the triangular face model of the mesoscopic dragon wing, use "texture UV animation" technology to make the scale texture stretch synchronously with the wing flapping angle (-35° ~ +35°), and the stretching ratio increases from 10% to 15%, dynamically matching the mesoscopic action.

[0272] Cross-scale space conflict detection: traverse the fused space data to check whether there is an element overlap conflict, for example, whether there is a "flame penetrating the snow mountain" conflict between the mesoscopic flame spouting range (440-470, 450, 270-300) and the microscopic snow mountain texture range (400-700, 300-600, 200-500), judge the distance between the flame particle and the snow mountain model through the collision detection algorithm, if the particle distance from the snow mountain surface ≤0.5m, automatically adjust the flame spouting direction (from horizontal to upward 10°), to ensure the rationality of the space logic.

[0273] Third step: multi-scale content time fusion and time sequence synchronization:

[0274] Frame-level timing unified scheduling: Based on the macroscopic time axis (30fps), allocate frame-level execution instructions for each scale dynamic element, such as starting the mesoscopic wing fan action in 91 frames (3 seconds), synchronously triggering the microscopic scale texture stretching calculation, and ensuring that the timing is completely synchronized with the capsule set.

[0275] Dynamic rhythm coordinated adjustment: According to the rhythm requirements of the macroscopic time axis stage, optimize the dynamic parameters of the mesoscopic and microscopic, for example, the macroscopic opening stage (0-3 seconds, 0-90 frames) rhythm is slow, the mesoscopic dragon flying into speed is reduced from 10 m / s to 5 m / s, and the microscopic scale texture stretching rate is also reduced (from 0.25% / frame to 0.1% / frame).

[0276] Timing conflict repair: If it is detected that there is a timing deviation in multi-scale dynamics (such as a 2-frame delay in the start of microscopic flame lighting), adjust it through the "frame compensation" mechanism.

[0277] Fourth step: global lighting fusion and style unified rendering:

[0278] Multi-light source superposition calculation: Integrate macroscopic main light source and microscopic local light source to generate global lighting field.

[0279] Style rendering parameter configuration: According to the macroscopic "fantasy dark tone" constraint, set the global rendering parameters to ensure the uniformity of the scene style.

[0280] Layered rendering and layer synthesis: Use the "from far to near" layered rendering strategy for rendering, so as to form a complete single frame picture.

[0281] Fifth step: preliminary rendering output and output

[0282] Full-frame sequence rendering: Execute the rendering process frame by frame according to the macroscopic time axis (0-15 seconds, 0-450 frames), and store each completed frame as a PNG format (with Alpha channel, for subsequent optimization). During the rendering process, real-time monitoring of resource occupation (such as GPU computing power, memory usage) is performed. If the computing power occupation of a certain frame exceeds 90% (such as the peak value of 180 frames of flame spraying), the rendering precision of non-core elements of that frame is automatically reduced (such as reducing the texture of distant rocks from 2K to 1K), to ensure smooth rendering (frame rate stable at 30fps).

[0283] Integrate the qualified PNG frame sequence (0-450 frames) into "preliminary rendered video frame sequence", label the sequence unique identifier, and submit the frame sequence to the physical-semantic joint constraint verification module for subsequent verification and optimization.

[0284] In this embodiment, the physical-semantic joint constraint verification module includes:

[0285] a joint constraint field construction module configured to generate a joint constraint field model according to the multi-scale spatio-temporal capsule modeling information and the intention graph with priorities and constraints;

[0286] a simulated annealing optimization module configured to optimize the joint constraint field model by a simulated annealing algorithm to obtain optimized generation parameters;

[0287] a real-time physics engine verification module configured to verify the optimized generation parameters to generate physical violation labels and correction suggestions;

[0288] a real-time semantic discriminator verification module configured to perform semantic scoring and evaluation on the optimized generation parameters to obtain semantic violation labels and correction suggestions;

[0289] a generation parameter adjustment module configured to adjust the optimized generation parameters according to the physical violation labels and correction suggestions and the semantic violation labels and correction suggestions to obtain adjusted generation parameters;

[0290] a re-rendering module configured to adjust the multi-scale spatio-temporal capsule modeling information according to the adjusted generation parameters to obtain verified modeling information.

[0291] In this embodiment, the physical-semantic joint constraint verification module can perform the following operations:

[0292] Physical constraint subfield construction: based on the physical verification parameters and the intention constraints, a physical law constraint model is built, for example, in the "Dragon Snow Mountain Battle" scene, a gravity constraint (gravity acceleration 9.8 m / s 2 , allowing the fantasy scene to be adjusted to 8 m / s 2 ), a collision constraint (impact force ≤30 MPa when the dragon collides with the snow mountain) and other constraints are constructed;

[0293] Semantic constraint subfield construction: based on the semantic verification parameters and the intention constraints, a semantic logic constraint model is built, for example, a style constraint subfield is constructed (all elements need to be derived from RGB(30,40,80) in tone, and the appearance of realistic style metal texture is prohibited), a correlation constraint subfield is constructed (flame special effects need to be always associated with the dragon's mouth coordinates with a deviation of ≤5m, and the snow mountain texture needs to maintain a fantasy mixed texture of ice crystals and rocks), and each semantic subfield establishes a matching standard through semantic vector coding (for example, "fantasy dark flame" is coded as a 768-dimensional vector by using the CLIP model).

[0294] The physical constraints are associated with semantic constraints to form a unified joint constraint field, for example, in a "flame colliding with snow mountain" scene, the physical constraint requires "flame temperature causing snow mountain to melt", and the semantic constraint requires "melting effect to conform to fantasy style (such as generating purple steam instead of realistic white steam)", in the joint constraint field, it is specified that "snow mountain melting rate needs to meet the physical temperature difference rule (physical sub-field), and the steam color needs to be RGB (100, 100, 180) (semantic sub-field)", and the weights of the two are defined (physical constraint weight 0.4, semantic constraint weight 0.6, because the semantic priority of the fantasy scene is slightly higher), to avoid verification bias caused by single-dimensional constraints.

[0295] Optimization by simulated annealing and parameter iteration adjustment:

[0296] Initial parameter deviation calculation: the modeling parameters corresponding to the preliminary rendering frame sequence are substituted into the joint constraint field, and the parameter deviation value is calculated, such as the calculation value of the lift of the dragon's wing flapping: 0.5×1.2×5 2 ×1200×1.2=21600N, which is far lower than the required 64000N, the physical deviation value is (64000-21600) / 64000=0.6625(66.25%), and all parameters with a deviation greater than 20% (such as the lift deviation corresponding to the wing flapping stage of frames 91-105) and the corresponding frame number are recorded.

[0297] Simulated annealing algorithm optimization: with "minimizing joint deviation value" as the goal, start simulated annealing optimization - set initial temperature T=100 (control the probability of accepting deviation), 100 iterations, adjust 1-2 parameters with a deviation greater than 20% in each iteration (such as the first iteration increases the wing flapping speed from 5m / s to 8m / s, recalculates the lift = 0.5×1.2×8 2 ×1200×1.2=55296N, the deviation is reduced to (64000-55296 / 64000=0.136(13.6%), which meets the threshold; the second iteration adjusts the flame injection starting coordinates by 3m towards the dragon's mouth, the deviation is reduced to 5m, which meets the semantic constraint); after each iteration, reduce the temperature T=0.9×T, until T<1 or the deviation value<5%, stop iteration, output the optimized generated parameters (such as wing flapping speed 8m / s, adjusted value of flame starting coordinates).

[0298] Real-time physics engine verification: input the optimized generated parameters into the real-time physics engine (such as NVIDIA PhysX), and perform physical simulation verification on the preliminary rendering frame sequence frame by frame, mark the frames with a physical parameter deviation greater than the threshold (such as a frame with a sudden rise in flame temperature to 1500℃, exceeding the allowed range of 1200℃) as "physical violation frames", record the violation type (temperature exceeds the standard), frame number (such as frame 180), violation parameter value (1500℃), and correction suggestion (reduce the temperature to 1100℃).

[0299] Real-time semantic discriminator verification: Use the "CLIP semantic matching + style discriminator" double model for semantic verification, for example, input the rendered frame corresponding to the optimized parameters (such as the 150 frames after adjusting the color of the flame) into the CLIP model, calculate the similarity with the "fantasy dark flame" semantic vector (such as similarity 0.85 ≥ threshold 0.7, consistent with semantics); input the style discriminator (pre-trained fantasy / realistic classification model), judge the element style matching degree (such as the style matching degree of the dark gold reflection of the dragon's scales 0.92, consistent with the requirements), mark the frames that do not match the semantics (such as a frame of snow mountain appearing RGB(220,220,220) bright color area) as "semantic violation frames", record the violation type (color tone exceeds the standard), frame number (such as 250 frames), violation parameter (RGB value) and correction suggestion (adjust the color tone to RGB(180,180,200)).

[0300] Generation parameter adjustment and re-rendering verification:

[0301] Violation parameter batch adjustment: Correct the optimized generation parameters, for example, for physical violation parameters, such as reducing the flame temperature from 1500°C to 1100°C at frame 180, synchronously adjusting the melting rate in the thermodynamic constraint (from 0.1 m / s to 0.08 m / s); for semantic violation parameters, such as adjusting the RGB value of the snow mountain area from (220,220,220) to (180,180,200) at frame 250, synchronously correcting the light reflection coefficient of the area (from 0.8 to 0.6, to avoid excessive brightness); generate "final corrected parameter set" after adjustment, mark the corresponding violation report ID for each adjustment to ensure traceability.

[0302] Local re-rendering: Only locally render the frame sequence with violations (instead of full sequence re-rendering, saving resources), for example, for frames 180-185 (flame temperature violation interval) and 250-255 (snow color tone violation interval), load the corrected parameters and re-execute the layered rendering (first render the background layer, then superimpose the dynamic layer and the detail layer) to generate the corrected PNG frames; monitor whether the parameters still violate the rules during rendering (such as the corrected temperature 1100°C at frame 180, which needs to be checked in real time whether it meets the constraints), if there are still violations, adjust the parameters again (such as reducing the temperature to 1050°C), until there are no violations in the locally rendered frames.

[0303] Final verification and modeling information output: integrate the local correction frame with the original qualified frame into a complete frame sequence, and perform joint constraint verification again (randomly select 20 frames, including 10 original qualified frames and 10 correction frames), check the physical parameter deviation ≤ 20%, the semantic matching degree ≥ 0.7, and no new violations; after verification, integrate the corrected generation parameters with the complete frame sequence to form “verified modeling information”, label the verification time, the number of violation correction times, the final parameter version (such as V2.0, corresponding to two parameter adjustments), submit to the spatiotemporal consistency optimization and content fusion module, and generate “verification summary report” at the same time, including violation type statistics (such as 2 physical violations and 1 semantic violation), correction effect (such as temperature deviation from 66.25% to 8.3%), for reference by subsequent modules.

[0304] In this embodiment, the spatiotemporal consistency optimization and content fusion module comprises:

[0305] A spatiotemporal consistency detection and marking module is configured to perform spatiotemporal consistency detection on the verified modeling information to obtain a spatiotemporal inconsistency mark, wherein the spatiotemporal inconsistency mark comprises a frame number and a region coordinate that need to be optimized.

[0306] An optical flow prediction and motion trajectory smoothing module is configured to process according to the spatiotemporal inconsistency mark to form a smoothed motion vector field, wherein the smoothed motion vector field comprises a corrected inter-frame motion parameter.

[0307] A temporal discriminator verification and continuity optimization module is configured to generate a video frame sequence that passes the temporal continuity verification according to the smoothed motion vector field.

[0308] Multi-scale content fusion and final rendering are configured to generate a spatiotemporal consistency optimized video frame sequence according to the video frame sequence that passes the temporal continuity verification, the macro scene capsule, the mesoscopic action and special effect capsule, and the microscopic detail and light and shadow capsule.

[0309] In this embodiment, the spatiotemporal consistency optimization and content fusion module can perform the following operations:

[0310] The verified modeling information received by the physical-semantic joint constraint verification module is archived in three categories: "frame sequence data", "parameter data", and "verification records". The coordinates of dynamic elements (e.g. dragon, flame, snow splashes) in the frame sequence, motion parameters (e.g. wing flapping angle change), and macro / meso / micro capsules' original space-time constraints (e.g. macro core zone coordinates, meso action time interval) are extracted. Key dynamic data is ensured to be complete (e.g. if there is no flame particle inter-frame motion trajectory, it is returned for supplementation). The complete frame sequence is unified in format and coordinates.

[0311] Space-time consistency detection and abnormal area marking:

[0312] Space consistency detection: The "dynamic element coordinate trajectory tracking + collision conflict judgment" method is used to detect whether the dynamic elements meet the macro space constraints and inter-frame space continuity frame by frame. For example, for the dragon model, the center coordinates of each frame are extracted (e.g. 91 frames (400, 450, 260), 92 frames (402, 450, 261)), the inter-frame coordinate offset is calculated (e.g. x-axis offset 2m, z-axis offset 1m), and if the offset of a certain frame suddenly increases (e.g. 95 frames of dragon coordinates from (410, 450, 263) to (450, 450, 263), offset 40m, far exceeding the normal flapping-induced 5m / frame offset), it is marked as a "space abnormal frame".

[0313] Time consistency detection: The "inter-frame motion parameter continuity analysis + time sequence constraint matching" is used to detect the time continuity of dynamic elements. For example, for the dragon wing flapping action, the change amount of the flapping angle of each frame is calculated (e.g. 91 frames -35°, 92 frames -32°, 93 frames -29°, normal change amount 3° / frame), and if the angle of 98 frames suddenly changes from -20° to +10° (change amount 30° / frame, far exceeding the normal range), it is marked as a "time motion mutation abnormality".

[0314] All detected abnormalities are arranged in the structure of "abnormal type (space / time) -> frame number interval -> abnormal area coordinates -> preliminary judgment of abnormal reason", and a "space-time abnormality report" is generated.

[0315] Optical flow prediction and motion trajectory smoothing: For time motion mutation abnormalities (e.g. 98 frames of wing angle mutation), an optical flow prediction model (e.g. FlowNet2.0) is used to calculate the motion vector field of the frames before and after the abnormal frame (97 frames, 99 frames). For example, the 97 frames of wing flapping angle are -23°, and the 99 frames are -17°. The reasonable angle of 98 frames is predicted by the optical flow model to be -20° (inter-frame change amount 3°), and the corrected motion parameters (angle -20°, flapping speed 8m / s) are generated. The smoothed motion vector field contains the corrected motion parameters (angle, speed, coordinates) of each dynamic element.

[0316] Temporal discriminator verification and continuity optimization: input the modified frame sequence into a pre-trained temporal discriminator (based on LSTM+CNN, used to judge the dynamic continuity between frames), the discriminator scores the motion parameters and temporal logic of each frame (0-1 score, 1 for complete continuity), for example, if the modified 95th, 98th, and 182nd frames all score ≥0.9 (the original abnormal frame scores ≤0.3), it means that the temporal continuity meets the standard; for frames that still do not meet the standard (such as the 211th flame frame that exceeds the period, with a score of 0.6), directly delete the flame particle layer of the frame to ensure that the temporal sequence and the mesoscopic constraints (151-210 frames) are completely matched; at the same time, add a gradual transition effect to the start / end transition frames of dynamic elements (such as the flame start frame 151 and the end frame 210) to avoid the temporal discomfort of "sudden appearance / disappearance", and finally generate a "temporally coherent video frame sequence".

[0317] Multi-scale content hierarchical fusion: in the order of "macroscopic scene layer → mesoscopic dynamic layer → microscopic detail layer → light and shadow effect layer", the verified frame sequence is deeply fused, for example, the macroscopic scene layer loads the basic model of the main peak of the snow mountain and the sky moon, and renders the background texture according to the macroscopic style constraint (dark color RGB(30, 40, 80)); the mesoscopic dynamic layer superimposes the modified dragon, flame, and snow splashing model to ensure that the dynamic element coordinates match the macroscopic space (such as the dragon always moving in the core area); the microscopic detail layer precisely attaches scale texture, ice crystal texture, and flame particle detail map to the corresponding model (such as the scale texture stretching with the wing flapping angle); the light and shadow effect layer integrates macroscopic moonlight (200 lux) and microscopic flame lighting (1100°C corresponding to 200 lux local light), calculates the global lighting distribution of each frame (such as 400 lux in the flame area and 180 lux in the background area), and adds dynamic shadows (such as the dragon shadow shifting with the position change between frames, with a shadow length ratio of 1:2), ensuring that the contents of each scale are completely fused in space and style without hierarchical fragmentation.

[0318] For the fused frame sequence, detail enhancement and style uniform adjustment are performed, the fused and optimized frame sequence is rendered into a frame sequence of MP4 format with a resolution of 1920x1080 (frame rate 30fps), frame-by-frame checking of spatiotemporal consistency (such as smoothness of dynamic element coordinate track, matching of timing and mesoscopic constraints), style consistency (tonality, light and shadow conform to macroscopic constraints), and detail integrity (microscopic texture is clear, special effect is reasonable); 30 frames (covering the opening, battle, and ending stages) are randomly extracted, 3 test personnel are invited to score the subjective coherence (1-5 points, 5 points being the best), and the average score is ≥4.5 points to determine that it is qualified; after being qualified, a "spatiotemporal consistency optimized video frame sequence" is generated, which is labeled with a unique identifier (such as "optimized frame sequence-dragon snow mountain-001"), optimization record (such as 3 places of spatial anomaly and 2 places of time anomaly), and submitted to a dynamic code rate video synthesis and output module, and an "optimization summary report" is generated, including a frame sequence comparison graph before and after optimization, and key parameter adjustment record (such as wing fan angle correction value), for subsequent module tracing.

[0319] In the embodiment, the dynamic code rate video synthesis and output module comprises:

[0320] a content importance analysis and region division module, which is configured to identify the spatiotemporal consistency optimized video frame sequence to obtain a content importance atlas, the content importance atlas comprising importance scores of each region of each frame;

[0321] a dynamic code rate allocation strategy formulation module, which is configured to perform dynamic code rate allocation according to the content importance atlas to generate a dynamic code rate allocation table, the dynamic code rate allocation table comprising code rate values of each region of each frame;

[0322] video encoding and dynamic code rate compression, configured to encode and compress the spatiotemporal consistency optimized video frame sequence according to the dynamic code rate allocation table to obtain a dynamic code rate compressed video stream.

[0323] In the embodiment, the dynamic code rate video synthesis and output module can perform the following processing:

[0324] receiving the spatiotemporal consistency optimized video frame sequence, extracting the basic attributes (resolution, frame rate, color space being RGB) of the frame sequence, and simultaneously obtaining spatiotemporal distribution information of key dynamic elements (such as the dragon battle core area being concentrated in frame pixel coordinates (800-1100, 400-700) and the flame special effect area being concentrated in (850-950, 450-550)) from the optimization report, to clearly determine the key areas and non-key areas for subsequent code rate allocation.

[0325] The PNG format frame sequence is converted to YUV420 color space in batches, noise removal processing is performed on the converted frame sequence, and the possible residual tiny pixel noise in the optimized frame sequence is eliminated to avoid waste of coding rate caused by noise. Meanwhile, the frame sequence is divided into 15 frames (0.5 seconds) as one "coding unit" (a total of 30 coding units) to facilitate subsequent analysis of content importance and allocation of code rate by unit.

[0326] The frame picture in each coding unit is calculated for regional importance score (0-10 points, 10 points for the highest importance) from three dimensions of "visual saliency, motion intensity, semantic relevance". For example, visual saliency uses a saliency map detection algorithm (such as Itti-Koch model) to calculate the visual attraction degree of each region in the frame. The core area of the dragon fight (800-1100, 400-700) has a high color contrast (red dragon on a dark background) and clear outline, with a score of 9-10 points; the flame special effect area (850-950, 450-550) has a dynamic light emitting characteristic, with a score of 8-9 points.

[0327] Motion intensity: the pixel change rate of each region is calculated by an inter-frame difference algorithm. The flame special effect area has a pixel change rate of 60%-80% (particle dynamic motion) between frames, with a score of 8-9 points; the dragon wing flapping area (880-920, 480-520) has a change rate of 40%-60%, with a score of 7-8 points.

[0328] Semantic relevance: combined with the intention map with priority, the core intention "dragon fight" associated area (dragon body, fight interaction area) has a score of 10 points, and the background area associated with the auxiliary intention "snow mountain scene" has a score of 3-4 points.

[0329] The final importance score of each region is "visual saliency x 40% + motion intensity x 30% + semantic relevance x 30%", with 1 decimal place.

[0330] Each frame picture is divided into grid regions according to the block size of "16x16 pixels" (1920x1080 resolution, a total of 120x67.5≈8100 block regions), and each block region is classified based on the above importance score—regions with a score of ≥8 points are marked as "core important regions", regions with a score of 5≤score<8 points are marked as "medium important regions", and regions with a score of <5 points are marked as "low important regions"; a "content importance map" is generated for each frame, marking the coordinates (such as (800-816, 400-416)) of each block region, the importance score (such as 9.2 points), and the region type (core / medium / low importance), for example, "frame 150 (flame injection peak frame): block region (850-866, 450-466) has a score of 9.1 points (core important region), and block region (200-216, 100-116) has a score of 2.3 points (low important region)".

[0331] Compare the importance map of adjacent frames, if the importance score of the same space region fluctuates more than 2 points (such as the score of a certain block region in frame 150 is 9.1 points, and frame 151 suddenly drops to 6.8 points), smooth the score by weighted average (previous frame weight 60%, current frame weight 40%) (adjusted score = 9.1 x 0.6 + 6.8 x 0.4 = 8.3 points), avoid the frequent switching of region type between adjacent frames leading to too large code rate fluctuation; at the same time, mark the "dynamic important region" (such as the trajectory of the flame special effect region moving with the injection action), keep the code rate continuity of this region in subsequent code rate allocation.

[0332] Total code rate budget and interval division: calculate the total code rate budget according to user output requirements (such as total video duration 15 seconds, target file size ≤100MB), for example, 100MB = 800Mbps, total code rate budget of 15 seconds video is 800Mbps / 15 ≈ 53.3Mbps, considering 10% redundancy reserved for encoding loss, the actual available code rate budget is 48Mbps (average code rate per frame ≈ 48Mbps / 30 ≈ 1.6Mbps under 30fps); set the code rate interval based on the importance type of the region: core important region code rate 7-8Mbps (per block region per frame), medium important region 4-5Mbps, low important region 1-2Mbps, ensure that the code rate of high important region is sufficient and the total code rate does not exceed the budget.

[0333] Frame-level dynamic code rate allocation table generation: combine the content importance map of each frame with the code rate interval to allocate specific code rate values to each region block of each frame, for example, frame 150 (flame injection peak frame): core important region (such as flame, dragon head) each block is allocated 7.8Mbps, medium important region (dragon body non-action part) each block is allocated 4.5Mbps, low important region (distant snow mountain) each block is allocated 1.2Mbps; calculate the total code rate of each frame (sum of code rates of each region block), if the total code rate of a certain frame exceeds the budget (such as frame 150 original total code rate 2.1Mbps > 1.6Mbps), down-regulate the code rate in the order of "low important region → medium important region", until the total code rate of the frame ≤1.6Mbps; arrange by frame to generate "dynamic code rate allocation table", including frame number, region block coordinates, allocated code rate, region type, adjustment record (such as "frame 150: block (200-216, 100-116) code rate from 1.2Mbps to 1.0Mbps").

[0334] Different coding parameters are configured for different importance regions, for example, the core important region adopts "high code rate + low compression ratio" parameters (such as the quantization parameter QP of H.265 coding is 20, and the omnidirectional prediction mode is adopted for intra prediction), and the details and textures (such as the scales of a dragon and the particles of a flame) are preserved; the medium important region adopts "medium code rate + medium compression ratio" parameters (QP is 26, and the bidirectional prediction mode is adopted for prediction); the low important region adopts "low code rate + high compression ratio" parameters (QP is 32, and the inter-frame reference compression is enabled to reduce redundant data); and meanwhile, the "B frame optimization" (one B frame is inserted every 5 frames) is enabled for the core region (such as the flame) that changes dynamically, so as to improve the coding efficiency and avoid the loss of details.

[0335] Dynamic code rate coding compression execution: load the dynamic code rate allocation table and the preprocessed YUV frame sequence, and call the H.265 / HEVC encoder (such as x265) to perform frame-by-frame coding, for example, the allocated code rate of each region of each frame is read in real time during the coding process, the core important region (such as the flame region of frame 150) is configured with the encoder parameters at a code rate of 7.8 Mbps, so as to ensure that the code rate deviation of the region after coding is ≤5% (such as the target 7.8 Mbps, and the actual 7.5-8.1 Mbps); the low important region (such as the static snow mountain) is compressed at a code rate of 1.0 Mbps, and the allowed deviation is ≤10%; the CPU / GPU resource occupation of the encoder is monitored during the coding process (such as CPU occupation ≤80%, and GPU computing power ≤70%), if the coding time of a certain frame exceeds 33 ms (30 fps single frame time limit), the code rate of the low important region of the frame is temporarily reduced (such as from 1.0 Mbps to 0.8 Mbps), so as to ensure that the coding is smooth and not stuck; and the dynamic code rate compressed video stream (the initial format is TS stream) is generated after the coding is completed.

[0336] In the embodiment, the hierarchical intention analysis and dynamic graph construction module further comprises:

[0337] The intention conflict detection and self-repair module detects and repairs the conflicts of the intention graph according to the intention graph after the dynamic adjustment, so as to obtain the intention graph after the repair; and the intention priority sorting and constraint generation module is used for processing the intention graph after the repair.

[0338] In the embodiment, the intention conflict detection and self-repair module comprises:

[0339] The intention conflict detection module detects the intention conflicts by using the ICOF formula, so as to obtain the conflict labels;

[0340] The intention conflict repair module repairs the intention conflicts according to the conflict labels, so as to obtain the intention graph after the repair.

[0341] In this embodiment, the ICOF formula is as follows:

[0342]

[0343] where I i ,I j are two intent nodes in the dynamic graph (e.g., "Dragon Fighting" and "Clear Sky"). v(I) is the semantic vector of intent I, generated by a pre-trained language model (e.g., CLIP-Text), containing keywords, extended words, and weights. α, β are weight coefficients of semantic and physical contradiction items, satisfying α + β = 1 (determined by dynamic programming optimization). C k (I) is the demand value of intent I for physical constraint k (e.g., the "Dragon Mass" requirement is 500 kg). R k is the actual upper limit of the physical constraint k (e.g., the "maximum load of the physical engine" is 800 kg).

[0344] In this embodiment, all intent pairs (I i ,I j ) in the graph are traversed, and the intent contradiction correlation optimization function is calculated. If ICOF(I i ,I j ) > θ (threshold θ is determined by user preference learning, e.g., θ = 0.5), mark the intent pair I i ,I j as a contradiction pair.

[0345] The present application replaces the traditional expert rule base with the self-designed ICOF formula, which can adapt to any combination of intents (e.g., "Dragon Fire" + "Clear Sky"). It considers both semantic and physical contradictions, avoiding new contradictions caused by local repair (e.g., only repairing semantic contradictions may violate physical laws). Through vector calculation and GAN generation, the repair speed is 3-5 times faster than traditional rule matching (experimental data).

[0346] The present application has the following advantages:

[0347] Traditional methods rely on static rule base to parse intents and cannot handle emerging combinations (e.g., implicit contradictions of "Dragon Fire" + "Clear Sky"). This method dynamically integrates multi-modal features (text, vision, audio) through a residual adjustment mechanism, automatically adjusts the intent dependency relationship based on dynamic priority sorting, and constructs an extensible intent graph. For example, when the user inputs "fantasy style dragon fighting", the system can automatically identify the contradiction between "fire needs clouds" and "sky without clouds", and prioritize solving core conflicts (e.g., adjust "clear sky" to "cloudy"), avoiding the limitations of hard-coded rules.

[0348] The traditional method adopts serial verification (physical first and semantic second), which is easy to cause new problems (such as repairing semantic contradiction to violate physical law) due to local repair. The method innovatively proposes an intention contradiction correlation optimization function (ICOF), combines semantic contradiction items (detected by semantic vector cosine similarity) and physical contradiction items (detected by the difference between physical constraint demand and resource upper limit), and dynamically detects and repairs intention contradiction. For example, when the "dragon mass + fire energy" exceeds the physical engine load, the ICOF formula can adjust "dragon mass" and "fire energy" at the same time to ensure the balance between physical feasibility and semantic expressiveness.

[0349] The traditional method generates scenes, actions and details in steps and then directly splices them, which is easy to cause spatiotemporal misplacement (such as sudden change of dragon position in adjacent frames). The method marks abnormal areas through spatiotemporal inconsistency detection (optical flow prediction + time sequence discriminator), combines Bézier curve fitting to smooth motion trajectory, and dynamically aligns macroscopic scenes, mesoscopic actions and microscopic details through a multi-scale content fusioner (such as spatial consistency of dragon wing flapping and fire effect). For example, when generating "dragon wing flapping", the system can automatically adjust the wing flapping frequency to ensure coherent action and compliance with aerodynamics.

[0350] The traditional method adopts fixed code rate compression, and the core area (such as dragon fire) is easy to lose details due to low code rate, and the background area (such as static mountains) is wasted in volume. The method generates a region importance map through content importance analysis (visual saliency + motion intensity), combines a dynamic code rate allocation strategy (high code rate for core area and low code rate for background area), and supports multi-format / multi-platform adaptation (such as TikTok vertical screen and YouTube horizontal screen). For example, when generating "dragon fire", the system can automatically allocate 8Mbps high code rate for fire effect and 2Mbps low code rate for background mountains to balance picture quality and file size.

[0351] The application also provides an AIAgentic multi-modal collaborative control automatic video generation method, which comprises:

[0352] Obtaining multi-modal generation requirement information provided by a user;

[0353] Analyzing the multi-modal generation requirement information to generate an intention graph with priority and constraint conditions;

[0354] Generating multi-scale spatiotemporal capsule modeling information according to the dynamically evolved intention graph;

[0355] Verifying the multi-scale spatiotemporal capsule modeling information to obtain verified modeling information;

[0356] The verified modeling information is optimized in space-time consistency and fused with content, so as to obtain a video frame sequence optimized in space-time consistency;

[0357] A video file is generated according to the video frame sequence optimized in space-time consistency.

[0358] Although the present application has been described in detail with general description and specific embodiments above, some modifications or improvements can be made on the basis of the present application, which is obvious to those skilled in the art. Therefore, these modifications or improvements made on the basis of not deviating from the spirit of the present application, all belong to the scope of the present application claimed.

Claims

1. An automated video generation system with AI Agentic multimodal collaborative control, characterized in that, The AIAgentic multimodal collaborative control automated video generation system includes: A requirement acquisition module, which is used to acquire multimodal generation requirement information provided by the user; A hierarchical intent parsing and dynamic graph construction module is used to parse the multimodal generation requirement information, thereby generating an intent graph with priority and constraints. A multi-scale spatiotemporal capsule scene generation module is used to generate multi-scale spatiotemporal capsule modeling information based on the dynamically evolving intent graph. The physical-semantic joint constraint verification module is used to verify the multi-scale spatiotemporal capsule modeling information, thereby obtaining the verified modeling information. The spatiotemporal consistency optimization and content fusion module is used to perform spatiotemporal consistency optimization and content fusion on the verified modeling information, thereby obtaining a spatiotemporally consistent optimized video frame sequence; A dynamic bitrate video synthesis and output module is used to generate video files based on a video frame sequence optimized for spatiotemporal consistency.

2. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 1, characterized in that, The requirement acquisition module includes: A text information acquisition module, which is used to acquire text information provided by the user; A visual information acquisition module, wherein the visual information acquisition module is used to acquire visual information provided by the user; An audio information acquisition module, wherein the audio information acquisition module is used to acquire audio information provided by the user; The text processing module is used to extract text features from the text information and expand the text features to obtain text modal features, which include text semantic vectors, expanded keyword lists and corresponding semantic weights. A visual processing module is used to extract features from visual information to obtain visual modal features, which include object-level visual feature vectors and visual dynamic trajectory feature vectors. An audio processing module is used to extract the audio information to obtain audio modal features, which include acoustic feature vectors and emotion tags. The dynamic modal weight allocation module is used to allocate weights to text modal features, visual modal features and audio modal features respectively, so as to obtain a multimodal feature vector after dynamic weight allocation. A cross-modal residual learning and alignment module is used to perform residual adjustment on the multimodal feature vectors after dynamic weight allocation, thereby obtaining residual-adjusted multimodal feature vectors, which serve as the required information.

3. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 2, characterized in that, The hierarchical intent parsing and dynamic graph construction module includes: An intent classification module is used to generate a core intent set based on the multimodal generation requirement information, the core intent set including at least one core intent information; A confidence scoring module is used to score the confidence of each core intent information. The sub-intent recursive decomposition module is used to decompose each core intent information separately to obtain executable sub-intents. A core intent information includes at least one sub-intent. A hierarchical intent graph construction module is used to generate and construct a hierarchical intent graph based on core intent information and sub-intents; wherein, the nodes of the constructed hierarchical intent graph include core intent information and sub-intents, and the edges represent dependencies; A dynamic intent graph evolution and adjustment module is used to generate a dynamically adjusted intent graph based on a hierarchical intent graph. The intent priority sorting and constraint generation module is used to adjust the priority and constraints of each node in the dynamically adjusted intent graph, thereby generating an intent graph with priorities and constraints.

4. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 3, characterized in that, The multi-scale spatiotemporal capsule scene generation module includes: A macro-scale scene layout generation module is used to generate macro-scale scene capsules based on the intent graph with priority and constraints. A mesoscale character motion and special effects generation module, which is used to generate mesoscale motion and special effects capsules based on the macroscale scene capsule and the requirement information; The microscale detail texture and lighting generation module is used to generate microscale detail and lighting capsules based on macroscale scene capsules, mesoscale action and special effects capsules, and requirement information. The spatiotemporal capsule routing and cross-scale connection module is used to merge macroscopic scene capsules, mesoscopic action and special effects capsules, and microscopic detail and light and shadow capsules into a spatiotemporal consistent capsule set. The multi-scale content fusion and preliminary rendering module is used to perform capsule fusion and preliminary rendering on the spatiotemporal consistency capsule set to obtain a preliminary rendered video frame sequence, which serves as multi-scale spatiotemporal capsule modeling information.

5. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 4, characterized in that, The physical-semantic joint constraint verification module includes: A joint constraint field construction module is used to generate a joint constraint field model based on the multi-scale spatiotemporal capsule modeling information and an intent graph with priority and constraint conditions. The simulated annealing optimization module is used to optimize the joint constraint field model using a simulated annealing algorithm, thereby obtaining the optimized generation parameters. A real-time physics engine verification module is used to verify the optimized generation parameters, thereby generating physics violation markers and correction suggestions. A real-time semantic discriminator verification module is used to perform semantic scoring and evaluation on the optimized generated parameters, thereby obtaining semantic violation tags and correction suggestions. A parameter adjustment module is provided, which is used to adjust the optimized generation parameters according to the physical violation markers and correction suggestions and the semantic violation markers and correction suggestions, so as to obtain the adjusted generation parameters. A re-rendering module is used to adjust the multi-scale spatiotemporal capsule modeling information according to the adjusted generation parameters, thereby obtaining verified modeling information.

6. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 5, characterized in that, The spatiotemporal consistency optimization and content fusion module includes: A spatiotemporal consistency detection and marking module is used to perform spatiotemporal consistency detection on the verified modeling information to obtain spatiotemporal inconsistency markers, which include the frame number to be optimized and the region coordinates. An optical flow prediction and motion trajectory smoothing module is used to process the spatiotemporal inconsistency marker to form a smoothed motion vector field, which includes corrected inter-frame motion parameters. A timing discriminator verification and coherence optimization module is used to generate a sequence of video frames that pass timing coherence verification based on a smoothed motion vector field. Multi-scale content fusion and final rendering are used to generate a spatiotemporally consistent video frame sequence based on the video frame sequence that has passed the temporal coherence verification, macro scene capsules, meso action and special effects capsules, and micro detail and light and shadow capsules.

7. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 6, characterized in that, The dynamic bitrate video synthesis and output module includes: The content importance analysis and region segmentation module is used to identify the video frame sequence after spatiotemporal consistency optimization, thereby obtaining a content importance map, which includes the importance score of each region in each frame; A dynamic bitrate allocation strategy formulation module is used to dynamically allocate bitrate based on the content importance graph, thereby generating a dynamic bitrate allocation table, which includes the bitrate values ​​of each region in each frame. Video encoding and dynamic bitrate compression are used to encode and compress the spatiotemporally consistent optimized video frame sequence according to the dynamic bitrate allocation table, thereby obtaining a dynamically bitrate compressed video file.

8. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 3, characterized in that, The hierarchical intent parsing and dynamic graph construction module further includes: The intent conflict detection and self-repair module detects and repairs conflicts based on the dynamically adjusted intent graph, thereby obtaining a repaired intent graph; the intent priority sorting and constraint generation module processes the repaired intent graph.

9. The automated video generation system with AI Agentic multimodal collaborative control as described in claim 8, characterized in that, The intent contradiction detection and self-repair module includes: An intent conflict detection module is used to detect intent conflicts using the ICOF formula, thereby obtaining conflict markers. An intent conflict repair module is used to repair intent conflicts based on the conflict markers, thereby obtaining a repaired intent graph.

10. An automated video generation method with AIAgentic multimodal cooperative control, characterized in that, The automated video generation method of AIAgentic multimodal collaborative control includes: Obtain multimodal generation requirements information provided by the user; The multimodal generation requirement information is parsed to generate an intent graph with priorities and constraints. Multi-scale spatiotemporal capsule modeling information is generated based on the dynamically evolving intent map; The modeling information of the multi-scale spatiotemporal capsule is validated to obtain validated modeling information; The validated modeling information is subjected to spatiotemporal consistency optimization and content fusion to obtain a spatiotemporally consistent optimized video frame sequence; Video files are generated based on the video frame sequence optimized for spatiotemporal consistency.

Citation Information

Patent Citations

  • Multi-source material fused video content generation method, system, equipment and medium

    CN119996786A

  • High-quality video content automatic generation method and related equipment

    CN120050487A

  • Multi-modal fusion real-time digital human driving method based on unified behavior vector mapping

    CN120339477A

  • Video plot generation and scene synthesis method and system based on natural language processing

    CN120339919A

  • Video generation system based on video large model

    CN120343361A

Cited By

  • Mixed content generation scheduling method and system for resource-constrained equipment

    CN121564171A