A text-driven video generation method based on causal kinematic chain reasoning
By constructing a structured motion graph and introducing a causal motion chain reasoning method that incorporates static, rigid, and non-rigid body guiding signals, this method addresses the shortcomings of existing text-driven video generation methods in motion detail modeling, physical consistency, and cross-modal semantic logic reasoning. It achieves high-quality and reliable video generation, adapting to diverse application needs.
Patent Information
- Application Number
- CN202610724120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-25
AI Technical Summary
Existing text-driven video generation methods have shortcomings in motion detail modeling, physical consistency, and cross-modal semantic logic reasoning, resulting in insufficient temporal consistency, motion accuracy, physical realism, and semantic consistency of the generated videos. Furthermore, they are highly dependent on training and cannot flexibly adapt to changing application requirements.
By constructing a structured motion graph, analyzing the instances and interaction relationships in the text prompts, generating a frame-level layout sequence, and introducing static, rigid, and non-rigid guiding signals into the diffusion sampling process, a large language model is used to perform causal motion chain reasoning, thereby achieving fine control over video generation.
It improves the consistency between inter-frame motion transitions, object interaction logic, and text semantics in video generation, reduces model deployment and maintenance costs, enhances the temporal consistency, physical rationality, and semantic accuracy of video generation, adapts to different generation architectures, and improves the quality and reliability of generated videos.
Smart Images

Figure CN122640604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence video generation technology, and in particular to a text-driven video generation method based on causal motion chain reasoning. Background Technology
[0002] With the development of generative artificial intelligence, text-driven video generation technology has made significant progress. Users only need to provide a text description, and the model can automatically generate dynamic video content that semantically corresponds to it, showing broad application prospects in digital content creation and film and animation production. Current mainstream text-to-video methods typically employ deep generative architectures such as diffusion models, learning the text-to-video mapping relationship from large-scale text-to-video data. However, in practical applications, these methods still face the following key challenges in motion modeling, physical consistency, and training dependency: Insufficient fine-grained motion modeling: Existing models lack explicit representation and constraints on fine-grained motion elements, and usually rely only on implicit guidance from text. It is difficult to accurately control the motion trajectory, speed and temporal evolution of objects, which easily leads to the phenomenon of discontinuous action between frames and blurred motion trajectory, resulting in insufficient temporal consistency and motion accuracy of the generated video.
[0003] Weak physical consistency constraints: Existing methods do not adequately depict objective physical relationships and lack systematic modeling of physical constraints such as rigid / non-rigid body properties, gravity, collision and occlusion. The generated results may exhibit phenomena such as objects intersecting each other, floating, severe rigid body deformation, or inconsistent size ratios between frames, affecting the realism and credibility of the video.
[0004] High dependence on additional training / fine-tuning: Most text-to-video methods require additional training or fine-tuning on specific datasets to learn new motion patterns or improve performance in specific scenes. This dependence on training reduces the generality of the method and significantly increases the cost of expansion. When the base model version is updated or migrated to a new scene, retraining is often required, making it difficult to flexibly meet the changing application requirements.
[0005] Insufficient cross-modal collaborative reasoning ability: There are complex temporal causal and multi-entity interaction correspondences between text descriptions and video content, but existing models mostly adopt simple text embedding conditionalization methods, lacking in-depth understanding of temporal causal chains and multi-entity interaction logic and explicit reasoning mechanisms, which makes it difficult for the generated video to accurately reflect the causal plot and multi-character action sequence in the text, affecting semantic consistency and narrative integrity.
[0006] In summary, existing text-driven video generation methods still have significant room for improvement in terms of motion detail modeling, adherence to physical laws, training efficiency, and cross-modal semantic logic reasoning.
[0007] Chinese patent application document CN121725084A, published on March 24, 2026, discloses a method and apparatus for generating a homepage video background based on page elements. The method includes: acquiring multimodal interface features of the homepage of a target website; receiving natural language instructions input by the user through an interactive dialog box; parsing the visual intent, style preferences, and dynamic requirements in the natural language instructions to generate a structured visual control description; extracting layout, color, and semantic entities from the multimodal interface features through a cross-modal attention model to form a generation condition vector; inputting the generation condition vector into a conditional decoupling video diffusion model; decoupling the generation condition vector into a control subspace of content, style, and motion; and synthesizing a video background sequence based on the control subspace and the structured visual control description to composite the dynamic background in the homepage of the target website.
[0008] The patent application discloses a method and apparatus for generating a homepage video background based on page elements. This method generates a dynamic video background by using the webpage interface itself as a condition and user language as a guide, avoiding the crude control issues of general text-based video models. However, the generated video is subpar in terms of inter-frame motion transitions, object interaction logic, and consistency with text semantics. Summary of the Invention
[0009] To overcome the shortcomings of the prior art, this invention provides a text-driven video generation method based on causal motion chain reasoning. This invention can improve the quality and reliability of text-to-video generation, and significantly enhance the generated video in terms of inter-frame motion transitions, object interaction logic, and semantic consistency with the text.
[0010] This invention is achieved through the following technical solution: A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1, receiving user text prompts, parsing the instances, action attributes and interaction relationships in the text prompts, and constructing a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0011] In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
[0012] In step S2, the frame-level layout sequence is represented as follows: ; in, For frame-level layout sequences, For structured motion graphs and text prompts The layout reasoning output; layout evolution constraints include: Static instances satisfy Rigid body instances satisfy Non-rigid instances satisfy ; in, For frame number, For the first Layout of frame instance nodes The layout of the instance nodes in frame 1. For the first Layout of frame instance nodes For velocity vectors, For acceleration vectors, This is a summary of layout-level displacement or deformation.
[0013] In step S4, the generation of the static guide includes: selecting a reference frame index in the latent space; ; in, For the first The feature mapping results of the frame For the first The feature mapping results of the frame Reference frame index; For feature distance measurement; Construct static guidance; ;in: For static guiding masks; This is an indicator function.
[0014] In step S4, the generation of rigid body guidance includes: constructing a shape template for the rigid body instance and performing geometric alignment, and defining a frame-independent template; ; in, For frame-independent shape templates, For the first Foreground mask after frame geometry alignment For aggregation operators; ; in, For deformation mapping operators, For thresholding or voting operators; Introduce displacement penalty; ; in, As a displacement penalty, Let f be the center coordinates of the instance in frame f. For hyperparameters; For the first Frame instance center coordinates; Define rigid body guidance; ; in, For rigid body guidance, for The transpose of the matrix, It is the element-wise product of the displacement penalty Γ.
[0015] In step S4, the generation of non-rigid body guidance includes: establishing pixel-level correspondence in the feature space; ; in, For pixels At the corresponding position in the matched frame, Position of frame f The feature representation of the location, For the first Frame position Feature representation of the location; Define perceived deformation; ; in, To sense deformation, These are the pixel position coordinates; Obtained by interpolation of the corner points of the layout box: ; in, To induce deformation of the frame, It is a bilinear interpolation operator; Structural deformation penalty; ; in, As a penalty for deformation; Define non-rigid body guidance; ; in, Non-rigid body guidance, In order to deal with deformation penalty Element-wise product.
[0016] In step S4, the injection diffusion sampling process adopts a gradient update method: ; ; in, For the updated sequence of latent variables, For the current step, the sequence of latent variables. The gradient of loss L with respect to latent variables, To guide the loss function, For guiding coefficients, For normalization term, For attention maps, For static guiding mask, It is a rigid body guiding mask.
[0017] In step S4, the injection diffusion sampling process achieves guided injection by modulating attention scores, which is suitable for video generation backbones of DiT architecture.
[0018] The motion categories include static instances, rigid body motion instances, and non-rigid body motion instances, which correspond to static guidance, rigid body guidance, and non-rigid body guidance, respectively, to achieve decoupled control of multiple instances and multiple motion categories.
[0019] The frame-level layout sequence is obtained by reasoning about causal motion chains based on structured motion graphs and text prompts using a large language model. The frame-level layout sequence is used to explicitly represent the spatial position and motion trajectory of each instance in time.
[0020] In step S1, the instance node includes an instance identifier, an attribute set, and a motion category label.
[0021] The beneficial effects of this invention are mainly reflected in the following aspects: 1. Compared with the prior art, the present invention can improve the quality and reliability of text-to-video generation, and significantly improve the generated video in terms of inter-frame motion transitions, object interaction logic, and semantic consistency with the text.
[0022] 2. This invention requires no additional training and is plug-and-play: It adopts a "planning-guided" approach, controlling only the inference stage, without the need to retrain or fine-tune the underlying generative model, which greatly reduces the cost of model deployment and maintenance.
[0023] 3. In this invention, the overall module is decoupled from the specific generation architecture, which can be easily integrated into video generation models with different architectures and has good engineering portability.
[0024] 4. This invention explicitly plans the event sequence by extracting motion facts and encoding causal chains, effectively alleviating the problems of inter-frame discontinuity, jitter, and target drift, making the motion trajectory more stable and natural, and significantly improving the consistency of the time sequence.
[0025] 5. This invention introduces physical consistency constraints and applies differentiated rules to static, rigid body, and non-rigid body motion, reducing physical violations such as object interpenetration, floating, abnormal deformation, and inconsistent scale, thereby enhancing physical rationality and realism.
[0026] 6. This invention, through structured parsing and cross-modal reasoning, more accurately models multi-entity relationships and causal events, making the generated content more consistent with the text narrative logic, avoiding semantic mismatch, and improving semantic matching accuracy.
[0027] 7. This invention designs differentiated guidance paths for three types of motion: static, rigid body, and non-rigid body, solving the problem of easy confusion between different motion categories in existing methods and achieving refined independent control.
[0028] 8. Under the static guidance path, the present invention can effectively constrain the feature consistency of the background or static objects in the video sequence, ensuring that they remain stable in dynamic scenes and do not change unexpectedly with the movement of the foreground, thereby improving the cross-frame stability of static instances.
[0029] 9. This invention, for rigid body motion examples, uses shape templates and displacement penalty mechanisms to ensure that the object maintains its overall shape during movement, avoiding unexpected stretching or twisting.
[0030] 10. This invention, for instances of non-rigid body motion, establishes pixel-level correspondences and constructs deformation penalties, allowing objects to undergo natural local deformations, thus balancing flexibility and physical constraints.
[0031] 11. This invention, through a decompositional control paradigm of "planning before generation", improves the temporal continuity, physical rationality and semantic accuracy of video generation without the need for training, and greatly enhances the ability to handle complex multi-entity interaction scenarios. Attached Figure Description
[0032] The present invention will now be further described in detail with reference to the accompanying drawings and specific embodiments, wherein: Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0033] Example 1 See Figure 1 A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0034] This embodiment is the most basic implementation method. Compared with the prior art, it can improve the quality and reliability of text-to-video generation, and significantly improve the generated video in terms of inter-frame motion transitions, object interaction logic, and consistency with text semantics.
[0035] Example 2 See Figure 1 A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0036] In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
[0037] This embodiment is a preferred implementation method that requires no additional training and is plug-and-play: it adopts a "planning-guided" approach, which only controls the inference stage and does not require retraining or fine-tuning of the underlying generative model, thus greatly reducing the cost of model deployment and maintenance.
[0038] The overall module is decoupled from the specific generation architecture, which can be easily integrated into video generation models with different architectures, and has good engineering portability.
[0039] Example 3 See Figure 1 A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0040] In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
[0041] In step S2, the frame-level layout sequence is represented as follows: ; in, For frame-level layout sequences, For structured motion graphs and text prompts The layout reasoning output; layout evolution constraints include: Static instances satisfy Rigid body instances satisfy Non-rigid instances satisfy ; in, For frame number, For the first Layout of frame instance nodes The layout of the instance nodes in frame 1. For the first Layout of frame instance nodes For velocity vectors, For acceleration vectors, This is a summary of layout-level displacement or deformation.
[0042] This embodiment is a preferred implementation method. By extracting motion facts and encoding causal chains, the event sequence is explicitly planned, which effectively alleviates the problems of inter-frame discontinuity, jitter and target drift, making the motion trajectory more stable and natural, and significantly improving the consistency of the time sequence.
[0043] By introducing physical consistency constraints, differentiated rules are applied to static, rigid body, and non-rigid body motions, reducing physical violations such as object interpenetration, floating, abnormal deformation, and inconsistent scale, thereby enhancing physical rationality and realism.
[0044] Example 4 See Figure 1 A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0045] In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
[0046] In step S2, the frame-level layout sequence is represented as follows: ; in, For frame-level layout sequences, For structured motion graphs and text prompts The layout reasoning output; layout evolution constraints include: Static instances satisfy Rigid body instances satisfy Non-rigid instances satisfy ; in, For frame number, For the first Layout of frame instance nodes The layout of the instance nodes in frame 1. For the first Layout of frame instance nodes For velocity vectors, For acceleration vectors, This is a summary of layout-level displacement or deformation.
[0047] In step S4, the generation of the static guide includes: selecting a reference frame index in the latent space; ; in, For the first The feature mapping results of the frame For the first The feature mapping results of the frame Reference frame index; For feature distance measurement; Construct static guidance; ;in: For static guiding masks; This is an indicator function.
[0048] In step S4, the generation of rigid body guidance includes: constructing a shape template for the rigid body instance and performing geometric alignment, and defining a frame-independent template; ; in, For frame-independent shape templates, For the first Foreground mask after frame geometry alignment For aggregation operators; ; in, For deformation mapping operators, For thresholding or voting operators; Introduce displacement penalty; ; in, As a displacement penalty, Let f be the center coordinates of the instance in frame f. For hyperparameters; For the first Frame instance center coordinates; Define rigid body guidance; ; in, For rigid body guidance, for The transpose of the matrix, It is the element-wise product of the displacement penalty Γ.
[0049] In step S4, the generation of non-rigid body guidance includes: establishing pixel-level correspondence in the feature space; ; in, For pixels At the corresponding position in the matched frame, Position of frame f The feature representation of the location, For the first Frame position Feature representation of the location; Define perceived deformation; ; in, To sense deformation, These are the pixel position coordinates; Obtained by interpolation of the corner points of the layout box: ; in, To induce deformation of the frame, It is a bilinear interpolation operator; Structural deformation penalty; ; in, As a penalty for deformation; Define non-rigid body guidance; ; in, Non-rigid body guidance, In order to deal with deformation penalty Element-wise product.
[0050] This embodiment is a preferred implementation method. By using structured parsing and cross-modal reasoning, it can more accurately model multi-entity relationships and causal events, making the generated content more consistent with the text narrative logic, avoiding semantic mismatch, and improving semantic matching accuracy.
[0051] Example 5 See Figure 1 A text-driven video generation method based on causal motion chain reasoning includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to the motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
[0052] In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
[0053] In step S2, the frame-level layout sequence is represented as follows: ; in, For frame-level layout sequences, For structured motion graphs and text prompts The layout reasoning output; layout evolution constraints include: Static instances satisfy Rigid body instances satisfy Non-rigid instances satisfy ; in, For frame number, For the first Layout of frame instance nodes The layout of the instance nodes in frame 1. For the first Layout of frame instance nodes For velocity vectors, For acceleration vectors, This is a summary of layout-level displacement or deformation.
[0054] In step S4, the generation of the static guide includes: selecting a reference frame index in the latent space; ; in, For the first The feature mapping results of the frame For the first The feature mapping results of the frame Reference frame index; For feature distance measurement; Construct static guidance; ;in: For static guiding masks; This is an indicator function.
[0055] In step S4, the generation of rigid body guidance includes: constructing a shape template for the rigid body instance and performing geometric alignment, and defining a frame-independent template; ; in, For frame-independent shape templates, For the first Foreground mask after frame geometry alignment For aggregation operators; ; in, For deformation mapping operators, For thresholding or voting operators; Introduce displacement penalty; ; in, As a displacement penalty, Let f be the center coordinates of the instance in frame f. For hyperparameters; For the first Frame instance center coordinates; Define rigid body guidance; ; in, For rigid body guidance, for The transpose of the matrix, It is the element-wise product of the displacement penalty Γ.
[0056] In step S4, the generation of non-rigid body guidance includes: establishing pixel-level correspondence in the feature space; ; in, For pixels At the corresponding position in the matched frame, Position of frame f The feature representation of the location, For the first Frame position Feature representation of the location; Define perceived deformation; ; in, To sense deformation, These are the pixel position coordinates; Obtained by interpolation of the corner points of the layout box: ; in, To induce deformation of the frame, It is a bilinear interpolation operator; Structural deformation penalty; ; in, As a penalty for deformation; Define non-rigid body guidance; ; in, Non-rigid body guidance, In order to deal with deformation penalty Element-wise product.
[0057] In step S4, the injection diffusion sampling process adopts a gradient update method: ; ; in, For the updated sequence of latent variables, For the current step, the sequence of latent variables. The gradient of loss L with respect to latent variables, To guide the loss function, For guiding coefficients, For normalization term, For attention maps, For static guiding mask, It is a rigid body guiding mask.
[0058] Preferably, in step S4, the injection diffusion sampling process achieves guided injection by modulating attention scores, which is suitable for video generation backbones of DiT architecture.
[0059] The motion categories include static instances, rigid body motion instances, and non-rigid body motion instances, which correspond to static guidance, rigid body guidance, and non-rigid body guidance, respectively, to achieve decoupled control of multiple instances and multiple motion categories.
[0060] The frame-level layout sequence is obtained by reasoning about causal motion chains based on structured motion graphs and text prompts using a large language model. The frame-level layout sequence is used to explicitly represent the spatial position and motion trajectory of each instance in time.
[0061] In step S1, the instance node includes an instance identifier, an attribute set, and a motion category label.
[0062] This embodiment represents the optimal implementation, designing differentiated guidance paths for three types of motion: static, rigid, and non-rigid. This addresses the issue of confusion between different motion categories in existing methods, achieving refined and independent control. Under the static guidance path, the feature consistency of the background or static objects in the video sequence is effectively constrained, ensuring stability in dynamic scenes and preventing unexpected changes with foreground movement, thus improving the cross-frame stability of static instances. For rigid motion instances, shape templates and displacement penalty mechanisms ensure that the object maintains its overall shape during movement, avoiding unexpected stretching or distortion. For non-rigid motion instances, establishing pixel-level correspondences and constructing deformation penalties allows for natural local deformation of the object, balancing flexibility and physical constraints.
[0063] By adopting a decompositional control paradigm of "planning before generation," the temporal continuity, physical rationality, and semantic accuracy of video generation are improved simultaneously without the need for training, greatly enhancing the ability to handle complex multi-entity interaction scenarios.
[0064] The basic principle of this invention is as follows: First, by parsing the user-input text prompts, instances, action attributes, and interaction relationships are extracted to construct a structured motion graph. This structured motion graph uses instances as nodes and relationships as edges, explicitly labeling the motion category of each instance, including stationary, rigid body, and non-rigid body types, providing structured prior knowledge for subsequent motion control. Based on this, a large language model is used for causal motion chain reasoning to generate frame-level layout sequences. These sequences not only define the spatial position of each instance in each frame but also apply corresponding evolutionary constraints based on different motion categories, such as stationary invariance, rigid body displacement constraints, and non-rigid body deformation constraints. This transforms the abstract text description into a temporally continuous and physically reasonable layout trajectory.
[0065] Secondly, frame-level layout sequences are mapped to three different types of guidance signals—static guidance, rigid guidance, and non-rigid guidance. For static instances, the object remains unchanged by constraining the consistency of its latent space features; for rigid instances, a shape template is constructed and a displacement penalty is introduced to ensure that the overall movement remains unchanged; for non-rigid instances, pixel-level correspondences are established in the feature space and a deformation penalty is constructed to allow natural shape changes. This categorized guidance mechanism allows multiple instances of different motion types in a scene to be controlled independently and precisely, avoiding mutual interference between different motion modes.
[0066] Finally, the three guiding signals mentioned above are injected into the sampling process of the diffusion model. By modulating the gradient update of the model's latent variables or the attention mechanism, the generation process is guided to strictly follow the preset layout trajectory and motion attributes. This injection method is compatible with mainstream video generation backbones based on the DiT architecture, i.e., the diffusion transformer, and can achieve fine control over the generated video content without compromising the generation capabilities of the pre-trained model. Overall, through "structural analysis - layout reasoning - classification guidance - controllable generation," the complex text-driven video generation problem is decoupled into a structured motion planning and decoupled motion control problem, thereby achieving high-quality and high-precision generation of multi-instance, multi-motion-category videos.
Claims
1. A text-driven video generation method based on causal motion chain reasoning, characterized in that, Includes the following steps: Step S1: Receive user text prompts, parse the instances, action attributes and interaction relationships in the text prompts, and construct a structured motion graph, which includes instance nodes and relationship edges; Step S2: Using the structured motion graph and text prompts as conditions, call the large language model to output a frame-level layout sequence. The frame-level layout sequence includes the spatial layout of each instance in each frame, and imposes constraints on the layout evolution according to different motion categories. Step S3: Map the frame-level layout sequence to the guidance mask in the generation stage, and generate static guidance, rigid guidance and non-rigid guidance according to motion category to achieve category-decoupled motion control. Step S4: Inject the static guidance, rigid guidance and non-rigid guidance into the diffusion sampling process, update the latent variables of the generation model through gradient or attention modulation, and output the video sequence.
2. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S1, the instance node is represented as follows: ; in, Label the sports category; The relation edge is represented as: ; in, For instance nodes, For relation edges.
3. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S2, the frame-level layout sequence is represented as follows: ; in, For frame-level layout sequences, For structured motion graphs and text prompts The layout reasoning output; layout evolution constraints include: Static instances satisfy Rigid body instances satisfy Non-rigid instances satisfy ; in, For frame number, For the first Layout of frame instance nodes The layout of the instance nodes in frame 1. For the first Layout of frame instance nodes For velocity vectors, For acceleration vectors, This is a summary of layout-level displacement or deformation.
4. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S4, the generation of the static guide includes: selecting a reference frame index in the latent space; ; in, For the first The feature mapping results of the frame For the first The feature mapping results of the frame Reference frame index; For feature distance measurement; Construct static guidance; ;in: For static guiding masks; This is an indicator function.
5. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S4, the generation of rigid body guidance includes: constructing a shape template for the rigid body instance and performing geometric alignment, and defining a frame-independent template; ; in, For frame-independent shape templates, For the first Foreground mask after frame geometry alignment For aggregation operators; ; in, For deformation mapping operators, For thresholding or voting operators; Introduce displacement penalty; ; in, As a displacement penalty, Let f be the center coordinates of the instance in frame f. For hyperparameters; For the first Frame instance center coordinates; Define rigid body guidance; ; in, For rigid body guidance, for The transpose of the matrix, It is the element-wise product of the displacement penalty Γ.
6. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S4, the generation of non-rigid body guidance includes: establishing pixel-level correspondence in the feature space; ; in, For pixels At the corresponding position in the matched frame, Position of frame f The feature representation of the location, For the first Frame position Feature representation of the location; Define perceived deformation; ; in, To sense deformation, These are the pixel position coordinates; Obtained by interpolation of the corner points of the layout box: ; in, To induce deformation of the frame, It is a bilinear interpolation operator; Structural deformation penalty; ; in, As a penalty for deformation; Define non-rigid body guidance; ; in, Non-rigid body guidance, In order to deal with deformation penalty Element-wise product.
7. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S4, the injection diffusion sampling process adopts a gradient update method: ; ; in, For the updated sequence of latent variables, For the current step, the sequence of latent variables. The gradient of loss L with respect to latent variables, To guide the loss function, For guiding coefficients, For normalization term, For attention maps, For static guiding mask, It is a rigid body guiding mask.
8. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S4, the injection diffusion sampling process achieves guided injection by modulating attention scores, which is suitable for video generation backbones of DiT architecture.
9. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: The motion categories include static instances, rigid body motion instances, and non-rigid body motion instances, which correspond to static guidance, rigid body guidance, and non-rigid body guidance, respectively, to achieve decoupled control of multiple instances and multiple motion categories.
10. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: The frame-level layout sequence is obtained by reasoning about causal motion chains based on structured motion graphs and text prompts using a large language model. The frame-level layout sequence is used to explicitly represent the spatial position and motion trajectory of each instance in time.
11. The text-driven video generation method based on causal motion chain reasoning according to claim 1, characterized in that: In step S1, the instance node includes an instance identifier, an attribute set, and a motion category label.
Citation Information
Patent Citations
Method and device for generating first screen video background based on page elements
CN121725084A