Instance-level motion-controllable video generation method, system, medium and equipment

By acquiring instance motion control conditions and using pre-trained image and video diffusion generation models, generating candidate keyframes and optimizing videos, the problem of difficult instance layout control in the prior art and poor video motion effect is solved, and high-quality instance-level motion controllable video generation is achieved.

CN120075549APending Publication Date: 2025-05-30SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510206227.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing controllable video generation method based on training is difficult to achieve instance layout control, and the generated video has poor motion effect, lacks realism and delicateness, and cannot meet the creation needs of high-quality visual content.

Method used

By obtaining instance motion control conditions, including instance layout information and motion trajectory, a pre-trained positioning image diffusion generation model and video diffusion generation model that introduces inter-frame mutual attention mechanism are used to generate candidate keyframes and optimize videos to achieve instance-level motion controllable video generation.

Benefits of technology

It improves the accuracy of layout control and trajectory control, improves the motion fluency and detail richness of generated videos, and realizes high-quality instance-level customized video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075549A_ABST
    Figure CN120075549A_ABST
Patent Text Reader

Abstract

The invention provides an instance-level motion-controllable video generation method and system, a medium and equipment, and the method comprises the steps: obtaining an instance motion control condition which comprises instance layout information and an instance motion track; augmenting the instance layout information along an instance movement track, and determining the augmented instance layout information; inputting the augmented instance layout information into a pre-trained positioning image diffusion generation model into which an inter-frame mutual attention mechanism is introduced, and generating candidate key frames; inputting the candidate key frames into a pre-trained video diffusion generation model in which an inter-frame mutual attention mechanism is introduced, and determining a first motion video; and performing optimization processing on the first motion video according to the motion priori of the pre-trained video diffusion generation model and the detail priori of the pre-trained image diffusion generation model, and determining a target motion video. According to the invention, the motion control capability of video generation is improved, the video generation quality is improved, and instance-level motion customized video generation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, system, medium, and device for generating videos with controllable instance-level motion. Background Art

[0002] With the rapid development of generative model technology, significant progress has also been made in the field of video generation. Among them, diffusion models and their derivative architectures have demonstrated excellent capabilities in video generation tasks. These models transform the original data into Gaussian noise through a gradually adding-noise process, and then reconstruct the video from the noise through the reverse process. This process not only enables the model to capture complex spatio-temporal dependencies but also accurately models the continuity and consistency between video frames. As the quality of video generation continues to improve, people have begun to explore more complex customized video generation tasks. Among them, one of the most challenging tasks is to achieve instance-level motion controllable video generation. This not only requires the model to create instances (such as people, objects) according to the specified layout information but also ensures that the motion of the instances in the video conforms to the given motion path.

[0003] To achieve instance-level motion controllable video generation, some research methods have explored methods of jointly training video diffusion models by taking layout and trajectory information as additional conditions. Among them, a representative work is the paper "Draganything: Motion control for anything using entity representation" published by W. Wu et al. at the European Conference on Computer Vision (ECCV 2024) in 2024. By introducing layout and trajectory annotations into the existing video dataset and introducing an additional attention mechanism layer into the decoder of the video diffusion model noise estimation backbone network for end-to-end training, the model's ability to understand and process trajectory inputs is achieved. However, for the model to further master the ability of layout controllable video generation, a large number of videos with layout information annotations need to be additionally produced, and additional computational resource overhead is required, resulting in the fact that the existing training-based controllable video generation methods can only achieve preliminary instance trajectory control and cannot achieve instance layout control.

[0004] To address this issue, some research methods have explored making full use of the capabilities of pre-trained video diffusion models to achieve instance layout control and trajectory control in a zero-shot manner. Among them, a representative work is the paper "Peekaboo: Interactive video generation via masked-diffusion" proposed by Y. Jain et al. at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR2024) in 2024. It presents an instance motion control method based on attention map manipulation, which can adaptively adjust the attention map of each frame according to the layout and trajectory, enabling the accurate presentation of specified instances in the specified position area while suppressing interference from irrelevant factors such as the background, thus achieving both instance layout control and trajectory control. However, such methods still have the following deficiencies: (1) The layout control and trajectory control are inaccurate, which not only limits the diversity and flexibility of the generated content but also makes the generated videos difficult to meet the application scenarios with high-precision requirements; (2) The motion effect of the generated videos is poor, there are missing details in the video frames, and the videos lack realism and fineness, which is not conducive to the creation needs of high-quality visual content; (3) The functions are single, only supporting layout information of the input bounding box type, and can only control the overall motion of the instance, unable to control the local motion, which limits the practicality and scalability of the method. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the purpose of the present disclosure is to provide a video generation method, system, medium, and device with controllable instance-level motion.

[0006] To achieve the above object, according to one aspect of the present disclosure, there is provided a video generation method with controllable instance-level motion, including:

[0007] Obtain instance motion control conditions, where the instance motion control conditions include instance layout information and instance motion trajectories;

[0008] Augment the instance layout information along the instance motion trajectory to determine the augmented instance layout information;

[0009] Input the augmented instance layout information into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to generate candidate key frames;

[0010] Input the candidate key frames into a pre-trained video diffusion generation model introducing an inter-frame mutual attention mechanism to determine the first motion video;

[0011] Optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video.

[0012] Optionally, the instance layout information includes position markers, and the position markers are in the form of bounding boxes or masks.

[0013] Optionally, augmenting the instance layout information along the instance motion trajectory to determine the augmented instance layout information includes:

[0014] At the endpoints of the instance motion trajectory where there are no position markers, copy the position markers of the nearest neighbor endpoints of the endpoints to the endpoints to determine the augmented instance layout information.

[0015] Optionally, the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism includes a noise estimation backbone network and a gated self-attention mechanism module, and the noise estimation backbone network includes an inter-frame mutual attention mechanism layer.

[0016] Optionally, inputting the augmented instance layout information into the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism to generate candidate key frames includes:

[0017] Input the augmented instance layout information into the noise estimation backbone network to determine the visual tokens of the noise estimation backbone network;

[0018] Input the augmented instance layout information into the encoder of the gated self-attention mechanism module to determine the encoded layout embedding;

[0019] Input the layout embedding into the tokenizer of the gated self-attention mechanism module to determine the layout tokens;

[0020] Input the visual tokens of the noise estimation backbone network and the layout tokens into the masked self-attention mechanism layer to determine the fused tokens;

[0021] Use the fused tokens as the input for the next denoising time step, and repeat the denoising operation to generate the candidate key frames.

[0022] Optionally, inputting the candidate key frames into the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism to determine the first motion video includes:

[0023] Generate a motion video according to random Gaussian noise using the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism;

[0024] Extract the motion information guidance and texture information guidance of the candidate key frames;

[0025] Use the motion information guidance and the texture information guidance of the candidate key frames to guide the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism in the denoising process of the motion video, and generate the first motion video.

[0026] Optionally, the extraction of the motion information guidance and texture information guidance of the candidate key frames includes:

[0027] Obtain the foreground of the candidate key frame;

[0028] Convert the foreground of the candidate key frame into an edge map using a preset edge detection algorithm;

[0029] Input the edge map into the pre-trained ControlNet to determine the motion information guidance;

[0030] Encode the first frame of the candidate key frame into an image embedding;

[0031] Add Gaussian noise to the image embedding to determine the image embedding with added Gaussian noise;

[0032] Input the image embedding with added Gaussian noise into the pre-trained IP-Adapter to determine the texture information guidance.

[0033] Optionally, the motion prior of the pre-trained video diffusion generation model is used to inject motion information into the first motion video, and the texture prior of the pre-trained localization image diffusion generation model is used to inject texture information into the first motion video.

[0034] Optionally, the optimization process of the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video includes:

[0035] Perform encoding processing on the first motion video to determine a video embedding;

[0036] Add Gaussian noise to the video embedding as the initial noise latent embedding for the denoising process of the pre-trained video diffusion model;

[0037] Perform a single-step denoising operation on the noise latent embedding at each denoising time step to determine a new noise latent embedding;

[0038] Perform denoising operations on the new noise latent embedding of the denoising time step with motion information injection for a first preset number of times, and perform weighted summation with the new noise latent embedding of the current denoising time step to determine the noise latent embedding for motion information injection, where the noise latent embedding for motion information injection is used for denoising in the next denoising time step or for texture information injection;

[0039] For the denoising time step with texture information injection, determine the noise-free latent embedding according to the new noise latent embedding corresponding to the current denoising time step or the noise latent embedding for motion information injection corresponding to the current denoising time step;

[0040] Add Gaussian noise to the noise-free latent embedding again to determine the image noise latent embedding;

[0041] Perform denoising operations on the image noise latent embedding for a second preset number of times to determine the denoised image noise latent embedding;

[0042] According to the denoised image noise latent embedding, determine the image noise-free latent embedding corresponding to the denoised image noise latent embedding;

[0043] Add Gaussian noise to the image noise-free latent embedding again to determine the denoised video noise latent embedding;

[0044] Perform weighted summation on the new noise latent embedding corresponding to the current denoising time step or the noise latent embedding for motion information injection corresponding to the current denoising time step and the denoised video noise latent embedding to determine the noise latent embedding for texture information injection at the current denoising time step, where the noise latent embedding for texture information injection is used for denoising in the next denoising time step until the denoising operations for all denoising time steps are completed to determine the target motion video.

[0045] According to a second aspect of the present disclosure, there is provided an instance-level motion controllable video generation system, including:

[0046] A first acquisition module, configured to acquire instance motion control conditions, where the instance motion control conditions include instance layout information and instance motion trajectories;

[0047] An instance layout information augmentation module, configured to augment the instance layout information along the instance motion trajectories to determine the augmented instance layout information;

[0048] A candidate key frame generation module, configured to input the augmented instance layout information into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to generate candidate key frames;

[0049] The first motion video generation module is configured to input the candidate key frames into a pre-trained video diffusion generation model that introduces an inter-frame mutual attention mechanism to determine a first motion video;

[0050] The target motion video generation module is configured to optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine a target motion video.

[0051] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.

[0052] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:

[0053] A memory having a computer program stored thereon;

[0054] A processor configured to execute the computer program in the memory to implement the steps of the method provided in the first aspect of the present disclosure.

[0055] Compared with the prior art, the embodiments of the present disclosure have at least one of the following beneficial effects:

[0056] Through the above technical solutions, the instance layout information and the instance motion trajectory are used as instance motion control conditions to realize the motion control of the instance as a whole and the instance part; by using a pre-trained localization image diffusion generation model that introduces an inter-frame mutual attention mechanism to generate candidate key frames based on the augmented instance layout information, the accuracy of layout control and trajectory control is improved; optimizing the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model can efficiently improve the motion smoothness and detail richness of the generated video at a low cost, realize high-quality instance-level motion customization video generation, and improve the video generation quality.

[0057] In the embodiments of the present disclosure, the instance motion control conditions support the input of various types of control conditions such as bounding boxes or masks, and at the same time, through the control of the instance motion trajectory, the motion control of the instance part is realized, thereby realizing more refined instance motion customization, which is beneficial to practical application and deployment.

[0058] In the embodiments of the present disclosure, during the process of generating the first motion video by using the pre-trained video diffusion generation model, motion information guidance and texture information guidance are extracted from the candidate key frames and used to guide the generation of the first motion video, ensuring texture consistency and improving the accuracy of layout control and trajectory control. Description of the Drawings

[0059] Other features, objects, and advantages of the present disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0060] Figure 1 is a schematic flowchart of a method for generating an instance-level motion-controllable video shown according to an exemplary embodiment.

[0061] Figure 2 is a schematic diagram of the effect of generating a video using a method for generating an instance-level motion-controllable video shown according to an exemplary embodiment.

[0062] Figure 3 is a block diagram of a system for generating an instance-level motion-controllable video shown according to an exemplary embodiment. Detailed Embodiments

[0063] The present disclosure will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present disclosure, but do not limit the present disclosure in any way. It should be noted that, for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can be made. These all belong to the protection scope of the present disclosure.

[0064] Figure 1 is a schematic flowchart of a method for generating an instance-level motion-controllable video shown according to an exemplary embodiment.

[0065] As Figure 1 shown, the present disclosure provides a method for generating an instance-level motion-controllable video, including S11 to S15.

[0066] S11, obtaining instance motion control conditions.

[0067] Among them, the instance motion control conditions include instance layout information and instance motion trajectories.

[0068] The instance layout information includes position markers, and the position markers are in the form of bounding boxes or masks. The instance layout information consists of a series of position markers and is used to mark the position layout of the instance in the video.

[0069] The instance motion trajectory is composed of a series of end points, and each adjacent two end points are connected by an arrow. The instance motion trajectory is used to mark the motion path of the instance.

[0070] There is a constraint relationship between the instance layout information and the instance motion trajectory. The geometric center of each position marker in the instance layout information belongs to the set of end points of the instance motion trajectory.

[0071] S12, augmenting the instance layout information along the instance motion trajectory to determine the augmented instance layout information.

[0072] S13. Input the augmented instance layout information into a pre-trained localization image diffusion generation model with an inter-frame mutual attention mechanism to generate candidate key frames.

[0073] Among them, an inter-frame mutual attention mechanism is introduced into the pre-trained localization image diffusion generation model to improve the texture consistency of the candidate key frames.

[0074] S14. Input the candidate key frames into a pre-trained video diffusion generation model with an inter-frame mutual attention mechanism to determine the first motion video.

[0075] Among them, the first motion video is represented as a rough motion video.

[0076] S15. Optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video.

[0077] Among them, the target motion video is represented as a motion video with high quality.

[0078] The pre-trained video diffusion generation model is a video diffusion generation model based on it, and the pre-trained image diffusion generation model is an image diffusion generation model based on it.

[0079] Through the above technical solutions, using the instance layout information and the instance motion trajectory as the instance motion control conditions, the motion control of the whole instance and the local part of the instance is realized; by using a pre-trained localization image diffusion generation model with an inter-frame mutual attention mechanism to generate candidate key frames based on the augmented strength layout information, the accuracy of layout control and trajectory control is improved; optimizing the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model can efficiently improve the motion smoothness and detail richness of the generated video at a low cost, realize high-quality instance-level motion customized video generation, and improve the video generation quality.

[0080] In a possible embodiment, S11. Obtaining the instance motion control conditions includes:

[0081] Obtain instance layout information {s i | i ∼ [1, 2,..., S]} and instance motion trajectory {l i | i ∼ [1, 2,..., L]}.

[0082] Among them, s i represents a position marker, such as a bounding box, S represents the number of position markers, and l iRepresents a segment of a motion trajectory, including two endpoints, connected by an arrow between the two endpoints, and L represents the number of trajectories.

[0083] The example motion control conditions of the present disclosure support multiple types of control condition inputs for bounding boxes or masks. Meanwhile, through the control of the example motion trajectory, the motion control of the local part of the example is realized, and then more refined example motion customization is achieved, which is beneficial to practical application and deployment.

[0084] In a possible embodiment, S12, augment the example layout information along the example motion trajectory to determine the augmented example layout information, including:

[0085] At the endpoints of the example motion trajectory where there is no position mark, copy the position mark of the nearest neighbor endpoint of the endpoint to the endpoint to determine the augmented example layout information.

[0086] Among them, for the endpoints in the example motion trajectory, if there is no position mark at the endpoint, augmentation is performed at this endpoint to determine the augmented example layout information {s i |i~[1,2,...,S']}, where S' represents the number of position marks in the augmented example layout information.

[0087] In a possible embodiment, the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism includes a noise estimation backbone network and a gated self-attention mechanism module, and the noise estimation backbone network includes an inter-frame mutual attention mechanism layer.

[0088] In the present disclosure, the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism is obtained after fine-tuning the pre-trained image diffusion generation model. Replace the self-attention mechanism layer in the noise estimation backbone network of the pre-trained image diffusion generation model with an inter-frame mutual attention mechanism layer, and add an additional learnable gated self-attention mechanism module on one side of the inter-frame mutual attention mechanism layer. The output of the inter-frame mutual attention mechanism layer and the output of the learnable gated self-attention mechanism module are summed residually to be adjusted to the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism.

[0089] For the inter-frame mutual attention mechanism layer of the noise estimation backbone network of the pre-trained localization image diffusion generation model introducing the inter-frame mutual attention mechanism:

[0090]

[0091] Among them, Q represents the query, K represents the key, V represents the value, and c represents the text embedding.

[0092] In the inter-frame mutual attention mechanism layer of the noise estimation backbone network, the first-frame candidate key frame is used to provide keys and values for subsequent candidate key frames, improving the texture consistency between candidate key frames.

[0093] In a possible embodiment, in S13, the augmented instance layout information is input into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to generate candidate key frames, including S131 to S135.

[0094] S131, input the augmented instance layout information into the noise estimation backbone network to determine the visual tokens of the noise estimation backbone network.

[0095] Among them, the augmented instance layout information {s i |i~[1,2,...,S']} is first input into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism, and then input into the noise estimation backbone network.

[0096] The visual tokens of the noise estimation backbone network represent the predicted values of the image noise by the noise estimation backbone network at the current denoising time step, and reflect high-level semantic features such as the object contour and scene structure of the image by modeling the spatial distribution and structural characteristics of the noise.

[0097] S132, input the augmented instance layout information into the encoder of the gated self-attention mechanism module to determine the encoded layout embedding.

[0098] Among them, the augmented instance layout information {s i |i~[1,2,...,S']} is first input into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism, and then input into the learnable gated self-attention mechanism module.

[0099] Specifically, sample the augmented instance layout information to determine a two-dimensional point sequence;

[0100] The two-dimensional point sequence is successively subjected to Fourier mapping and passed through the encoder to be encoded into a layout embedding.

[0101] S133, input the layout embedding into the tokenizer of the gated self-attention mechanism module to determine the layout tokens.

[0102] Among them, the layout tokens represent the encoded values of the spatial position, category information, and attributes of the instance, and are used to guide the denoising process of the pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to ensure that the generated image conforms to the input layout constraints.

[0103] S134, input the visual tokens and layout tokens of the noise estimation backbone network into the masked self-attention mechanism layer of the gated self-attention mechanism to determine the fused tokens.

[0104] S135. Use the fusion token as the input for the next denoising time step, and repeat the denoising operation to generate candidate key frames.

[0105] Based on the above steps S131 to S135, perform the denoising process of the pre-trained localization image diffusion generation model with an inter-frame mutual attention mechanism. The above steps S131 to S134 can be repeated for denoising. After the denoising process is completed, the generated noise-free image is the candidate key frame {f i |i ∼ [1, 2,..., S']}.

[0106] In a possible embodiment, S14. Input the candidate key frames into the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism to determine the first motion video, including: S141 to S143.

[0107] S141. According to the random Gaussian noise, use the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism to generate a motion video.

[0108] Among them, the first motion video is generated starting from the random Gaussian noise using the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism, and the generated motion video is a rough motion video.

[0109] S142. Extract the motion information guidance and texture information guidance of the candidate key frames.

[0110] This step S142 is executed during the generation process of the motion video in step S141. The motion information guidance of the candidate key frames is also represented as motion information embedding, which is used to guide the rough motion video to meet the instance motion control condition input; the texture information guidance is also represented as texture information embedding, which is used to guide the texture consistency of the rough motion video frames.

[0111] S143. Use the motion information guidance and texture information guidance of the candidate key frames to guide the denoising process of the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism for the motion video, and generate the first motion video.

[0112] Among them, the first motion video is also a rough motion video and needs to be further optimized to obtain a high-quality motion video.

[0113] This disclosure extracts the motion information guidance and texture information guidance from the candidate key frames during the process of generating a motion video using the pre-trained video diffusion generation model with an inter-frame mutual attention mechanism, and guides the generation of the first motion video for the motion video, ensuring texture consistency and improving the accuracy of layout control and trajectory control.

[0114] In a possible embodiment, S142, extracting the motion information guidance and texture information guidance of candidate key frames includes: S101 to S106.

[0115] S101, obtaining the foreground of the candidate key frame.

[0116] Among them, removing the background of the candidate key frames {f i |i ~ [1, 2,..., S']}, and extracting the foreground of the candidate key frames {f i |i ~ [1, 2,..., S']}.

[0117] S102, converting the foreground of the candidate key frame into an edge map by using a preset edge detection algorithm.

[0118] S103, inputting the edge map into the pre-trained ControlNet to determine the motion information guidance.

[0119] Among them, the motion information guidance is also expressed as motion information embedding, and the motion information embedding is added to the embedding of the noise estimation backbone network of the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism in the form of weighted summation to guide the denoising process.

[0120] S104, encoding the first frame of the candidate key frame into an image embedding.

[0121] Among them, the first frame of the candidate key frame is expressed as f 1 .

[0122] S105, adding Gaussian noise to the image embedding to determine the image embedding with added Gaussian noise.

[0123] Among them, the image embedding with added Gaussian noise is expressed as

[0124] As an example, for the initial Gaussian noise, that is, the random Gaussian noise generated in step S141 for generating the motion video, independent and identically distributed Gaussian noise can be sampled where N represents the length of the video to be generated, and and are concatenated in the time dimension to serve as the initial Gaussian noise for the denoising process of the pre-trained video diffusion model introducing the inter-frame mutual attention mechanism

[0125] S106, inputting the image embedding with added Gaussian noise into the pre-trained IP-Adapter to determine the texture information guidance.

[0126] Among them, texture information guidance is also expressed as texture information embedding, and the texture information embedding is added to the embedding of the noise estimation backbone network of the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism in the form of weighted summation to guide the denoising process.

[0127] In the present disclosure, the self-attention mechanism layer included in the intermediate block of the noise estimation backbone network of the pre-trained video diffusion generation model is replaced with an inter-frame mutual attention mechanism layer as the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism.

[0128] Adopt the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism, and use the first-frame motion video to provide keys and values for subsequent frames to generate a rough first motion video, improving the texture consistency of the rough motion video.

[0129] In a possible embodiment, S143, using the motion information guidance and texture information guidance of the candidate key frame to guide the denoising process of the motion video by the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism to generate the first motion video may include:

[0130] Add the motion information embedding and the texture information embedding to the embedding of the noise estimation backbone network of the pre-trained video diffusion generation model introducing the inter-frame mutual attention mechanism in the form of weighted summation to guide the denoising process and generate the first motion video.

[0131] In a possible embodiment, the motion prior of the pre-trained video diffusion generation model is used to inject motion information into the first motion video, and the texture prior of the pre-trained image diffusion generation model is used to inject texture information into the first motion video.

[0132] Inject motion information into the first motion video through the motion prior of the pre-trained video diffusion generation model, that is, inject motion information into the rough motion video to improve the motion smoothness of the rough motion video.

[0133] Inject texture information into the first motion video through the texture prior of the pre-trained image diffusion generation model, that is, inject texture information into the rough motion video to improve the detail richness of the frames of the rough motion video.

[0134] In a possible embodiment, S15, optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video, which may include S151 to S160.

[0135] S151, perform encoding processing on the first motion video to determine the video embedding.

[0136] Among them, the encoder of the pre-trained video diffusion generation model can be used to encode the first motion video.

[0137] S152, Add Gaussian noise to the video embedding as the initial noise latent embedding for the denoising process of the pre-trained video diffusion model.

[0138] Among them, the initial noise latent embedding is denoted as z T .

[0139] S153, Perform a single-step denoising operation on the noise latent embedding at each denoising time step to determine the new noise latent embedding.

[0140] As an example, for each denoising time step t, perform a denoising operation on the current noise latent embedding z t to obtain the new noise latent embedding z t-1 .

[0141] If there is no motion information injection and texture information injection within the denoising time step, the new noise latent embedding of the current denoising time step can be used for denoising in the next denoising time step.

[0142] S154, Perform the first preset number of denoising operations on the new noise latent embedding of the denoising time step with motion information injection, and perform residual summation with the new noise latent embedding of the current denoising time step to determine the noise latent embedding for motion information injection.

[0143] Among them, the noise latent embedding for motion information injection is used for denoising in the next denoising time step or for texture information injection.

[0144] If there is only motion information injection within the denoising time step, the noise latent embedding for motion information injection can be used for denoising in the next denoising time step.

[0145] If there is not only motion information injection but also texture information injection within the denoising time step, the noise latent for motion information injection is used for texture information injection.

[0146] As an example, if motion information injection needs to be performed, input the new noise latent embedding z t-1 into the noise estimation backbone network of the pre-trained video diffusion model to perform M steps of denoising to obtain z t-1-M , and perform weighted summation of z t-1-M with the new noise latent embedding z t-1 of the current denoising time step to determine the noise latent embedding for motion information injection

[0147]

[0148] Among them, λ1 represents the first weighting coefficient.

[0149] S155. For the denoising time step with texture information injection, determine the noise-free latent embedding according to the new noise latent embedding corresponding to the current denoising time step or the noise latent embedding with motion information injection.

[0150] Among them, if there is no motion information injection in the current denoising time step, estimate the noise-free latent embedding according to the new noise latent embedding corresponding to the current denoising time step. If there is motion information injection in the current denoising time step, estimate the noise-free latent embedding according to the noise latent embedding with motion information injection corresponding to the current denoising time step.

[0151] As an example, if there is motion information injection in the current denoising time step, then according to the noise latent embedding with motion information injection corresponding to the current denoising time step t estimate the noise-free latent embedding

[0152]

[0153] Among them, represents the noise-free latent embedding corresponding to the noise latent embedding with motion information injection at the current denoising time step t, and α represents the noise coefficient of the pre-trained video diffusion model, and ∈ t-1 represents the noise coefficient of the pre-trained video diffusion model, and c represents the text embedding. θ represents the noise estimation backbone network of the pre-trained video diffusion model, and c represents the text embedding.

[0154] S156. Re-add Gaussian noise to the noise-free latent embedding to determine the image noise latent embedding.

[0155] As an example, the noise scheduler of the noise estimation backbone network of the pre-trained image diffusion model can be used to re-add Gaussian noise to the noise-free latent embedding to obtain the image noise latent embedding that is, obtain the image noise embedding corresponding to the video noise latent embedding in the pre-trained image diffusion model

[0156]

[0157] Among them, represents the image noise latent embedding, and α t-1 represents the noise coefficient of the pre-trained video diffusion model, and ∈ θ represents the noise estimation backbone network of the pre-trained video diffusion model.

[0158] S157. Perform the denoising operation for the second preset number of times on the image noise latent embedding to determine the denoised image noise latent embedding.

[0159] As an example, the image noise is potentially embedded into the noise estimation backbone network of the pre-trained image diffusion model, and denoising is performed D times to determine the denoised image noise potential embedding

[0160] S158. According to the denoised image noise potential embedding, determine the image noise-free potential embedding corresponding to the denoised image noise potential embedding

[0161] The method in step S155 above can be referred to estimate the image noise-free potential embedding corresponding to the denoised image noise potential embedding, which will not be elaborated here

[0162] S159. Add Gaussian noise to the image noise-free potential embedding again to determine the denoised video noise potential embedding

[0163] As an example, the noise scheduler of the noise estimation backbone network of the pre-trained video diffusion model can be used to add Gaussian noise to the image noise-free potential embedding again to determine the denoised video noise potential embedding

[0164] S160. Perform weighted summation on the new noise potential embedding corresponding to the current denoising time step or the noise potential embedding injected with motion information corresponding to the current denoising time step and the denoised video noise potential embedding to determine the noise potential embedding injected with texture information corresponding to the current denoising time step

[0165] Among them, the noise potential embedding injected with texture information is used for denoising in the next denoising time step. Until the denoising operations for all denoising time steps are completed, the target motion video is determined

[0166] The denoised video noise potential embedding corresponding to the denoised image noise potential embedding can be obtained by performing the inverse process of steps S156 to S157, that is, obtaining the denoised image noise potential embedding corresponding video noise embedding in the pre-trained video diffusion generation model and the new noise potential embedding corresponding to the current denoising time step or the noise potential embedding injected with motion information corresponding to the current denoising time step and the video noise embedding in the pre-trained video diffusion generation model Perform weighted summation. In this embodiment, the noise potential embedding injected with motion information corresponding to the current denoising time step is used and the video noise embedding in the pre-trained video diffusion generation model Perform weighted summation:

[0167]

[0168] Among them, λ 2 represents the second weighting coefficient.

[0169] By repeatedly executing the above steps S153 to S158 until all denoising time steps are completed, a target motion video, that is, a high-quality motion video, is obtained.

[0170] In a possible embodiment, a method for generating an instance-level motion-controllable video provided by the present disclosure, as well as the methods of T2V-Zero, Peekaboo, and TrailBlazer, are verified on the public datasets DAVIS-17 and GOT10k.

[0171] The test protocol for the experiment is as follows: Extract example motion control conditions inputs and text descriptions from the video data of the dataset, generate videos according to the instance motion trajectories and text descriptions, and each test protocol evaluates the differences between the generated videos and the real videos.

[0172] Experiments are conducted on the DAVIS-17 and GOT10k datasets. The evaluation metrics include the Fréchet Video Distance (FVD) between videos, the Kernel Inception Distance (KID), the CLIP Similarity (CLIPSim), the Mean Intersection of Union (mIoU), and the Centroid Distance (CD). Among them, the smaller the FVD, KID, and CD between videos, and the larger the CLIPSim and mIoU, the better the performance of the method. The specific comparison results are shown in Table 1:

[0173]

[0174] Table 1

[0175] As shown in Table 1, a method for generating an instance-level motion-controllable video provided by the present disclosure has more excellent evaluation results on the quality of test protocol video generation on the DAVIS-17 and GOT10k datasets, better inter-frame consistency, and a higher matching degree with the instance motion control condition inputs, achieving more excellent video generation quality.

[0176] Figure 2 is a schematic diagram of the effect of generating a video using a method for generating an instance-level motion-controllable video shown according to an exemplary embodiment.

[0177] As Figure 2The generated video using the instance-level motion controllable video generation method provided by the present disclosure is shown. The instance-level motion controllable video generation method provided by the present disclosure can accurately generate a video that meets the input of the instance motion control conditions, generate more realistic actions and forms, and also greatly improve the consistency and detail richness of the generated video. At the same time, the instance-level motion controllable video generation method provided by the present disclosure can not only control the overall motion of the instance, but also control the local motion of the instance; it can not only support bounding box input, but also support mask input, realizing higher customized video generation.

[0178] The above experiments show that the instance-level motion controllable video generation method proposed by the present disclosure can more accurately generate a video that meets the instance motion control conditions in different scenarios, support the motion control of the whole and local parts of the instance, support different input types, and at the same time achieve higher video generation quality, motion consistency and detail richness, thus having a wider application scenario.

[0179] Figure 3 It is a block diagram of an instance-level motion controllable video generation system shown according to an exemplary embodiment.

[0180] Based on the same concept, the present disclosure also provides an instance-level motion controllable video generation system 100, as Figure 3 shown, including: a first acquisition module 110, an instance layout information augmentation module 120, a candidate key frame generation module 130, a first motion video generation module 140, and a target motion video generation module 150.

[0181] The first acquisition module 110 is used to acquire instance motion control conditions, and the instance motion control conditions include instance layout information and instance motion trajectories;

[0182] The instance layout information augmentation module 120 is used to augment the instance layout information along the instance motion trajectory to determine the augmented instance layout information;

[0183] The candidate key frame generation module 130 is used to input the augmented instance layout information into a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to generate candidate key frames;

[0184] The first motion video generation module 140 is used to input the candidate key frames into a pre-trained video diffusion generation model introducing an inter-frame mutual attention mechanism to determine the first motion video;

[0185] The target motion video generation module 150 is used to optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video.

[0186] Through the above technical solution, the instance layout information and the instance motion trajectory are used as the instance motion control conditions to achieve the motion control of the whole instance and the local instance; by using a pre-trained localization image diffusion generation model introducing an inter-frame mutual attention mechanism to generate candidate key frames based on the augmented instance layout information, the accuracy of layout control and trajectory control is improved; according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model, the first motion video is optimized, which can improve the motion smoothness and detail richness of the generated video at a low cost and efficiently, realize high-quality instance-level motion-customized video generation, and improve the video generation quality.

[0187] Regarding the embodiments of the above system, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0188] Based on the same concept as above, in another embodiment of the present disclosure, an electronic device is further provided, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the program, it is used to execute the instance-level motion controllable video generation method.

[0189] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random access memory (English: random-access memory, abbreviation: RAM), such as static random access memory (English: static random-access memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as application programs and functional modules for implementing the above method), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in one or more memories in partitions. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0190] The above computer programs, computer instructions, etc. can be stored in one or more memories in partitions. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0191] A processor is configured to execute a computer program stored in a memory to implement each step in the method described in the above embodiments. For specific details, reference may be made to the relevant descriptions in the foregoing method embodiments.

[0192] The processor and the memory may be of an independent structure or an integrated structure integrated together. When the processor and the memory are of an independent structure, the memory and the processor may be coupled through a bus.

[0193] In an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a method for generating a video with instance-level motion controllability in any of the above embodiments.

[0194] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0195] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0196] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0197] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or steps for implementing the functions specified in one block or more blocks.

[0198] Although the preferred embodiments of the present disclosure have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0199] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.

Claims

1. A method for generating video with instance-level motion controllable, characterized in that: include: Acquire instance motion control conditions, where the instance motion control conditions include instance layout information and instance motion trajectory; Augmenting the instance layout information along the instance motion trajectory to determine augmented instance layout information; The augmented instance layout information is input into a pre-trained positioning image diffusion generation model of an inter-frame mutual attention mechanism to generate candidate key frames; The candidate key frame input is introduced into a pre-trained video diffusion generation model of an inter-frame mutual attention mechanism to determine a first motion video; The first motion video is optimized according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video.

2. The method according to claim 1, characterized in that The instance layout information includes a position mark, and the position mark is a bounding box or a mask; The augmenting the instance layout information along the instance motion trajectory to determine the augmented instance layout information includes: At an endpoint of the instance motion trajectory where no position mark exists, the position mark of the nearest neighboring endpoint of the endpoint is copied to the endpoint to determine the augmented instance layout information.

3. The method according to claim 1, characterized in that The pre-trained positioning image diffusion generation model that introduces the inter-frame mutual attention mechanism includes a noise estimation backbone network and a gated self-attention mechanism module, and the noise estimation backbone network includes an inter-frame mutual attention mechanism layer; The step of introducing the augmented instance layout information input into a pre-trained positioning image diffusion generation model of an inter-frame mutual attention mechanism to generate candidate key frames includes: Inputting the augmented instance layout information into the noise estimation backbone network to determine visual tokens of the noise estimation backbone network; Inputting the augmented instance layout information into the encoder of the gated self-attention mechanism module to determine the encoded layout embedding; The layout embedding is input into the tokenizer of the gated self-attention mechanism module to determine the layout token; Inputting the visual token of the noise estimation backbone network and the layout token into the mask self-attention mechanism layer of the gated self-attention mechanism module to determine a fusion token; The fused token is used as the input of the next denoising time step, and the denoising operation is repeated to generate the candidate key frame.

4. The method according to claim 1, characterized in that: The step of introducing the candidate key frame input into the pre-trained video diffusion generation model of the inter-frame mutual attention mechanism to determine the first motion video includes: Generate motion video using the pre-trained video diffusion generation model that introduces the inter-frame mutual attention mechanism according to random Gaussian noise; Extracting motion information guidance and texture information guidance of the candidate key frame; The motion information guidance and the texture information guidance of the candidate key frame are used to guide the pre-trained video diffusion generation model that introduces the inter-frame mutual attention mechanism to denoise the motion video, so as to generate the first motion video.

5. The method according to claim 4, characterized in that The step of extracting motion information guidance and texture information guidance of the candidate key frame comprises: Get the foreground of the candidate keyframes; Using a preset edge detection algorithm to convert the foreground of the candidate key frame into an edge map; Inputting the edge map into a pre-trained ControlNet to determine the motion information guidance; encoding a first frame of the candidate key frames as an image embedding; Adding Gaussian noise to the image embedding to determine the image embedding with the Gaussian noise added; The image with added Gaussian noise is embedded into the input pre-trained IP-Adapter to determine the texture information guidance.

6. The method according to claim 1, characterized in that The motion prior of the pre-trained video diffusion generation model is used to inject motion information into the first motion video, and the texture prior of the pre-trained image diffusion generation model is used to inject texture information into the first motion video.

7. The method according to claim 6, characterized in that The step of optimizing the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video includes: Performing encoding processing on the first motion video to determine video embedding; adding Gaussian noise to the video embedding as an initial noise potential embedding for a denoising process of the pre-trained video diffusion model; Perform a single-step denoising operation on the noise latent embedding of each denoising time step to determine a new noise latent embedding; Performing a denoising operation for a first preset number of times on a new noise potential embedding at a denoising time step with motion information injection, and performing a weighted summation with the new noise potential embedding at a current denoising time step to determine a noise potential embedding injected with motion information, wherein the noise potential embedding injected with motion information is used for denoising at a next denoising time step or for texture information injection; For a denoising time step with texture information injection, determining a noise-free latent embedding according to a new noise latent embedding corresponding to the current denoising time step or a noise latent embedding injected with motion information corresponding to the current denoising time step; Re-adding Gaussian noise to the noise-free latent embedding to determine an image noise latent embedding; Performing a denoising operation on the image noise potential embedding for a second preset number of times to determine the image noise potential embedding after denoising; Determining, according to the denoised image noise potential embedding, an image noise-free potential embedding corresponding to the denoised image noise potential embedding; Re-adding Gaussian noise to the noise-free latent embedding of the image to determine a denoised video noise latent embedding; A weighted sum is performed on the new noise potential embedding corresponding to the current denoising time step or the noise potential embedding injected with motion information corresponding to the current denoising time step and the noise potential embedding of the denoised video to determine the noise potential embedding injected with texture information at the current denoising time step. The noise potential embedding injected with texture information is used for denoising at the next denoising time step until the denoising operations of all denoising time steps are completed to determine the target motion video.

8. An instance-level motion controllable video generation system, characterized in that: include: A first acquisition module, used for acquiring instance motion control conditions, wherein the instance motion control conditions include instance layout information and instance motion trajectory; An instance layout information augmentation module, configured to augment the instance layout information along the instance motion trajectory to determine augmented instance layout information; A candidate keyframe generation module, used for introducing the augmented instance layout information input into a pre-trained positioning image diffusion generation model of an inter-frame mutual attention mechanism to generate candidate keyframes; A first motion video generation module is used to introduce the candidate key frame input into a pre-trained video diffusion generation model of an inter-frame mutual attention mechanism to determine a first motion video; The target motion video generation module is used to optimize the first motion video according to the motion prior of the pre-trained video diffusion generation model and the detail prior of the pre-trained image diffusion generation model to determine the target motion video.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Training method, video generation method and device, electronic equipment and program product

    CN120706231A

  • Training method, video generation method and apparatus, electronic device, program product

    CN120706231B

  • Camera and illumination combined controllable 4D video generation method, device and equipment

    CN121567935A