Motion video generation method and device, electronic equipment and storage medium
By injecting LoRA components into the spatial and temporal modules of the base model and adopting an asynchronous training strategy, the motion consistency problem in motion video generation in the prior art is solved, and a high controllability and robust motion video generation is achieved.
Patent Information
- Application Number
- CN202510735996.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing motion video generation methods are difficult to maintain the motion consistency of subjects before and after editing in creative marketing scenarios, especially when modeling small amplitude motion and camera movements.
By injecting LoRA components into the space module and time module of the base model, and using asynchronous training strategies, the target motion video is generated, and the decoupling of the time and space modules is achieved, the space module is avoided learning timing parameters, and the time module is enhanced to learn pure motion information.
It improves the controllability and robustness of the target motion video, realizes the generation of motion video without trajectory guidance, and enhances the controllability and robustness of the motion video.
Smart Images

Figure CN120264038A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for generating motion videos. Background Art
[0002] In recent years, significant progress has been made in VDM (Video Diffusion Model) technology, which has promoted the rapid development of a large number of cutting-edge applications such as advertisement generation and virtual fitting. However, most current methods have limitations in motion control. Especially in the main body editing task of creative marketing scenarios, maintaining the motion consistency of the main body before and after editing is a very challenging task.
[0003] Most current motion controllable methods mainly guide the production process based on trajectories specified by users. However, such methods still have the following problems: trajectory information is difficult to model small-scale motions, such as head shaking and hand raising of a person; and trajectory information cannot describe camera motions, such as zooming in, zooming out, panning, and rotating. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method, device, electronic device and storage medium for generating motion videos to solve the problem of poor automatic generation effect of motion videos in the prior art.
[0005] In the first aspect of the embodiments of this application, a method for generating a motion video is provided, including: Obtain a reference video, where the reference video includes at least the motion information of a first moving object; Obtain a base model for generating a motion video, inject low-rank adaptation LoRA components into the spatial module and the temporal module of the base model respectively, and train the LoRA component of the spatial module and the LoRA component of the temporal module using an asynchronous training strategy to obtain a trained LoRA component of the spatial module and a trained LoRA component of the temporal module; Receive an instruction to generate a target motion video, where the instruction to generate a target motion video includes at least the motion information of a second moving object in the target motion video, and the motion information of the second moving object is the same as the motion information of the first moving object in the reference video; Determine the target weights of the LoRA component of the spatial module and the target weights of the LoRA component of the temporal module; Use the target weights of the LoRA component of the spatial module and the target weights of the LoRA component of the temporal module to fuse the trained LoRA component of the spatial module and the trained LoRA component of the temporal module with the base model to obtain a fused video generation model; Generate a target motion video based on the instruction to generate a target motion video and using the fused video generation model.
[0006] In the second aspect of the embodiments of the present application, a motion video generation device is provided, including: An acquisition module, configured to acquire a reference video, where the reference video includes at least the motion information of a first moving object; A training module, configured to acquire a base model for motion video generation, inject low-rank adaptation (LoRA) components into the spatial module and the temporal module of the base model respectively, and train the LoRA component of the spatial module and the LoRA component of the temporal module using an asynchronous training strategy to obtain a trained LoRA component of the spatial module and a trained LoRA component of the temporal module; A receiving module, configured to receive an instruction for generating a target motion video, where the instruction for generating a target motion video includes at least the motion information of a second moving object in the target motion video, and the motion information of the second moving object is the same as the motion information of the first moving object in the reference video; A determination module, configured to determine the target weights of the LoRA component of the spatial module and the target weights of the LoRA component of the temporal module; A fusion module, configured to fuse the trained LoRA component of the spatial module and the trained LoRA component of the temporal module with the base model using the target weights of the LoRA component of the spatial module and the target weights of the LoRA component of the temporal module to obtain a fused video generation model; A generation module, configured to generate a target motion video based on the instruction for generating a target motion video and using the fused video generation model.
[0007] In the third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the above method are implemented.
[0008] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0009] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: By obtaining a reference video and a base model generated from a motion video, the embodiments of the present application inject LoRA components into the spatial module and the temporal module of the base model respectively, train the LoRA components using an asynchronous training strategy, then receive an instruction to generate a target motion video, determine the target weights of the LoRA components, fuse the LoRA components with the base model using the target weights, and generate the target motion video based on the motion video instruction using the fused base model. This realizes the generation of a motion video using the motion information provided by the reference video without trajectory guidance. Moreover, by training and fine-tuning the parameters of the LoRA components of the spatial module and the temporal module in sequence using the asynchronous training strategy, it is possible to decouple the temporal and spatial modules, prevent the LoRA components of the spatial module from learning the temporal parameters of the reference video, and enable the LoRA components of the temporal module to learn pure motion information to the greatest extent, thereby improving the controllability and robustness of the target motion video. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 It is a flowchart showing a method for generating a motion video provided by an embodiment of the present application.
[0012] Figure 2 It is a flowchart showing a method for training the LoRA components of the spatial module and the temporal module using an asynchronous training strategy provided by an embodiment of the present application.
[0013] Figure 3 It is a flowchart showing a method for training the LoRA components of the spatial module using a reference video provided by an embodiment of the present application.
[0014] Figure 4 It is a flowchart showing a method for training the LoRA components of the temporal module using a reference video provided by an embodiment of the present application.
[0015] Figure 5 It is a flowchart showing a method for determining the target weights of the LoRA components of the spatial module provided by an embodiment of the present application.
[0016] Figure 6 It is a flowchart showing a method for determining the target weights of the LoRA components of the temporal module provided by an embodiment of the present application.
[0017] Figure 7 It is a schematic flowchart of another method for generating a motion video provided by an embodiment of the present application.
[0018] Figure 8 It is a schematic diagram of a motion video generation device provided by an embodiment of the present application.
[0019] Figure 9 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0020] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0021] A method and a device for generating a motion video according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0022] As mentioned above, most current motion control methods are based on the trajectory specified by the user to guide the production process. However, such methods still have the following problems: it is difficult to model small-amplitude motions, such as head shaking and hand raising, with trajectory information; and trajectory information cannot describe camera motions, such as zooming in, zooming out, panning, and rotating.
[0023] In view of this, an embodiment of the present application provides a method for generating a motion video. By obtaining a reference video and obtaining a base model for generating the motion video, LoRA components are respectively injected into the spatial module and the temporal module of the base model, and the LoRA components are trained using an asynchronous training strategy. Then, upon receiving an instruction to generate a target motion video, the target weights of the LoRA components are determined, and the LoRA components are fused with the base model using the target weights. Based on this motion video instruction, the fused base model is used to generate the target motion video. This realizes generating a motion video without trajectory guidance using the motion information provided by the reference video, and by sequentially training and fine-tuning the parameters of the LoRA components in the spatial module and the temporal module using the asynchronous training strategy, it is possible to decouple the temporal and spatial modules, avoid the LoRA components in the spatial module from learning the temporal parameters of the reference video, and enable the LoRA components in the temporal module to learn pure motion information to the greatest extent, thereby improving the controllability and robustness of the target motion video.
[0024] Figure 1 It is a schematic flowchart of a method for generating a motion video provided by an embodiment of the present application. As Figure 1 shown, the method includes the following steps: In step S101, a reference video is obtained.
[0025] Among them, the reference video includes at least the motion information of the first moving object.
[0026] In step S102, a base model generated from a motion video is obtained. Low-Rank Adaptation (LoRA) components are injected into the spatial module and the temporal module of the base model respectively, and the LoRA components of the spatial module and the LoRA components of the temporal module are trained using an asynchronous training strategy to obtain the trained LoRA components of the spatial module and the trained LoRA components of the temporal module.
[0027] In step S103, a command to generate a target motion video is received.
[0028] Among them, the command to generate a target motion video includes at least the motion information of the second moving object in the target motion video, and the motion information of the second moving object is consistent with the motion information of the first moving object in the reference video.
[0029] In step S104, the target weights of the LoRA components of the spatial module and the target weights of the LoRA components of the temporal module are determined.
[0030] In step S105, the trained LoRA components of the spatial module and the trained LoRA components of the temporal module are fused with the base model using the target weights of the LoRA components of the spatial module and the target weights of the LoRA components of the temporal module to obtain a fused video generation model.
[0031] In step S106, based on the command to generate a target motion video, a target motion video is generated using the fused video generation model.
[0032] In some embodiments of the present application, this method can be executed by a terminal device or a server. In some embodiments, this method can be used to generate motion videos in application scenarios such as advertising design with creative migration, virtual clothing try-on for videos, animation production, and video face swapping.
[0033] In some embodiments of the present application, a reference video can be obtained first, and the reference video includes at least the motion information of the first moving object. For example, the reference video can be a virtual clothing try-on video, which includes the motion information of the first object trying on one or several pieces of clothing. Another example is that the reference video can be an advertising video, which includes the motion information of the first displayed product.
[0034] In some embodiments of the present application, a base model for generating a motion video can also be obtained, and LoRA components can be injected into the spatial module and the temporal module of the base model respectively. The base model can be a conventional motion video generation model, and its model parameters are usually fixed. To improve the accuracy and flexibility of the base model in generating motion videos, LoRA components can be injected into it, and the parameters of the LoRA components can be trained in a bypass working mode. Then, the trained LoRA components are fused with the base model, and finally, the fused model is used to generate motion videos.
[0035] In some embodiments of the present application, to improve the model training efficiency of the LoRA components, so that after the trained LoRA components are fused with the base model, they can better generate motion videos with higher consistency and robustness, an asynchronous training strategy can be used to train the LoRA components of the spatial module and the LoRA components of the temporal module, so as to achieve the decoupling of the temporal and spatial modules, avoid the LoRA components of the spatial module learning the timing parameters of the reference video, and enable the LoRA components of the temporal module to learn pure motion information to the greatest extent.
[0036] In some embodiments, a reference video can be used to train the LoRA components of the spatial module and the LoRA components of the temporal module, and when training, the reference video and the text description of the reference video can be obtained, and the LoRA components are trained using the reference video and the text description of the reference video together. The text description of the reference video can be input by the user or generated according to the reference video, and there is no limitation here.
[0037] In some embodiments of the present application, a generation target motion video instruction input by the user can also be received. The generation target motion video instruction includes at least the motion information of the second moving object in the target motion video, and the motion information of the second moving object is consistent with the motion information of the first moving object in the reference video.
[0038] For example, if the reference video is the above virtual fitting video, the target motion video that the user hopes to generate can be a video of the second object trying on the same clothing as in the reference video. Another example is that if the reference video is the above advertisement video, the target motion video that the user hopes to generate can be a video of the second product moving in the same way as in the reference video.
[0039] In some embodiments of the present application, the target weights of the spatial module LoRA component and the temporal module LoRA component can be determined. Among them, the target weight of the spatial module LoRA component can be determined according to the generated target motion video instruction, and the target weight of the temporal module LoRA component can be dynamically adjusted during the process of generating the motion video. For the detailed methods of determining the target weights of the spatial module LoRA component and the temporal module LoRA component, please refer to the detailed description below, and will not be elaborated here.
[0040] After determining the target weights of the spatial module LoRA component and the temporal module LoRA component, the target weights of the spatial module LoRA component and the temporal module LoRA component can be used to fuse the trained spatial module LoRA component and the trained temporal module LoRA component with the base model to obtain a fused video generation model. Finally, based on the generated target motion video instruction and using the fused video generation model, the target motion video can be generated.
[0041] According to the technical solution provided by the embodiments of the present application, by obtaining a reference video and obtaining a base model for generating a motion video, injecting LoRA components into the spatial module and the temporal module of the base model respectively, training the LoRA components using an asynchronous training strategy, then receiving a generated target motion video instruction, determining the target weights of the LoRA components, using the target weights to fuse the LoRA components with the base model, and generating the target motion video using the fused base model based on the motion video instruction, it is possible to generate a motion video without trajectory guidance using the motion information provided by the reference video, and by training and fine-tuning the parameters of the spatial module LoRA component and the temporal module LoRA component in sequence using an asynchronous training strategy, it is possible to decouple the time and space modules, avoid the spatial module LoRA component from learning the temporal parameters of the reference video, and enable the temporal module LoRA component to learn pure motion information to the greatest extent, thereby improving the controllability and robustness of the target motion video.
[0042] In some embodiments of the present application, the base model can be a diffusion model. At this time, the spatial module of the base model can be the spatial self-attention and text-spatial cross-attention modules in the diffusion model, where the text is the text corresponding to the generated target motion video instruction, and this text can be input by the user, that is, the generated target motion video instruction is directly input by the user in the form of a text instruction. On the other hand, this text can also be generated by parsing other instructions such as voice instructions input by the user, and no limitation is made here.
[0043] At the same time, the temporal module of the base model can be the temporal self-attention module in the diffusion model.
[0044] Figure 2It is a schematic flowchart of a method for training a spatial module LoRA component and a temporal module LoRA component using an asynchronous training strategy provided by an embodiment of the present application. As Figure 2 shown, the method includes the following steps: In step S201, the spatial module LoRA component is trained using a reference video to obtain a trained spatial module LoRA component.
[0045] In step S202, in response to determining that the training of the spatial module LoRA component is completed, the temporal module LoRA component is trained using the reference video to obtain a trained temporal module LoRA component.
[0046] In some embodiments of the present application, when training the spatial module LoRA component and the temporal module LoRA component using an asynchronous training strategy, the spatial module LoRA component can be trained first using a reference video to obtain a trained spatial module LoRA component. Then, under the condition that it is determined that the training of the spatial module LoRA component is completed, the temporal module LoRA component is trained using the reference video to obtain a trained temporal module LoRA component.
[0047] That is to say, when training the spatial module LoRA component and the temporal module LoRA component using an asynchronous training strategy, the spatial module LoRA component can be trained first in each round of training to learn the appearance information of the reference video, and then the temporal module LoRA component. This asynchronous training strategy enables the temporal module LoRA component to further learn to generate the motion pattern of the reference video based on the content generated by the spatial module LoRA component without additional learning of appearance information.
[0048] Figure 3 It is a schematic flowchart of a method for training a spatial module LoRA component using a reference video provided by an embodiment of the present application. As Figure 3 shown, the method includes the following steps: In step S301, a random frame image in the reference video is obtained as a target training image.
[0049] In step S302, the spatial module LoRA component is trained using the target training image to obtain a trained spatial module LoRA component.
[0050] In some embodiments of the present application, when training the spatial module LoRA component, a random frame image in the reference video can be obtained as a target training image, and then the spatial module LoRA component is trained using the target training image, thereby obtaining a trained spatial module LoRA component.
[0051] Among them, the above training process can be repeated N times. Each time, a random frame image in the reference video is randomly selected as the target training image to train the spatial module LoRA component. N is a positive integer.
[0052] In this way, by using the diffusion model including the spatial module LoRA component to denoise a random frame in the reference video to train the spatial module LoRA component, it can enable the spatial module LoRA component to only learn the fixed appearance content in the reference video and avoid the spatial module LoRA component learning the temporal information in the reference video. At the same time, using random frame denoising can prevent the spatial module LoRA component from only learning the content of a certain frame and hindering the subsequent learning of temporal information.
[0053] Figure 4 It is a schematic flowchart of the method for training the temporal module LoRA component using the reference video provided by the embodiment of the present application. As Figure 4 shown, the method includes the following steps: In step S401, any two frame images in the reference video are obtained.
[0054] In step S402, the difference between any two frame images is predicted to obtain the cross-frame difference between any two frame images.
[0055] In step S403, using each frame image and the difference between it and the cross-frame difference with any other frame image as training data, the temporal module LoRA component is trained to obtain the trained temporal module LoRA component.
[0056] In some embodiments of the present application, when training the temporal module LoRA component, any two frame images in the reference video can be obtained first, and the difference between the any two frame images is predicted to obtain the cross-frame difference between the any two frame images. Then, using each frame image and the difference between it and the cross-frame difference with any other frame image as training data, the temporal module LoRA component is trained, and the trained temporal module LoRA component can be obtained.
[0057] That is to say, when training the temporal module LoRA component, all frames in the reference video can be trained by denoising so that the temporal module LoRA component can fully learn the motion information in the reference video. In addition, the appearance information in the reference video can be eliminated by predicting the cross-frame difference so that the temporal module LoRA component can learn pure motion information to the greatest extent.
[0058] For example, if the reference video includes n frames of images, the cross-frame differences between the i-th frame of image and the j-th frame of image can be predicted respectively. Then, subtract the cross-frame difference from the i-th frame of image as the j-th value of the i-th frame of image, and subtract the cross-frame difference from the j-th frame of image as the i-th value of the j-th frame of image. Finally, train the time module LoRA component according to all the values of the n frames of images. Wherein, n is a positive integer greater than 1, both i and j are positive integers greater than or equal to 1 and less than or equal to n, and i is different from j.
[0059] Figure 5 It is a schematic flowchart of the method for determining the target weight of the spatial module LoRA component provided by the embodiments of the present application. As Figure 5 shown, the method includes the following steps: In step S501, obtain the received instruction type for generating the target motion video.
[0060] In step S502, determine the editing type based on the instruction type.
[0061] In step S503, determine the weight of the spatial module LoRA component corresponding to the editing type as the target weight of the spatial module LoRA component.
[0062] In some embodiments of the present application, when determining the target weight of the spatial module LoRA component, the received instruction type for generating the target motion video can be obtained first, and the editing type can be determined based on this instruction type. Then, determine the weight of the spatial module LoRA component corresponding to the editing type as the target weight of the spatial module LoRA component.
[0063] That is to say, when generating a motion video, different types of the generated target motion video instructions may lead to differences in the size of the area for editing the video. For example, some types of instructions only need to change the video subject, and in this case, the video generation model needs to retain most of the areas outside the subject; another example is that some types of instructions also need to change the video style, and in this case, the video generation model needs to change the entire area of the video. Therefore, different target weights can be set for the spatial module LoRA component according to the instruction type.
[0064] In one example, the weights of the spatial module LoRA component corresponding to different editing types can be predefined in advance, and then the text corresponding to the generated target motion video instruction is classified, for example, using a pre-trained text classification model for classification, to obtain the instruction type of the generated target motion video instruction, and determine the editing type corresponding to this instruction type. Then, look up the weight of the spatial module LoRA component corresponding to this editing type, and the target weight of the corresponding spatial module LoRA component can be determined.
[0065] Figure 6It is a schematic flowchart of a method for determining the target weight of the LoRA component of the time module provided by an embodiment of the present application. As Figure 6 shown, the method includes the following steps: In step S601, determine the initial weight of the LoRA component of the time module.
[0066] In step S602, based on generating a target motion video instruction, use the base model without fusing the weight of the LoRA component of the time module to generate a first motion video.
[0067] In step S603, based on generating a target motion video instruction, use the base model that has fused the initial weight of the LoRA component of the time module to generate a second motion video.
[0068] In step S604, in response to determining that the content consistency between the first motion video and the second motion video is less than or equal to a first preset threshold, increase the initial weight to obtain the updated weight of the LoRA component of the time module.
[0069] In step S605, in response to determining that the content consistency between the first motion video and the second motion video is greater than or equal to a second preset threshold, decrease the initial weight to obtain the updated weight of the LoRA component of the time module; the second preset threshold is greater than the first preset threshold.
[0070] In step S606, iteratively execute the operation of generating an updated second motion video based on generating a target motion video instruction and using the base model that has fused the updated initial weight of the LoRA component of the time module until the iteration termination condition is met.
[0071] Among them, the iteration termination condition includes that the content consistency between the first motion video and the updated second motion video is greater than the first preset threshold and less than the second preset threshold.
[0072] In step S607, determine that the updated weight of the LoRA component of the time module when the iteration termination condition is met is the target weight of the LoRA component of the time module.
[0073] In some embodiments of the present application, when determining the target weight of the LoRA component of the time module, the initial weight of the LoRA component of the time module can be determined first. Among them, the initial weight can be an empirical value or a fixed value, which is not limited here.
[0074] In some implementation manners, a first motion video can be generated based on generating a target motion video instruction and using the base model without fusing the weight of the LoRA component of the time module, and a second motion video can be generated based on generating a target motion video instruction and using the base model that has fused the initial weight of the LoRA component of the time module.
[0075] The content consistency between the first motion video and the second motion video can be compared. If the content consistency between the first motion video and the second motion video is less than or equal to the first preset threshold, it can be considered that the motion information of the reference video has not been transferred to the newly generated motion video. At this time, the initial weight can be increased to obtain the updated weight of the time module LoRA component.
[0076] On the other hand, if it is determined that the content consistency between the first motion video and the second motion video is greater than or equal to the second preset threshold, it can be considered that the difference between the main body of the newly generated motion video and the main body in the reference video is not enough. At this time, the initial weight is reduced to obtain the updated weight of the time module LoRA component.
[0077] After updating the weight of the time module LoRA component, the operation of generating the updated second motion video using the base model that incorporates the updated initial weight of the time module LoRA component based on the instruction to generate the target motion video can be iteratively executed until the content consistency between the first motion video and the updated second motion video is greater than the first preset threshold and less than the second preset threshold. Finally, the updated weight of the time module LoRA component at the end of the iteration can be used as the target weight of the time module LoRA component.
[0078] That is to say, the target weight of the time module LoRA component can be adaptively updated based on the content consistency of the generated video in different situations during the video generation process until the optimal value is reached. Since the weight of the time module LoRA component is too small to transfer the learned motion information to the newly generated video, and too large will limit the flexibility of the model to generate the main body, restricting the main body to be highly consistent with the reference video. Therefore, during the motion video generation process, the content consistency between the first motion video generated by the base model and the second motion video generated by the base model with the time module LoRA component added can be compared. When the consistency is low, it indicates that the motion has not been injected into the newly generated video, and at this time, a larger weight can be given to the time module LoRA component. If the consistency is too high, the weight of the time module LoRA component can be correspondingly reduced to support the model to flexibly generate a new main body.
[0079] The technical solution provided by the embodiments of the present application efficiently fine-tunes the time module and the space module of the base model by adding trainable LoRA modules respectively. The reference video is used as training data to train the space module LoRA component to memorize the video appearance, and train the time module LoRA component to memorize the motion pattern ability, so as to obtain a motion controllable video generation model without trajectories. During the parameter fusion process of the base model and the LoRA component, a LoRA component weight adaptive calculation method is also adopted, enabling the model to have both motion controllability and the ability to flexibly generate the main body.
[0080] Meanwhile, the technical solution provided by the embodiments of this application adopts a dual-granularity activation motion customization technology to efficiently fine-tune a small number of parameters in the LoRA component to "memorize" the motion pattern in the reference video and adaptively inject it into the generation process of the new video, thereby improving the controllability and robustness of the motion. The technical solution provided by the embodiments can effectively train a motion customization model to batch generate creative videos with controllable motion, controllable scenes, and different contents. The model only needs to fine-tune a small number of parameters, ensuring the efficiency of the technology.
[0081] Figure 7 It is a schematic flowchart of another motion video generation method provided by the embodiments of this application. As Figure 7 shown, the user can first input a reference video and a text description of the reference video, and the text description of the reference video can also be automatically parsed from the reference video. In one example, the user can directly upload the reference video for customization and customize the text description corresponding to the reference video. Alternatively, the text description of the reference video can also be automatically generated by calling a multimodal understanding large model.
[0082] Then, obtain a video base model and inject a time module LoRA component and a space module LoRA component into it. The LoRA components of the base model can be trained using the reference video to achieve the motion customization ability for the reference video.
[0083] On the other hand, the user can also input an editing instruction, which can be an editing instruction for the target motion video that the user wants to newly generate. The multimodal understanding large model can be called to optimize the editing instruction input by the user and, after being fed back to the user for confirmation, used as the input to the base model.
[0084] The target weight of the space module LoRA component can be obtained by calling a classification model according to the editing type of the editing instruction. The time module LoRA component and the space module LoRA component are trained using the user input video and text description. During the training process, the target motion video can also be generated according to the editing instruction, and the target weight of the time module LoRA component can be adaptively adjusted based on the content consistency of the target motion videos generated under different conditions until the trained space module LoRA component and time module LoRA component are obtained. That is, it can Finally, fuse the trained space module LoRA component and time module LoRA component with the base model, and then use the fused base model to generate the target motion video, thus realizing the automatic generation of a motion controllable video.
[0085] Adopting the technical solution provided by the embodiment of the present application can address the problem that trajectory guidance in motion video generation is insensitive to small-scale motions and camera motions, and provide an efficient training for a motion controllable model. Meanwhile, by adjusting the target weights of the time module LoRA component and the spatial module LoRA component, the model can have both a robust motion controllable ability and a flexible subject generation ability.
[0086] Any combination of the above all optional technical solutions can form an optional embodiment of the present application, which will not be elaborated one by one here.
[0087] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0088] Figure 8 It is a schematic diagram of a motion video generation device provided by an embodiment of the present application. As Figure 8 shown, the device includes: An acquisition module 801, configured to acquire a reference video, where the reference video at least includes motion information of a first moving object.
[0089] A training module 802, configured to acquire a base model for motion video generation, inject low-rank adaptation LoRA components into the spatial module and the time module of the base model respectively, and train the spatial module LoRA component and the time module LoRA component using an asynchronous training strategy to obtain a trained spatial module LoRA component and a trained time module LoRA component.
[0090] A receiving module 803, configured to receive a target motion video generation instruction, where the target motion video generation instruction at least includes motion information of a second moving object in the target motion video, and the motion information of the second moving object is the same as the motion information of the first moving object in the reference video.
[0091] A determination module 804, configured to determine the target weight of the spatial module LoRA component and the target weight of the time module LoRA component.
[0092] A fusion module 805, configured to fuse the trained spatial module LoRA component and the trained time module LoRA component with the base model using the target weight of the spatial module LoRA component and the target weight of the time module LoRA component to obtain a fused video generation model.
[0093] A generation module 806, configured to generate a target motion video based on the target motion video generation instruction and using the fused video generation model.
[0094] According to the technical solution provided by the embodiments of the present application, by obtaining a reference video and a base model generated from a motion video, injecting LoRA components into the spatial module and the temporal module of the base model respectively, training the LoRA components using an asynchronous training strategy, then receiving a command to generate a target motion video, determining the target weights of the LoRA components, fusing the LoRA components with the base model using the target weights, and generating the target motion video using the fused base model based on the motion video command, it is possible to generate a motion video using the motion information provided by the reference video without trajectory guidance. Moreover, by training and fine-tuning the parameters of the LoRA components of the spatial module and the temporal module in sequence using the asynchronous training strategy, it is possible to decouple the temporal and spatial modules, prevent the LoRA components of the spatial module from learning the temporal parameters of the reference video, and enable the LoRA components of the temporal module to learn pure motion information to the greatest extent, thereby improving the controllability and robustness of the target motion video.
[0095] In some embodiments, the base model is a diffusion model; the spatial module of the base model is the spatial self-attention and text-spatial cross-attention modules in the diffusion model, and the text is the text corresponding to the command to generate the target motion video; the temporal module of the base model is the temporal self-attention module in the diffusion model.
[0096] In some embodiments, training the LoRA components of the spatial module and the temporal module using an asynchronous training strategy includes: training the LoRA components of the spatial module using the reference video to obtain the trained LoRA components of the spatial module; in response to determining that the training of the LoRA components of the spatial module is completed, training the LoRA components of the temporal module using the reference video to obtain the trained LoRA components of the temporal module.
[0097] In some embodiments, training the LoRA components of the spatial module using the reference video includes: obtaining a random frame image in the reference video as the target training image; training the LoRA components of the spatial module using the target training image to obtain the trained LoRA components of the spatial module.
[0098] In some embodiments, training the LoRA components of the temporal module using the reference video includes: obtaining any two frame images in the reference video; predicting the difference between the two frame images to obtain the cross-frame difference between the two frame images; using each frame image and the difference between it and the cross-frame difference with any other frame image as training data to train the LoRA components of the temporal module to obtain the trained LoRA components of the temporal module.
[0099] In some embodiments, determining the target weight of the spatial module LoRA component includes: obtaining the type of the received instruction for generating the target motion video; determining the editing type based on the instruction type; and determining the weight of the spatial module LoRA component corresponding to the editing type as the target weight of the spatial module LoRA component.
[0100] In some embodiments, determining the target weight of the temporal module LoRA component includes: determining the initial weight of the temporal module LoRA component; generating a first motion video based on the instruction for generating the target motion video and using the base model without fusing the weight of the temporal module LoRA component; generating a second motion video based on the instruction for generating the target motion video and using the base model that has fused the initial weight of the temporal module LoRA component; in response to determining that the content consistency between the first motion video and the second motion video is less than or equal to a first preset threshold, increasing the initial weight to obtain the updated weight of the temporal module LoRA component; or in response to determining that the content consistency between the first motion video and the second motion video is greater than or equal to a second preset threshold, decreasing the initial weight to obtain the updated weight of the temporal module LoRA component; the second preset threshold is greater than the first preset threshold; iteratively performing the operation of generating an updated second motion video based on the instruction for generating the target motion video and using the base model that has fused the updated initial weight of the temporal module LoRA component until an iteration termination condition is met; the iteration termination condition includes that the content consistency between the first motion video and the updated second motion video is greater than the first preset threshold and less than the second preset threshold; determining the updated weight of the temporal module LoRA component when the iteration termination condition is met as the target weight of the temporal module LoRA component.
[0101] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0102] Figure 9 is a schematic diagram of an electronic device provided by an embodiment of the present application. As Figure 9 shown, the electronic device 9 in this embodiment includes: a processor 901, a memory 902, and a computer program 903 stored in the memory 902 and executable on the processor 901. When the processor 901 executes the computer program 903, the steps in the above method embodiments are implemented. Alternatively, when the processor 901 executes the computer program 903, the functions of each module / unit in the above device embodiments are implemented.
[0103] The electronic device 9 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 9 may include, but is not limited to, the processor 901 and the memory 902. Those skilled in the art can understand, Figure 9This is only an example of the electronic device 9, which does not constitute a limitation on the electronic device 9. It may include more or fewer components than those shown in the figure, or different components.
[0104] The processor 901 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0105] The memory 902 may be an internal storage unit of the electronic device 9. For example, the hard disk or memory of the electronic device 9. The memory 902 may also be an external storage device of the electronic device 9. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 9. The memory 902 may also include both an internal storage unit and an external storage device of the electronic device 9. The memory 902 is used to store computer programs and other programs and data required by the electronic device.
[0106] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0107] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0108] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating a sports video, characterized in that, Including: Obtain a reference video, where the reference video at least includes the motion information of a first moving object; Obtain a base model generated from a motion video, inject low-rank adaptation (LoRA) components into the spatial module and the temporal module of the base model respectively, and train the LoRA component of the spatial module and the LoRA component of the temporal module using an asynchronous training strategy to obtain a trained LoRA component of the spatial module and a trained LoRA component of the temporal module; Receive an instruction to generate a target motion video, where the instruction to generate a target motion video at least includes the motion information of a second moving object in the target motion video, and the motion information of the second moving object is consistent with the motion information of the first moving object in the reference video; Determine the target weights of the LoRA component of the spatial module and the LoRA component of the temporal module; Use the target weights of the LoRA component of the spatial module and the LoRA component of the temporal module to fuse the trained LoRA component of the spatial module and the trained LoRA component of the temporal module with the base model to obtain a fused video generation model; Generate the target motion video based on the instruction to generate a target motion video and using the fused video generation model.
2. The method according to claim 1, characterized in that, The base model is a diffusion model; The spatial module of the base model is the spatial self-attention and text-spatial cross-attention modules in the diffusion model, and the text is the text corresponding to the instruction to generate a target motion video; The temporal module of the base model is the temporal self-attention module in the diffusion model.
3. The method according to claim 2, characterized in that, Training the LoRA component of the spatial module and the LoRA component of the temporal module using an asynchronous training strategy includes: Training the LoRA component of the spatial module using the reference video to obtain a trained LoRA component of the spatial module; In response to determining that the training of the LoRA component of the spatial module is completed, training the LoRA component of the temporal module using the reference video to obtain a trained LoRA component of the temporal module.
4. The method according to claim 3, wherein Training the LoRA component of the spatial module using the reference video includes: Obtain a random frame image in the reference video as a target training image; Train the LoRA component of the spatial module using the target training image to obtain a trained LoRA component of the spatial module.
5. The method according to claim 3, wherein Training the LoRA component of the temporal module using the reference video includes: Obtain any two frame images in the reference video; Predict the difference between the any two frame images to obtain the cross-frame difference between the any two frame images; Use each frame image and the difference between it and the cross-frame difference with any other frame image as training data to train the LoRA component of the temporal module to obtain a trained LoRA component of the temporal module.
6. The method according to claim 1, wherein Determining the target weights of the LoRA component of the spatial module includes: Obtain the type of the instruction to receive and generate a target motion video; Determine the editing type based on the instruction type; Determine the weight of the LoRA component of the spatial module corresponding to the editing type as the target weight of the LoRA component of the spatial module.
7. The method according to claim 1, characterized in that Determining the target weights of the LoRA component of the temporal module includes: Determine the initial weights of the time module LoRA component; Based on the generated target motion video instruction, generate a first motion video using a base model that has not fused the weights of the time module LoRA component; Based on the generated target motion video instruction, generate a second motion video using a base model that has fused the initial weights of the time module LoRA component; In response to determining that the content consistency between the first motion video and the second motion video is less than or equal to a first preset threshold, increase the initial weights to obtain updated weights of the time module LoRA component; or In response to determining that the content consistency between the first motion video and the second motion video is greater than or equal to a second preset threshold, decrease the initial weights to obtain updated weights of the time module LoRA component; the second preset threshold is greater than the first preset threshold; Iteratively execute the operation of generating an updated second motion video based on the generated target motion video instruction and using a base model that has fused the updated initial weights of the time module LoRA component until an iteration termination condition is met; the iteration termination condition includes that the content consistency between the first motion video and the updated second motion video is greater than the first preset threshold and less than the second preset threshold; Determine that the updated weights of the time module LoRA component when the iteration termination condition is met are the target weights of the time module LoRA component.
8. A sports video generation device, characterized in that, Includes: An acquisition module configured to acquire a reference video, the reference video including at least the motion information of a first moving object; A training module configured to acquire a base model for motion video generation, inject low-rank adaptation LoRA components into the spatial module and the time module of the base model respectively, and train the spatial module LoRA component and the time module LoRA component using an asynchronous training strategy to obtain a trained spatial module LoRA component and a trained time module LoRA component; A receiving module configured to receive a generated target motion video instruction, the generated target motion video instruction including at least the motion information of a second moving object in the target motion video, and the motion information of the second moving object is consistent with the motion information of the first moving object in the reference video; A determination module configured to determine the target weights of the spatial module LoRA component and the time module LoRA component; A fusion module configured to fuse the trained spatial module LoRA component and the trained time module LoRA component with the base model using the target weights of the spatial module LoRA component and the target weights of the time module LoRA component to obtain a fused video generation model; A generation module configured to generate the target motion video based on the generated target motion video instruction and using the fused video generation model.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video generation method, device and equipment
CN117788656A
Model processing method, video processing method and model processing device
CN118314440A
Video generation model training method and device, electronic equipment and storage medium
CN118658032A
Text-based motion video generation method and device, storage medium and equipment
CN119450164A
Multi-diffusion model fused image and video customization method, system and device
CN119676532A