Video generation method and related equipment

By using the feature representation of reference images and action videos to guide the video hidden diffusion model to generate action videos, the problem of lack of diversity and flexibility in video generation in the prior art is solved, and high-quality and efficient video generation is achieved.

CN120111318APending Publication Date: 2025-06-06BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311648969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video generation technology mainly relies on text guidance, resulting in the generated video lacking diversity and flexibility and low quality.

Method used

By using the image containing the reference image and the video of the reference action as a guide condition, combined with the timing maintenance module in the video implicit diffusion model, the action video of the target object is generated. The method includes encoding the reference image and action video, generating feature representations, and denoising the process through the hidden diffusion network, and finally decoding to generate action video.

Benefits of technology

The quality and efficiency of generated videos are improved, and the generated videos are more in line with the diversity and flexibility of reference images and actions, and the convenience and low cost of content creation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111318A_ABST
    Figure CN120111318A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method. The video generation method comprises the following steps: encoding an image containing a reference image to obtain a feature representation of the reference image; coding the video containing the reference action to obtain the feature representation of the reference action; injecting the feature representation of the reference image and the feature representation of the reference action into an implicit diffusion network in the video implicit diffusion model through at least one time sequence maintaining module in the trained video implicit diffusion model; wherein the implicit diffusion network comprises a plurality of cross attention and denoising units; a time sequence maintaining module is connected in series between the at least two cross attention and denoising units; the time sequence holding module comprises a time converter and a space converter; performing denoising processing on an input noise sequence by using the implicit diffusion network to obtain a denoised implicit vector sequence; and decoding the implicit vector sequence to obtain an action video of the target object. The invention further provides a video generation device, electronic equipment, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a video generation method and related equipment. Background Art

[0002] Computer vision is a simulation of biological vision using computers and related equipment. Its main task is to obtain three-dimensional information of the corresponding scene by processing the collected pictures or videos. At this stage, the continuous development of artificial intelligence technology has injected new vitality into computer vision technology. Among them, the emerging video generation technology provides many content creators with new tools, making video content creation easier and more cost-effective. However, most of the current video generation technologies use text as a guiding condition, so the generated videos lack diversity and flexibility, resulting in low video quality. Summary of the invention

[0003] In view of this, an embodiment of the present disclosure provides a video generation method, which can generate an action video of a target object using an image containing a reference image and a video containing a reference action as guiding conditions to improve the quality and efficiency of the generated video.

[0004] The video generation method described in the embodiment of the present disclosure includes: encoding an image containing a reference image to obtain a feature representation of the reference image; encoding a video containing a reference action to obtain a feature representation of the reference action; injecting the feature representation of the reference image and the feature representation of the reference action into a latent diffusion network in the video latent diffusion model through at least one timing preservation module in the trained video latent diffusion model; wherein the latent diffusion network includes multiple cross-attention and denoising units; a timing preservation module is connected in series between at least two of the cross-attention and denoising units; and the timing preservation module includes a time converter and a space converter; using the latent diffusion network to denoise the input noise sequence to obtain a denoised latent vector sequence; and decoding the latent vector sequence to obtain an action video of the target object.

[0005] In an embodiment of the present disclosure, the encoding of an image containing a reference image to obtain a feature representation of the reference image includes: performing semantic encoding and / or pixel encoding on the image containing the reference image to obtain a feature representation of the reference image.

[0006] In an embodiment of the present disclosure, encoding the video containing the reference action to obtain the feature representation of the reference action includes: semantically encoding the video containing the reference action to obtain the feature representation of the reference action.

[0007] In an embodiment of the present disclosure, the above-mentioned method of injecting the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network in the video latent diffusion model through at least one timing keeping module in the video latent diffusion model includes: inputting the feature representation of the reference image into the spatial converter of the timing keeping module to generate a spatial guiding condition corresponding to the reference image; and inputting the feature representation of the reference action into the temporal converter of the timing keeping module to generate a temporal guiding condition corresponding to the reference action.

[0008] In an embodiment of the present disclosure, the above-mentioned denoising processing of the input noise sequence using the hidden diffusion network includes: using the time-guided condition and the space-guided condition to perform cross-attention processing on the input noise sequence to obtain predicted noise; and using the predicted noise to denoise the input noise sequence to obtain a denoised sequence.

[0009] In an embodiment of the present disclosure, the noise sequence is a Gaussian noise sequence; or, the noise sequence is a noise sequence obtained by performing noise processing on an image containing a reference image.

[0010] In an embodiment of the present disclosure, the above method may further include: extracting the last predetermined number of N image frames of the action video of the target object to replace the first N image frames of the video containing the reference action; wherein N is a positive integer; executing the video generation method again to obtain the next action video of the target object; and splicing the action video of the target object with the next action video of the target object to obtain the spliced ​​action video of the target object.

[0011] In an embodiment of the present disclosure, the above-mentioned splicing of the action video of the target object with the action video of the next target object includes: removing the first N image frames of the action video of the next target object; and directly splicing the action video of the target object with the action video of the next target object from which the first N image frames are removed to obtain the spliced ​​action video of the target object.

[0012] In an embodiment of the present disclosure, the above method may further include: performing smoothing processing on the spliced ​​action video of the target object.

[0013] In an embodiment of the present disclosure, the time converter and the space converter are connected in series.

[0014] In an embodiment of the present disclosure, a timing keeping module is connected in series between any two of the cross attention and denoising units.

[0015] Based on the above video generation method, an embodiment of the present disclosure further provides a video generation device, including:

[0016] An image encoding module, used for encoding an image containing a reference image to obtain a feature representation of the reference image;

[0017] A video encoding module, used for encoding a video containing a reference action to obtain a feature representation of the reference action;

[0018] A video latent diffusion model, comprising a latent diffusion network and at least one timing preservation module; wherein the latent diffusion network comprises a plurality of cross attention and denoising units, which are used to denoise an input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross attention and denoising units; the timing preservation module comprises a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network; and

[0019] The decoder is used to decode the latent vector sequence to obtain the action video of the target object.

[0020] In an embodiment of the present disclosure, the time converter and the space converter are connected in series.

[0021] In an embodiment of the present disclosure, a timing keeping module is connected in series between any two of the cross attention and denoising units.

[0022] In an embodiment of the present disclosure, the spatial converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; and the temporal converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto.

[0023] In an embodiment of the present disclosure, the space converter is implemented by a converter; and / or the time converter is implemented by a converter.

[0024] In an embodiment of the present disclosure, the image encoder includes: a semantic encoder and / or a pixel encoder.

[0025] In an embodiment of the present disclosure, the video generating device further comprises:

[0026] A video splicing module is used to extract the last predetermined number of N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; and splice the action video of the target object with the next segment of the action video of the target object to obtain a spliced ​​action video of the target object; wherein N is a positive integer.

[0027] In an embodiment of the present disclosure, the video generating device further comprises: a smoothing module, which is used to perform smoothing processing on the spliced ​​action video of the target object.

[0028] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned video generation method when executing the program.

[0029] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned video generation method.

[0030] An embodiment of the present disclosure further provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer executes the above-mentioned video generating method.

[0031] The above-mentioned video generation method and related equipment can use an image containing a reference image and a video containing a reference action as guiding conditions to generate an action video of a target object that conforms to the above-mentioned reference image and the above-mentioned reference action, thereby improving the quality of the generated video and the efficiency of video generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0033] Figure 1 The implementation process of the video generation method described in the embodiment of the present disclosure is shown;

[0034] Figure 2 The structure of the video implicit diffusion model described in the embodiment of the present disclosure is shown;

[0035] Figure 3 The implementation process of the method for denoising a noise sequence using an implicit diffusion network according to an embodiment of the present disclosure is shown;

[0036] Figure 4 The structure of the video generation model described in the embodiment of the present disclosure is shown;

[0037] Figure 5 The internal structure of the video generating device according to some embodiments of the present disclosure is shown;

[0038] Figure 6A more specific schematic diagram of the hardware structure of an electronic device described in some embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0040] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0041] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0042] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0043] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0044] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0045] As mentioned earlier, although the development of artificial intelligence technology has made video content creation easier and more cost-effective, most of the current video generation technologies are based on text as a guiding condition, resulting in the generated videos lacking diversity and flexibility and having low quality.

[0046] In order to solve the above problems, an embodiment of the present disclosure provides a video generation method, which can use an image containing a reference image and a video containing a reference action as guiding conditions to generate an action video of a target object that conforms to the above reference image and the above reference action, so as to improve the quality and efficiency of the generated video.

[0047] Figure 1 The implementation process of the video generation method described in the embodiment of the present disclosure is shown. Figure 1 As shown, the above method may include the following steps:

[0048] In step 110, the image containing the reference image is encoded to obtain a feature representation of the reference image.

[0049] In step 120, the video containing the reference action is encoded to obtain a feature representation of the reference action.

[0050] It should be noted that the embodiment of the present disclosure does not limit the execution order of the above steps 110 and 120, that is, the above steps 110 and 120 can be executed in parallel or in any order.

[0051] In step 130, the feature representation of the reference image and the feature representation of the reference action are injected into the latent diffusion network in the latent diffusion model through at least one time-preserving module in the trained video latent diffusion model.

[0052] In an embodiment of the present disclosure, the implicit diffusion network may include multiple cross attention and denoising units. A timing keeping module is connected in series between at least two of the cross attention and denoising units. Moreover, the timing keeping module may include a time converter and a space converter.

[0053] In step 140, the input noise sequence is denoised using the latent diffusion network to obtain a denoised latent vector sequence.

[0054] In step 150, the latent vector sequence is decoded to obtain an action video of the target object.

[0055] It can be seen that, through the above-mentioned video generation method, an action video of a target object that conforms to the above-mentioned reference image and the above-mentioned reference action can be generated by using an image containing a reference image and a video containing a reference action as guiding conditions.

[0056] Each step of the above video generation method will be described in detail below with reference to the accompanying drawings and specific examples.

[0057] In an embodiment of the present disclosure, the reference image described in the above step 110 generally refers to the image of the target object whose movement needs to be manipulated, and its main function is to guide the image of the target object in the generated video. In some examples of the present disclosure, the above target object may be a human body. Generally, the above reference image may generally include the face and body of the target object. In a specific application, the above reference image may be a real person image or a virtual person image wearing any clothing. For example, different images may be selected as the above reference image in different application scenarios. For example, in an e-commerce scenario, an image containing a model may be selected as the above image; in a social scenario, a meme containing any target object may be selected as the above image; and in a creative scenario, a material image containing the target object may be selected as the above image.

[0058] In the embodiment of the present disclosure, the coding described in the above step 110 may include: semantic coding and / or pixel coding. The above semantic coding mainly focuses on the global information of the reference image, for example, the abstract information of the human body region such as the skeleton, skin color, posture, fatness and thinness of the reference image; the above pixel coding mainly focuses on the local information of the reference image, for example, the clothing color and texture of the reference image and other detail information. After the above semantic coding and / or pixel coding, the above image can obtain the feature representation of the above reference image.

[0059] In some specific examples, the above semantic coding and / or pixel coding can be implemented by a transformer or other existing image encoders. It should be noted that the embodiments of the present disclosure do not limit the specific implementation methods of semantic coding and pixel coding.

[0060] In some embodiments of the present disclosure, the encoding described in step 110 above may include both semantic encoding and pixel encoding. In this case, the image containing the reference image may be input into a semantic encoder and a pixel encoder respectively, and the semantic encoding result output by the semantic encoder - the semantic feature representation of the reference image and the pixel encoding result output by the pixel encoder - the pixel feature representation of the reference image may be further fused to obtain the feature representation of the reference image. For example, the semantic feature representation of the reference image may be directly superimposed with the pixel feature representation of the reference image to obtain the feature representation of the reference image. In the above example, through the combination of multiple encoders at different levels, richer and more accurate features of the reference image can be extracted from the image, so that the video latent diffusion model can be more accurately and effectively guided to generate the action video of the target object, so as to achieve better results in the synthesis of the action of the target object.

[0061] In the embodiment of the present disclosure, the reference action described in step 120 above represents a series of actions of the target object, for example, it can be a video containing a human body movement process. The reference action is mainly used to guide the action of the target object in the generated action video of the target object.

[0062] It should be noted that, in the embodiments of the present disclosure, the object corresponding to the reference action in the video and the object corresponding to the reference image in the image described in step 110 may belong to the same object or different objects.

[0063] In the embodiments of the present disclosure, the encoding described in the above step 120 may generally refer to semantic encoding, which is used to extract global information such as the posture sequence of the target object from the video, for example, a sequence composed of information such as the position and posture of each key point of the human body. In some specific embodiments, the above encoding can be implemented by Transformer. It should be noted that the embodiments of the present disclosure do not limit the specific implementation method adopted by the encoding in the above step 120.

[0064] In the above steps 130 and 140, the video latent diffusion model can be a deep learning network model that has been trained to generate an action video of the target object using the feature representation of the reference image and the feature representation of the reference action as a guide condition. The generated action video of the target object should generally conform to the reference image in the above image and conform to the reference action in the above video.

[0065] In the embodiment of the present disclosure, the video latent diffusion model can be based on the existing latent diffusion model (LDM) structure as the backbone network. In order to better inject the feature representation of the reference image and the feature representation of the reference action into the LDM as the spatial guidance condition and the temporal guidance condition, the embodiment of the present disclosure further adds at least one timing preservation module to the LDM to complete the spatiotemporal modeling of the spatial guidance condition and the temporal guidance condition.

[0066] Figure 2 FIG. 2 shows an exemplary structure of the video latent diffusion model described in the embodiment of the present disclosure. Figure 2As shown, in an embodiment of the present disclosure, the above-mentioned video latent diffusion model may include: a latent diffusion network 210 and at least one timing preservation module 220. Among them, the above-mentioned latent diffusion network 210 is used to denoise the input noise sequence (usually considered to be a noise latent vector sequence) to obtain a denoised latent vector sequence. The above-mentioned at least one timing preservation module 220 is used to generate spatial guidance conditions based on the feature representation of the above-mentioned reference image and generate temporal guidance conditions based on the feature representation of the above-mentioned reference action and inject them into the above-mentioned latent diffusion network 210 to complete the spatiotemporal modeling of spatial guidance conditions and temporal guidance conditions, thereby guiding the latent diffusion network 210 to generate an action video of the target object that conforms to the above-mentioned reference image and the above-mentioned reference action.

[0067] Specifically, Figure 2 As shown, the implicit diffusion network 210 may include a plurality of cross-attention and denoising units connected in series corresponding to different time steps. The cross-attention and denoising units are used to perform cross-attention processing on the input noise sequence in the corresponding time step to obtain the predicted noise, and then denoise the input noise sequence based on the predicted noise. It can be understood that after being processed by multiple levels of cross-attention and denoising units, the implicit diffusion network 210 can output a denoised latent vector sequence based on the input noise sequence. Furthermore, after decoding the denoised latent vector sequence, each image frame of the action video of the target object can be obtained. And after splicing the image frames, the action video of the target object can be obtained.

[0068] In the embodiment of the present disclosure, in order to inject the above-mentioned reference image features and the above-mentioned reference action features into the implicit diffusion network 210 and complete the spatiotemporal modeling of the spatial guidance conditions and the temporal guidance conditions, a timing preservation module 220 can be added in series between at least two cross-attention and denoising units to generate spatial guidance conditions based on the feature representation of the above-mentioned reference image and temporal guidance conditions based on the feature representation of the above-mentioned reference action, and inject them into the cross-attention and denoising units in the above-mentioned implicit diffusion network 210, so as to fuse them with the noise sequence input thereto. For example, Figure 2 As shown, a time sequence keeping module 220 can be connected in series between every two cross attention and denoising units of the implicit diffusion network 210. In practical applications, the time sequence keeping modules 220 can also be connected in series between any number of cross attention and denoising units.

[0069] In the embodiments of the present disclosure, Figure 2 As shown, the timing keeping module 220 may specifically include: a time converter and a space converter. The time converter and the space converter may be connected in series between the two cross attention and denoising units. Figure 2In the embodiment, the time converter is located before the space converter, and the embodiment of the present disclosure does not limit the order of the above two converters. In practical applications, the space converter can also be placed before the time converter.

[0070] It should be noted that the input of the above-mentioned time converter is the feature representation of the above-mentioned reference action, and the above-mentioned time converter is used to convert the feature representation of the above-mentioned reference action into a time-guided condition to guide the subsequent cross-attention and denoising unit to predict noise in the time dimension, thereby denoising the noise sequence. The input of the above-mentioned spatial converter is the feature representation of the above-mentioned reference image, which is used to convert the feature representation of the above-mentioned reference image into a spatial-guided condition to guide the subsequent cross-attention and denoising unit to predict noise in the spatial dimension, thereby denoising the noise sequence.

[0071] In a specific example of the present disclosure, both the time converter and the space converter can be implemented by Transformer, that is, in the above example, the timing keeping module 220 can be a two-layer Transformer network. At this time, each layer of the Transformer in the timing keeping module 220 can determine whether the Transformer is a time converter or a space converter according to its own input.

[0072] It can be seen that by adding one or more groups of two-layer Transformer networks in series inside the implicit diffusion network 210, the spatiotemporal modeling can be completed inside the implicit diffusion network 210, without adding any additional spatiotemporal fusion modules outside the implicit diffusion network to inject time or space guiding conditions into the implicit diffusion network. Such a processing method can effectively avoid the problem of increasing the amount of spatiotemporal fusion calculations due to the increase in the number of guiding conditions. In addition, the computing capacity of at least one timing preservation module 220 added inside the above-mentioned implicit diffusion network 210 is very large, and by connecting the timing preservation module 220 in series between the cross-attention and denoising units of the implicit diffusion network 210, its computing capacity can be shared in the implicit diffusion network 210, so that the video implicit diffusion model described in the present disclosure has a stronger timing preservation capability.

[0073] Based on the above Figure 2 The structure of the video latent diffusion model shown in the figure, in the embodiment of the present disclosure, the step 130 of injecting the reference image features and the reference action features into the latent diffusion network through the timing preservation module of the video latent diffusion model can specifically include the following steps:

[0074] Inputting the feature representation of the reference image into the spatial converter of the timing holding module to generate a spatial guidance condition corresponding to the reference image; and

[0075] The characteristic representation of the reference action is input into the time converter of the timing keeping module to generate a time guidance condition corresponding to the reference action.

[0076] For the above step 140, Figure 3 The implementation process of the method for denoising a noise sequence using the implicit diffusion network described in the embodiment of the present disclosure is shown. Figure 3 As shown, the above step 140 may specifically include:

[0077] In step 310, the input noise sequence is cross-attention processed using the above-mentioned time guidance condition and the above-mentioned space guidance condition to obtain the predicted noise.

[0078] In the embodiment of the present disclosure, the noise sequence may be a randomly generated noise sequence. Generally, the randomly generated noise sequence conforms to a normal distribution, so the noise sequence may also be a Gaussian noise sequence.

[0079] In other embodiments of the present disclosure, the above-mentioned noise sequence may also be a noise sequence obtained by performing noise processing on the above-mentioned image containing the reference image. It can be understood that using the noise sequence obtained by performing noise processing on the image containing the reference image as the input of the above-mentioned video latent diffusion model can further ensure that the output action video of the target object can better match the above-mentioned reference image, thereby further improving the quality and efficiency of the generated video. Specifically, the above-mentioned noise processing refers to continuously superimposing Gaussian noise on the above-mentioned image containing the reference image, for example, continuously superimposing Gaussian noise 50 times or more on the above-mentioned image containing the reference image, thereby obtaining a noise sequence similar to Gaussian noise.

[0080] In step 320, the input noise sequence is denoised using the predicted noise to obtain a denoised latent vector sequence.

[0081] It can be understood that based on the structural characteristics of the above-mentioned video latent diffusion model, the prediction noise and denoising processing described in the above-mentioned steps 310 and 320 are respectively performed in the cascaded cross-attention and denoising units corresponding to multiple time steps, so that the final output of the above-mentioned video latent diffusion model will be a denoised latent vector sequence. In addition, since the prediction noise and denoising processing described in the above-mentioned steps 310 and 320 are performed under the guidance of the above-mentioned time-guided conditions and space-guided conditions, the action video of the target object generated based on the denoised latent vector sequence will basically conform to the above-mentioned reference image and the above-mentioned reference action.

[0082] In the above step 150, the decoder can be used to decode the above denoised latent vector sequence to obtain each image frame of the action video of the target object; then, the image frames of the action video of the target object are further spliced ​​to obtain the action video of the target object. It should be noted that the embodiments of the present disclosure do not limit the specific implementation method used for decoding in the above step 150.

[0083] It can be seen that the above-mentioned video generation method given in the embodiment of the present disclosure can generate spatial and temporal guidance conditions based on an image containing a reference image and a video containing a reference action, and then generate an action video of the target object that conforms to the above-mentioned reference image and the above-mentioned reference action under the guidance of the above-mentioned guidance conditions, thereby improving the quality of the generated video and the efficiency of video generation.

[0084] In some embodiments of the present disclosure, steps 110 to 150 in the above-mentioned video generation method can be considered to be implemented by a video generation model. Figure 4 The structure of the video generation model described in the embodiment of the present disclosure is shown. Figure 4 As shown, the above-mentioned video generation model may include: an image encoding module 410 , a video encoding module 420 , a video latent diffusion model 430 and a decoder 440 .

[0085] The image encoding module 410 is used to encode the image containing the reference image to obtain a feature representation of the reference image. In some embodiments, the image encoding module 410 may include: a semantic encoder and / or a pixel encoder. As mentioned above, the semantic encoder and / or the pixel encoder may be implemented by a Transformer or other image encoders.

[0086] As an example, Figure 4 The image encoding module 410 shown in FIG. 4 includes a semantic encoder and a pixel encoder. In this case, the image encoding module 410 will further include a feature fusion module for fusing the semantic feature representation of the reference image output by the semantic encoder and the pixel feature representation of the reference image output by the pixel encoder. The feature fusion module Figure 4 It is represented by ⊕.

[0087] The video encoding module 420 is used to encode the video containing the reference action to obtain the feature representation of the reference action. In some embodiments, the video encoding module 420 can be implemented by Transformer or other encoders.

[0088] The video implicit diffusion model 430 may be as described above. Figure 2The structure shown includes a latent diffusion network and at least one timing preservation module. The latent diffusion network includes a plurality of cross attention and denoising units, which are used to denoise the input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross attention and denoising units; the timing preservation module includes a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network.

[0089] The decoder 440 is used to decode the latent vector sequence output by the video latent diffusion model 430 to obtain the action video of the target object.

[0090] In practical applications, the video generation model needs to be trained in advance. The training sample can usually be selected to include a human action video and its corresponding human image image. It can be understood that the human action video used as a sample is usually a selected high-quality human action video. In some embodiments, the training loss function of the video generation model can be set to the mean square error (MSE) between the output of the video generation model and the true distribution, and the training process is aimed at reducing the MSE loss to improve the effect of the video generation model.

[0091] Furthermore, since a large number of human action videos in the training samples also contain images of the human body at different perspectives during the motion process, the video generation module can also learn the images of the human body at different perspectives through the training process. In this way, although the video generation model only inputs a single frame of an image containing a reference image, the video generation model can still accurately predict the image of the reference image at different perspectives during the motion process.

[0092] In addition, since high-quality human action videos are relatively difficult to obtain, in practical applications, multiple sets of image data of the human body at different perspectives can also be used to assist in training the above video generation model. Using multiple sets of image data for auxiliary training can ensure that there are enough sample data in the training process, thereby improving the accuracy of the video generation model.

[0093] It can be understood that, because the backbone network of the above video generation model is LDM, the length of the action video of the target object generated by the above video generation model is usually short. For example, the video generated by LDM usually only contains about 30 frames of images and the time length is about 2 seconds. Obviously, such a video length usually cannot meet the business needs.

[0094] In order to solve the problem that the action video of the target object generated by the above video generation model is too short, the embodiments of the present disclosure at least further provide a method for generating an action video of the target object with a longer length based on the above video generation method.

[0095] Specifically, in the above method, first, the first target object action video is obtained by executing the above video generation method. Then, the first N image frames of the above reference action video are replaced with the last predetermined number of N image frames of the generated first target object action video, where N is a positive integer. After the above preparation, the above method is executed again. Figure 1 The video generation method shown generates a new segment of action video of the target object, such as the second segment of action video of the target object. Specifically, when executing step 140 of the above-mentioned video generation method, for the predetermined number of N image frames in the front of the action video of the target object to be generated, the above-mentioned video latent diffusion model can determine the real noise of the above-mentioned N image frames based on the input video and the noise sequence, and then denoise the above-mentioned noise sequence based on the real noise of the above-mentioned N image frames; and for the remaining image frames in the back of the action video of the target object to be generated, the above-mentioned video latent diffusion model will continue to share the real noise of the previous N image frames during the reasoning process, thereby achieving the purpose of transferring the content of the action video of the target object generated last time to the action video of the target object generated subsequently, so as to enhance the temporal continuity of the generated video. After the second segment of the action video of the target object is generated, the above-mentioned first segment of the action video of the target object can be spliced ​​with the above-mentioned second segment of the action video of the target object to obtain the spliced ​​action video of the target object. It can be understood that in the process of generating the action video of the target object, N image frames are repeated in the two action videos of the target object. Therefore, when splicing the videos, the repeated image frames should be removed before splicing. For example, the first N image frames of the action video of the second target object can be removed first; then, the action video of the first target object and the action video of the second target object with the first N image frames removed can be directly spliced ​​together to obtain the spliced ​​action video of the target object.

[0096] Next, the above operation can be repeated, that is, the last predetermined number of N image frames of the generated second segment of the target object's action video are used to replace the first N image frames of the above video containing the reference action of the target object, and the real noise of the above N image frames is determined based on the above video and the noise sequence, and then the above video generation method is executed again to generate a third segment of the target object's action video. Then, the third segment of the target object's action video is spliced ​​with the previous video. And so on, repeat the above operation, after splicing M segments of the target object's action video, a longer segment of the target object's action video can be obtained.

[0097] In some embodiments of the present disclosure, the above-mentioned predetermined number N can generally be set according to the number of image frames contained in the action video of the target object generated by the video generation model, for example, it is set to half the number of image frames contained in the action video of the target object generated by the video generation model. In a specific example, assuming that the action video of the target object generated by the video generation model contains 32 image frames, the above-mentioned predetermined number N can be set to 16. It can be understood that the setting of the above-mentioned predetermined number N is only an example, and the above-mentioned predetermined number N can also be set according to other methods, and the embodiments of the present disclosure are not limited to this.

[0098] It can be seen from the above method that after generating the first segment of the target object's action video, the target object's action video of other segments will not be directly generated, but the second half of the previous segment of the target object's action video will be selected as the first half of the next segment of the target object's action video. By adding the second half of the previous segment of the target object's action video to the inference process of the video latent diffusion model, the timing preservation module of the video latent diffusion model is used to transfer the content of the previous segment of the target object's action video to the subsequently generated target object's action video, thereby improving the generated timing continuity.

[0099] Corresponding to the above-mentioned video generating method, the embodiment of the present disclosure further discloses a video generating device. Figure 5 The internal structure of the video generation device described in some embodiments of the present disclosure is shown. Figure 5 As shown, the above video generation device may include the following modules:

[0100] An image encoding module 510, for encoding an image containing a reference image to obtain a feature representation of the reference image;

[0101] A video encoding module 520, configured to encode a video containing a reference action to obtain a feature representation of the reference action;

[0102] Video latent diffusion model 530; wherein the video latent diffusion model comprises a latent diffusion network and at least one timing preservation module. The latent diffusion network comprises a plurality of cross attention and denoising units, which are used to denoise the input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross attention and denoising units; the timing preservation module comprises a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network;

[0103] The decoder 540 is used to decode the latent vector sequence to obtain the action video of the target object.

[0104] In an embodiment of the present disclosure, the spatial converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; the temporal converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto. Specifically, the spatial converter can be implemented by a converter; and / or the temporal converter can be implemented by a converter. In an embodiment of the present disclosure, the temporal converter and the spatial converter are connected in series.

[0105] In an embodiment of the present disclosure, a timing keeping module is connected in series between any two of the cross attention and denoising units.

[0106] In an embodiment of the present disclosure, the above-mentioned image encoder includes: a semantic encoder and / or a pixel encoder.

[0107] In an embodiment of the present disclosure, the above-mentioned video generation device may further include: a video splicing module, used to extract the last predetermined number N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; and splice the action video of the target object with the next segment of the action video of the target object to obtain a spliced ​​action video of the target object; wherein N is a positive integer.

[0108] In an embodiment of the present disclosure, the video generating device may further include: a smoothing module, configured to perform smoothing processing on the spliced ​​action video of the target object.

[0109] The specific implementation of each of the above modules can refer to the above methods and drawings, and will not be repeated here. For the convenience of description, the above device is described by function and is divided into various modules and described separately. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or more software and / or hardware. The device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0110] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video generation method described in any of the above embodiments is implemented.

[0111] Figure 6A more specific hardware structure diagram of an electronic device provided in this embodiment is shown, and the device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. The processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are connected to each other in communication within the device through the bus 2050.

[0112] The processor 2010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0113] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 2020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.

[0114] The input / output interface 2030 is used to connect input / output devices to realize information input and output. The input / output devices can be configured in the device as components, or can be externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.

[0115] The communication interface 2040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0116] The bus 2050 includes a path that transmits information between the various components of the device (eg, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).

[0117] It should be noted that, although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0118] The electronic device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0119] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the video generation method described in any of the above embodiments.

[0120] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0121] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0122] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0123] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0124] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0125] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A video generation method, include: Encoding the image containing the reference image to obtain a feature representation of the reference image; Encoding the video containing the reference action to obtain a feature representation of the reference action; The feature representation of the reference image and the feature representation of the reference action are injected into the latent diffusion network in the video latent diffusion model through at least one timing preservation module in the trained video latent diffusion model; wherein the latent diffusion network includes a plurality of cross attention and denoising units; a timing preservation module is connected in series between at least two of the cross attention and denoising units; and the timing preservation module includes a time converter and a space converter; De-noising the input noise sequence using the latent diffusion network to obtain a denoised latent vector sequence; and The latent vector sequence is decoded to obtain an action video of the target object.

2. The method according to claim 1, in, The encoding of the image containing the reference image to obtain the feature representation of the reference image includes: performing semantic encoding and / or pixel encoding on the image containing the reference image to obtain the feature representation of the reference image.

3. The method according to claim 1, in, The step of encoding the video containing the reference action to obtain the feature representation of the reference action includes: semantically encoding the video containing the reference action to obtain the feature representation of the reference action.

4. The method according to claim 1, in, The step of injecting the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network of the video latent diffusion model through at least one timing preservation module in the video latent diffusion model comprises: Inputting the feature representation of the reference image into the spatial converter of the timing holding module to generate a spatial guidance condition corresponding to the reference image; and The characteristic representation of the reference action is input into the time converter of the timing keeping module to generate a time guiding condition corresponding to the reference action.

5. The method according to claim 4, in, The step of using the implicit diffusion network to denoise the input noise sequence includes: Using the temporal guidance condition and the spatial guidance condition to perform cross-attention processing on the input noise sequence to obtain predicted noise; and The predicted noise is used to perform denoising on the input noise sequence to obtain a denoised sequence.

6. The method according to claim 5, in, The noise sequence is a Gaussian noise sequence; or, the noise sequence is a noise sequence obtained by performing noise processing on an image containing a reference image.

7. The method according to claim 1, further comprising: include: Extracting the last predetermined number of N image frames of the action video of the target object to replace the first N image frames of the video containing the reference action; wherein N is a positive integer; Executing the video generation method again to obtain the next action video of the target object; and The action video of the target object is spliced ​​with the action video of the next target object to obtain a spliced ​​action video of the target object.

8. The method according to claim 7, in, The step of splicing the action video of the target object with the action video of the next target object comprises: Removing the first N image frames of the next target object's action video; and The action video of the target object is directly spliced ​​with the next segment of the action video of the target object with the first N image frames removed to obtain the spliced ​​action video of the target object.

9. The method according to claim 7, further comprising: include: The spliced ​​action video of the target object is smoothed.

10. The method according to claim 1, in, The time converter and the space converter are connected in series.

11. The method according to claim 1, in, A timing preservation module is connected in series between any two of the cross attention and denoising units.

12. A video generating device, include: An image encoding module, used for encoding an image containing a reference image to obtain a feature representation of the reference image; A video encoding module, used for encoding a video containing a reference action to obtain a feature representation of the reference action; A video latent diffusion model, comprising a latent diffusion network and at least one timing preservation module; wherein the latent diffusion network comprises a plurality of cross attention and denoising units, which are used to denoise an input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross attention and denoising units; the timing preservation module comprises a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network; as well as The decoder is used to decode the latent vector sequence to obtain the action video of the target object.

13. The video generating device according to claim 12, in, The spatial transformer is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; and the temporal transformer is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto.

14. The video generating device according to claim 13, in, The space transformer is implemented by a transformer; and / or The time converter is implemented by a converter.

15. The video generating device according to claim 12, in, The time converter and the space converter are connected in series.

16. The video generating device according to claim 12, in, A timing preservation module is connected in series between any two of the cross attention and denoising units.

17. The video generating device according to claim 12, in, The image encoder includes: a semantic encoder and / or a pixel encoder.

18. The video generating device according to claim 12, further comprising: include: A video splicing module is used to extract the last predetermined number of N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; And the action video of the target object is spliced ​​with the action video of the next target object to obtain the spliced ​​action video of the target object; wherein N is a positive integer.

19. The video generating device according to claim 17, further comprising: include: The smoothing module is used to smooth the action video of the target object after the splicing.

20. An electronic device, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video generation method according to any one of claims 1 to 11 is implemented.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the video generation method according to any one of claims 1 to 11.

22. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the video generating method according to any one of claims 1 to 11.

Citation Information

Cited By

  • System and method for efficient text-guided generation of high-resolution videos

    US12749234B2

  • System and method for efficient text-guided generation of high-resolution videos

    US20250111552A1