Video generation method and related device
By feature encoding of videos with reference images and actions, and using video hidden diffusion model for denoising processing, the problem of lack of diversity and flexibility in video generation in the prior art is solved, and high-quality and efficient video generation is achieved.
Patent Information
- Application Number
- PCT/CN2024/135036
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-12
AI Technical Summary
Existing video generation technology mainly relies on text guidance, resulting in the generated video lacking diversity and flexibility and low quality.
By encoding the video containing the reference image and the reference action, the feature representation is obtained, and the timing maintenance module in the trained video hidden diffusion model is used to inject these features into the hidden diffusion network for denoising processing, and finally the action video of the target object is generated.
The quality and diversity of generated videos are improved, the flexibility and efficiency of video generation are enhanced, and the generated videos are more in line with the characteristics of reference images and actions.
Smart Images

Figure CN2024135036_12062025_PF_FP_ABST
Abstract
Description
Video generation method and related equipment
[0001] This application claims priority to the Chinese invention patent application with application number 202311648969.5 filed on December 4, 2023 and titled “Video Generation Method and Related Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of computer vision technology, and in particular to a video generation method and related equipment. Background Art
[0003] Computer vision is a simulation of biological vision using computers and related equipment. Its primary task is to process captured images or videos to obtain three-dimensional information about the corresponding scene. Currently, the continuous development of artificial intelligence has injected new vitality into computer vision technology. Emerging video generation technologies, in particular, have provided content creators with new tools, making video content creation easier and more cost-effective. However, most current video generation technologies use text as a guide, resulting in a lack of diversity and flexibility in the generated videos, resulting in low video quality. Summary of the Invention
[0004] The video generation method described in the embodiment of the present disclosure includes: encoding an image containing a reference image to obtain a feature representation of the reference image; encoding a video containing a reference action to obtain a feature representation of the reference action; injecting the feature representation of the reference image and the feature representation of the reference action into a latent diffusion network in the video latent diffusion model through at least one timing preservation module in the trained video latent diffusion model; wherein the latent diffusion network includes multiple cross-attention and denoising units; a timing preservation module is connected in series between at least two of the cross-attention and denoising units; and the timing preservation module includes a time converter and a space converter; using the latent diffusion network to denoise the input noise sequence to obtain a denoised latent vector sequence; and decoding the latent vector sequence to obtain an action video of the target object.
[0005] In an embodiment of the present disclosure, the encoding of the image containing the reference image to obtain the feature representation of the reference image includes: performing semantic encoding and / or pixel encoding on the image containing the reference image to obtain the feature representation of the reference image.
[0006] In an embodiment of the present disclosure, encoding the video containing the reference action to obtain the feature representation of the reference action includes: semantically encoding the video containing the reference action to obtain the feature representation of the reference action.
[0007] In an embodiment of the present disclosure, the above-mentioned step of injecting the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network of the video latent diffusion model through at least one timing preservation module in the video latent diffusion model includes: inputting the feature representation of the reference image into the spatial converter of the timing preservation module to generate a spatial guidance condition corresponding to the reference image; and inputting the feature representation of the reference action into the temporal converter of the timing preservation module to generate a temporal guidance condition corresponding to the reference action.
[0008] In an embodiment of the present disclosure, the above-mentioned denoising processing of the input noise sequence using the hidden diffusion network includes: using the time guidance condition and the spatial guidance condition to perform cross-attention processing on the input noise sequence to obtain predicted noise; and using the predicted noise to denoise the input noise sequence to obtain a denoised sequence.
[0009] In an embodiment of the present disclosure, the noise sequence is a Gaussian noise sequence; or, the noise sequence is a noise sequence obtained by performing noise processing on an image containing a reference image.
[0010] In an embodiment of the present disclosure, the above method may further include: extracting the last predetermined number N image frames of the action video of the target object to replace the first N image frames of the video containing the reference action; wherein N is a positive integer; executing the video generation method again to obtain the next segment of the action video of the target object; and splicing the action video of the target object with the next segment of the action video of the target object to obtain the spliced action video of the target object.
[0011] In an embodiment of the present disclosure, the above-mentioned splicing of the action video of the target object with the action video of the next target object includes: removing the first N image frames of the action video of the next target object; and directly splicing the action video of the target object with the action video of the next target object from which the first N image frames are removed to obtain the spliced action video of the target object.
[0012] In an embodiment of the present disclosure, the above method may further include: performing smoothing processing on the spliced action video of the target object.
[0013] In an embodiment of the present disclosure, the time converter and the space converter are connected in series.
[0014] In an embodiment of the present disclosure, a timing preservation module is connected in series between any two of the cross attention and denoising units.
[0015] Based on the above-mentioned video generation method, an embodiment of the present disclosure also provides a video generation device, including: an image encoding module, used to encode an image containing a reference image to obtain a feature representation of the reference image; a video encoding module, used to encode a video containing a reference action to obtain a feature representation of the reference action; a video latent diffusion model, including a latent diffusion network and at least one timing preservation module; wherein the latent diffusion network includes multiple cross-attention and denoising units, used to denoise the input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross-attention and denoising units; the timing preservation module includes a time converter and a space converter, used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network; and a decoder, used to decode the above-mentioned latent vector sequence to obtain an action video of the target object.
[0016] In an embodiment of the present disclosure, the time converter and the space converter are connected in series.
[0017] In an embodiment of the present disclosure, a timing preservation module is connected in series between any two of the cross attention and denoising units.
[0018] In an embodiment of the present disclosure, the spatial converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; and the temporal converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto.
[0019] In an embodiment of the present disclosure, the space converter is implemented by a converter; and / or the time converter is implemented by a converter.
[0020] In an embodiment of the present disclosure, the image encoder includes: a semantic encoder and / or a pixel encoder.
[0021] In an embodiment of the present disclosure, the above-mentioned video generation device further includes: a video splicing module, used to extract the last predetermined number N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; and splice the action video of the target object with the next segment of the action video of the target object to obtain a spliced action video of the target object; wherein N is a positive integer.
[0022] In an embodiment of the present disclosure, the video generating apparatus further comprises: a smoothing module for smoothing the spliced action video of the target object.
[0023] In addition, an embodiment of the present disclosure further provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned video generation method when executing the program.
[0024] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned video generation method.
[0025] An embodiment of the present disclosure further provides a computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the above-mentioned video generation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] FIG1 shows the implementation process of the video generation method according to an embodiment of the present disclosure;
[0028] FIG2 shows the structure of the video latent diffusion model according to an embodiment of the present disclosure;
[0029] FIG3 shows the implementation process of the method for denoising a noise sequence using a hidden diffusion network according to an embodiment of the present disclosure;
[0030] FIG4 shows the structure of the video generation model according to an embodiment of the present disclosure;
[0031] FIG5 shows the internal structure of a video generating device according to some embodiments of the present disclosure;
[0032] FIG6 shows a more specific schematic diagram of the hardware structure of an electronic device according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0034] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.
[0035] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0036] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0037] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0038] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0039] As mentioned earlier, although the development of artificial intelligence technology has made video content creation easier and more cost-effective, most current video generation technologies are based on text as a guide, resulting in a lack of diversity and flexibility in the generated videos, and low quality.
[0040] In order to solve the above problems, an embodiment of the present disclosure provides a video generation method, which can use an image containing a reference image and a video containing a reference action as guiding conditions to generate an action video of a target object that conforms to the above reference image and the above reference action, so as to improve the quality and efficiency of the generated video.
[0041] FIG1 shows the implementation process of the video generation method according to an embodiment of the present disclosure. As shown in FIG1 , the method may include the following steps 110 , 120 , 130 , 140 , and 150 .
[0042] In step 110, the image containing the reference image is encoded to obtain a feature representation of the reference image.
[0043] In step 120 , the video containing the reference action is encoded to obtain a feature representation of the reference action.
[0044] It should be noted that the embodiment of the present disclosure does not limit the execution order of the above steps 110 and 120, that is, the above steps 110 and 120 can be executed in parallel or in any order.
[0045] In step 130 , the feature representation of the reference image and the feature representation of the reference action are injected into the latent diffusion network in the latent diffusion model through at least one time-series preserving module in the trained latent diffusion model.
[0046] In an embodiment of the present disclosure, the implicit diffusion network may include multiple cross-attention and denoising units. A timing preservation module is connected in series between at least two of the cross-attention and denoising units. Furthermore, the timing preservation module may include a time converter and a space converter.
[0047] In step 140, the input noise sequence is denoised using the latent diffusion network to obtain a denoised latent vector sequence.
[0048] In step 150, the latent vector sequence is decoded to obtain an action video of the target object.
[0049] It can be seen that through the above-mentioned video generation method, an image containing a reference image and a video containing a reference action can be used as guiding conditions to generate an action video of the target object that conforms to the above-mentioned reference image and the above-mentioned reference action, thereby improving the quality of the generated video and the video generation efficiency.
[0050] The following will further describe in detail the various steps of the above video generation method with reference to the accompanying drawings and specific examples.
[0051] In an embodiment of the present disclosure, the reference image described in the above step 110 generally refers to the image of the target object that needs to be manipulated to move, and its main function is to guide the image of the target object in the generated video. In some examples of the present disclosure, the above target object may be a human body. Generally, the above reference image may generally include the face and body of the target object. In specific applications, the above reference image may be a real person image or a virtual person image wearing any clothing. For example, different images may be selected as the above reference image in different application scenarios. For example, in an e-commerce scenario, an image containing a model may be selected as the above image; in a social scenario, a meme image containing any target object may be selected as the above image; and in a creative scenario, a material image containing the target object may be selected as the above image.
[0052] In embodiments of the present disclosure, the encoding described in step 110 may include semantic encoding and / or pixel encoding. In some embodiments, the semantic encoding focuses on global information of the reference image, such as abstract information about the reference image's skeleton, skin color, posture, and body shape, while the pixel encoding focuses on local information of the reference image, such as detailed information about the reference image's clothing color and texture. After the semantic encoding and / or pixel encoding of the image, a feature representation of the reference image can be obtained.
[0053] In some specific examples, the above semantic coding and / or pixel coding can be implemented by a transformer or other existing image encoders. It should be noted that the embodiments of the present disclosure do not limit the specific implementation of semantic coding and pixel coding.
[0054] In some embodiments of the present disclosure, the encoding described in step 110 above may include both semantic encoding and pixel encoding. In this case, the above-mentioned image containing the reference image can be input into the semantic encoder and the pixel encoder respectively, and the semantic encoding result output by the above-mentioned semantic encoder - the semantic feature representation of the reference image and the pixel encoding result output by the above-mentioned pixel encoder - the pixel feature representation of the reference image are further fused to obtain the feature representation of the above-mentioned reference image. For example, the semantic feature representation of the above-mentioned reference image and the pixel feature representation of the above-mentioned reference image can be directly superimposed to obtain the feature representation of the above-mentioned reference image. In the above example, through the combination of multiple encoders at different levels, richer and more accurate reference image features can be extracted from the image, so that the video latent diffusion model can be more accurately and effectively guided to generate the action video of the target object, so as to achieve better results in the synthesis of the action of the target object.
[0055] In the embodiment of the present disclosure, the reference motion described in step 120 above represents a series of motions of the target object, for example, a video of a human body movement process. The reference motion is mainly used to guide the motion of the target object in the generated motion video of the target object.
[0056] It should be noted that, in the embodiment of the present disclosure, the object corresponding to the reference action in the video and the object corresponding to the reference image in the image described in step 110 may belong to the same object or different objects.
[0057] In embodiments of the present disclosure, the encoding described in step 120 above may generally refer to semantic encoding, which is used to extract global information such as the posture sequence of the target object from the video, for example, a sequence consisting of information such as the position and posture of each key point of the human body. In some specific embodiments, the above encoding can be implemented using a Transformer. It should be noted that the embodiments of the present disclosure do not limit the specific implementation method used for the encoding in step 120 above.
[0058] In steps 130 and 140, the video latent diffusion model can be a trained deep learning network model for generating a target object's motion video using a reference image feature representation and a reference motion feature representation as guidance. In some embodiments, the generated target object's motion video should generally conform to the reference image in the image and the reference motion in the video.
[0059] In embodiments of the present disclosure, the video latent diffusion model can utilize the existing latent diffusion model (LDM) structure as its underlying backbone network. To better incorporate the feature representations of the reference image and action into the LDM as spatial and temporal guidance conditions, embodiments of the present disclosure further incorporate at least one timing preservation module into the LDM to complete the spatiotemporal modeling of these spatial and temporal guidance conditions.
[0060] FIG2 shows an exemplary structure of the video latent diffusion model described in an embodiment of the present disclosure. As shown in FIG2 , in an embodiment of the present disclosure, the video latent diffusion model may include: a latent diffusion network 210 and at least one timing preservation module 220. In some embodiments, the latent diffusion network 210 is used to denoise the input noise sequence (which can usually be considered as a noise latent vector sequence) to obtain a denoised latent vector sequence. The at least one timing preservation module 220 is used to generate spatial guidance conditions based on the feature representation of the reference image and to generate temporal guidance conditions based on the feature representation of the reference action, and inject them into the latent diffusion network 210 to complete the spatiotemporal modeling of the spatial guidance conditions and the temporal guidance conditions, thereby guiding the latent diffusion network 210 to generate an action video of the target object that conforms to the reference image and the reference action.
[0061] Specifically, as shown in FIG2 , the latent diffusion network 210 may include a plurality of cross-attention and denoising units connected in series, each corresponding to a different time step. The cross-attention and denoising unit is used to perform cross-attention processing on the input noise sequence within its corresponding time step to obtain predicted noise, and then perform denoising processing on the input noise sequence based on the predicted noise. It can be understood that after being processed by multiple levels of cross-attention and denoising units, the latent diffusion network 210 can output a denoised latent vector sequence based on the input noise sequence. Furthermore, after decoding the denoised latent vector sequence, the individual image frames of the action video of the target object can be obtained. And after splicing the individual image frames, the action video of the target object can be obtained.
[0062] In an embodiment of the present disclosure, in order to inject the above-mentioned reference image features and the above-mentioned reference action features into the implicit diffusion network 210 and complete the spatiotemporal modeling of the spatial guidance conditions and the temporal guidance conditions, a timing preservation module 220 can be added in series between at least two cross-attention and denoising units to generate spatial guidance conditions based on the feature representation of the above-mentioned reference image and temporal guidance conditions based on the feature representation of the above-mentioned reference action, and inject them into the cross-attention and denoising units in the above-mentioned implicit diffusion network 210, thereby fusing them with the noise sequence input thereto. For example, as shown in Figure 2, a timing preservation module 220 can be connected in series between every two cross-attention and denoising units in the implicit diffusion network 210. In practical applications, the above-mentioned timing preservation module 220 can also be connected in series between any number of cross-attention and denoising units.
[0063] In an embodiment of the present disclosure, as also shown in FIG2 , the timing maintenance module 220 may specifically include: a time converter and a space converter. In some embodiments, the time converter and the space converter may be connected in series between two cross-attention and denoising units. It should be noted that although the time converter is located before the space converter in FIG2 , the embodiment of the present disclosure does not limit the order of the two converters. In practical applications, the space converter may also be placed before the time converter.
[0064] It should be noted that the input of the temporal transformer is the feature representation of the reference action. The temporal transformer is used to convert the feature representation of the reference action into a temporal guidance condition to guide the subsequent cross-attention and denoising units in the temporal dimension to predict noise, thereby denoising the noise sequence. The input of the spatial transformer is the feature representation of the reference image. The spatial transformer is used to convert the feature representation of the reference image into a spatial guidance condition to guide the subsequent cross-attention and denoising units in the spatial dimension to predict noise, thereby denoising the noise sequence.
[0065] In a specific example of the present disclosure, both the temporal transformer and the spatial transformer can be implemented using a Transformer. That is, in the above example, the timing preservation module 220 can be a two-layer Transformer network. In this case, each Transformer layer in the timing preservation module 220 can determine whether it is a temporal transformer or a spatial transformer based on its own input.
[0066] It can be seen that by adding one or more two-layer Transformer networks in series within the latent diffusion network 210, spatiotemporal modeling can be completed within the latent diffusion network 210, without adding any additional spatiotemporal fusion modules outside the latent diffusion network to inject time or space guidance conditions into the latent diffusion network. This processing method can effectively avoid the problem of needing to further add spatiotemporal fusion modules due to the increase in the number of guidance conditions, which in turn leads to an increase in the amount of spatiotemporal fusion calculations. In addition, the computing capacity of the at least one timing preservation module 220 added to the above-mentioned latent diffusion network 210 is very large, and by connecting the timing preservation module 220 in series between the cross-attention and denoising units of the latent diffusion network 210, its computing capacity can be shared in the latent diffusion network 210, thereby making the video latent diffusion model described in the present disclosure have a more powerful timing preservation capability.
[0067] Based on the structure of the video latent diffusion model shown in FIG2 , in an embodiment of the present disclosure, the step 130 of injecting the reference image features and the reference action features into the latent diffusion network through the timing preservation module of the video latent diffusion model may specifically include the following steps: inputting the feature representation of the reference image into the spatial converter of the timing preservation module to generate a spatial guidance condition corresponding to the reference image; and inputting the feature representation of the reference action into the temporal converter of the timing preservation module to generate a temporal guidance condition corresponding to the reference action.
[0068] Regarding step 140, FIG3 shows the implementation process of the method for denoising a noise sequence using the implicit diffusion network according to an embodiment of the present disclosure. As shown in FIG3, step 140 may specifically include steps 310 and 320.
[0069] In step 310, the input noise sequence is subjected to cross-attention processing using the above-mentioned temporal guidance condition and the above-mentioned spatial guidance condition to obtain predicted noise.
[0070] In the embodiment of the present disclosure, the noise sequence may be a randomly generated noise sequence. Generally, the randomly generated noise sequence conforms to a normal distribution, and thus, the noise sequence may also be a Gaussian noise sequence.
[0071] In other embodiments of the present disclosure, the above-mentioned noise sequence can also be a noise sequence obtained by performing noise processing on the above-mentioned image containing the reference image. It can be understood that using the noise sequence obtained by performing noise processing on the image containing the reference image as the input of the above-mentioned video latent diffusion model can further ensure that the output action video of the target object can better match the above-mentioned reference image, thereby further improving the quality and efficiency of the generated video. Specifically, the above-mentioned noise processing refers to continuously superimposing Gaussian noise on the above-mentioned image containing the reference image, for example, continuously superimposing Gaussian noise 50 times or even more times on the above-mentioned image containing the reference image, thereby obtaining a noise sequence similar to Gaussian noise.
[0072] In step 320, the input noise sequence is denoised using the predicted noise to obtain a denoised latent vector sequence.
[0073] It can be understood that, based on the structural characteristics of the aforementioned video latent diffusion model, the noise prediction and denoising processes described in steps 310 and 320 are performed separately within the cascaded cross-attention and denoising units corresponding to multiple time steps. As such, the final output of the aforementioned video latent diffusion model is a denoised latent vector sequence. Furthermore, because the noise prediction and denoising processes described in steps 310 and 320 are performed under the guidance of the aforementioned temporal and spatial guidance conditions, the target object's action video generated based on the denoised latent vector sequence will substantially conform to the aforementioned reference image and reference action.
[0074] In step 150, a decoder can be used to decode the denoised latent vector sequence to obtain individual image frames of the target object's motion video. The image frames of the target object's motion video are then spliced together to obtain the target object's motion video. It should be noted that the embodiments of this disclosure do not limit the specific implementation method used in step 150.
[0075] It can be seen that the above-mentioned video generation method given in the embodiment of the present disclosure can generate spatial and temporal guidance conditions based on an image containing a reference image and a video containing a reference action, and then generate an action video of the target object that conforms to the above-mentioned reference image and the above-mentioned reference action under the guidance of the above-mentioned guidance conditions, thereby improving the quality of the generated video and the efficiency of video generation.
[0076] In some embodiments of the present disclosure, steps 110 through 150 of the above-described video generation method can be considered to be implemented using a video generation model. FIG4 illustrates the structure of the video generation model described in an embodiment of the present disclosure. As shown in FIG4 , the video generation model may include an image encoding module 410, a video encoding module 420, a video latent diffusion model 430, and a decoder 440.
[0077] In some embodiments, the image encoding module 410 is configured to encode an image containing a reference image to obtain a feature representation of the reference image. In some embodiments, the image encoding module 410 may include a semantic encoder and / or a pixel encoder. As previously described, the semantic encoder and / or pixel encoder may be implemented using a Transformer or other image encoder.
[0078] As an example, the image encoding module 410 shown in Figure 4 includes a semantic encoder and a pixel encoder. In this case, the image encoding module 410 further includes a feature fusion module for fusing the semantic feature representation of the reference image output by the semantic encoder with the pixel feature representation of the reference image output by the pixel encoder. The feature fusion module is represented by ⊕ in Figure 4.
[0079] The video encoding module 420 is used to encode the video containing the reference action to obtain a feature representation of the reference action. In some embodiments, the video encoding module 420 can be implemented by a Transformer or other encoders.
[0080] The video latent diffusion model 430 may be structured as shown in FIG2 , comprising a latent diffusion network and at least one time-preserving module. In some embodiments, the latent diffusion network includes multiple cross-attention and denoising units for denoising an input noise sequence to generate a denoised latent vector sequence; a time-preserving module is connected in series between at least two of the cross-attention and denoising units; and the time-preserving module includes a temporal transformer and a spatial transformer for injecting the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network.
[0081] The decoder 440 is used to decode the latent vector sequence output by the video latent diffusion model 430 to obtain the action video of the target object.
[0082] In practical applications, the video generation model needs to be trained in advance. The training samples can usually be selected from human action videos and their corresponding human image images. It is understandable that the human action videos used as samples are usually selected high-quality human action videos. In some embodiments, the training loss function of the video generation model can be set to the mean square error (MSE) between the output of the video generation model and the true distribution. The training process aims to reduce the MSE loss to improve the effect of the video generation model.
[0083] Furthermore, since a large number of human motion videos in the training samples also include images of the human body from different perspectives during motion, the video generation module can also learn to image the human body from different perspectives through training. Thus, even though the video generation model only receives a single frame of a reference image as input, it can still accurately predict the reference image from different perspectives during motion.
[0084] Furthermore, since high-quality human action videos are relatively difficult to obtain, in practical applications, multiple sets of images of the human body from different perspectives can be used to assist in training the aforementioned video generation model. Using multiple sets of image data for auxiliary training ensures sufficient sample data during the training process, thereby improving the accuracy of the video generation model.
[0085] It's understandable that, because the backbone network of the video generation model is LDM, the target object's action videos generated by this model are typically short. For example, a video generated using LDM typically contains only about 30 frames and lasts about 2 seconds. Clearly, such a video length often fails to meet business needs.
[0086] In order to solve the problem that the action video of the target object generated by the above video generation model is too short, the embodiments of the present disclosure further provide a method for generating an action video of the target object with a longer length based on the above video generation method.
[0087] Specifically, in the above method, first, a first target object action video is obtained by executing the above video generation method. Then, the first N image frames of the reference action video are replaced with the last predetermined number N image frames of the generated first target object action video, where N is a positive integer. After the above preparation, the video generation method shown in Figure 1 is executed again to generate a new target object action video, such as a second target object action video. Specifically, when executing step 140 of the above video generation method, for the first predetermined number N image frames of the target object action video to be generated, the video latent diffusion model can determine the true noise of the N image frames based on the input video and the noise sequence, and then denoise the noise sequence based on the true noise of the N image frames. For the remaining image frames of the target object action video to be generated, the video latent diffusion model will continue to share the true noise of the first N image frames during the inference process, thereby achieving the purpose of transferring the content of the previously generated target object action video to the subsequently generated target object action video, thereby enhancing the temporal continuity of the generated video. After generating the second segment of the target object's action video, the first segment of the target object's action video and the second segment of the target object's action video can be spliced together to obtain a spliced target object's action video. It can be understood that during the generation of the target object's action video, N image frames are repeated in the two target object's action videos. Therefore, when performing video splicing, the repeated image frames should be removed before splicing. For example, the first N image frames of the second segment of the target object's action video can be removed first; then, the first segment of the target object's action video and the second segment of the target object's action video with the first N image frames removed can be directly spliced together to obtain the spliced target object's action video.
[0088] Next, the above steps can be repeated, using the last predetermined number of N frames of the second target object motion video to replace the first N frames of the reference target object motion video. The actual noise level of these N frames is determined based on the video and the noise sequence. The above video generation method is then repeated to generate a third target object motion video. This third target object motion video is then spliced together with the previous video. Repeating the above steps in this manner, after splicing together M target object motion videos, a longer target object motion video can be obtained.
[0089] In some embodiments of the present disclosure, the predetermined number N can generally be set according to the number of image frames contained in the action video of the target object generated by the video generation model, for example, it can be set to half the number of image frames contained in the action video of the target object generated by the video generation model. In a specific example, assuming that the action video of the target object generated by the video generation model contains 32 image frames, the predetermined number N can be set to 16. It can be understood that the setting of the predetermined number N is only an example, and the predetermined number N can also be set according to other methods, and the embodiments of the present disclosure are not limited to this.
[0090] As can be seen from the above method, after generating the first segment of the target object's action video, the method does not directly generate further segments of the target object's action video. Instead, it uses the second half of the previous segment of the target object's action video as the first half of the next segment of the target object's action video. By incorporating the second half of the previous segment of the target object's action video into the inference process of the video latent diffusion model, the temporal sequence preservation module of the video latent diffusion model transfers the content of the previous segment of the target object's action video to the subsequently generated target object's action video, thereby improving the temporal continuity of the generated target object's action video.
[0091] Corresponding to the above-mentioned video generation method, embodiments of the present disclosure further disclose a video generation device. FIG5 shows the internal structure of the video generation device described in some embodiments of the present disclosure. As shown in FIG5, the above-mentioned video generation device may include the following modules.
[0092] The image encoding module 510 is used to encode an image containing a reference image to obtain a feature representation of the reference image; the video encoding module 520 is used to encode a video containing a reference action to obtain a feature representation of the reference action; the video latent diffusion model 530; wherein the above-mentioned video latent diffusion model includes a latent diffusion network and at least one timing preservation module. In some embodiments, the above-mentioned latent diffusion network includes multiple cross-attention and denoising units, which are used to denoise the input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross-attention and denoising units; the above-mentioned timing preservation module includes a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the above-mentioned latent diffusion network; the decoder 540 is used to decode the above-mentioned latent vector sequence to obtain an action video of the target object.
[0093] In an embodiment of the present disclosure, the spatial converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; and the temporal converter is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto. Specifically, the spatial converter can be implemented by a converter; and / or the temporal converter can be implemented by a converter. In an embodiment of the present disclosure, the temporal converter and the spatial converter are connected in series.
[0094] In an embodiment of the present disclosure, a timing preservation module is connected in series between any two of the cross attention and denoising units.
[0095] In an embodiment of the present disclosure, the above-mentioned image encoder includes: a semantic encoder and / or a pixel encoder.
[0096] In an embodiment of the present disclosure, the above-mentioned video generation device may further include: a video splicing module, used to extract the last predetermined number N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; and splice the action video of the target object with the next segment of the action video of the target object to obtain a spliced action video of the target object; wherein N is a positive integer.
[0097] In an embodiment of the present disclosure, the video generating apparatus may further include: a smoothing module configured to perform smoothing processing on the spliced action video of the target object.
[0098] The specific implementation of each of the above modules can be referenced to the aforementioned methods and accompanying drawings, and will not be repeated here. For ease of description, the above device is described by function as a separate module. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware. The device of the above embodiment is used to implement the corresponding video generation method of any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0099] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the video generation method described in any of the above embodiments is implemented.
[0100] 6 shows a schematic diagram of the hardware structure of a more specific electronic device provided in this embodiment. The device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. The processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are connected to each other within the device via the bus 2050.
[0101] The processor 2010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0102] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 2020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 220 and is called and executed by the processor 2010.
[0103] The input / output interface 2030 is used to connect input / output devices to enable information input and output. Input / output devices can be configured as components within the device or externally connected to the device to provide corresponding functions. Input devices may include microphones and various sensors, while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0104] The communication interface 2040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0105] The bus 2050 comprises a path for transmitting information between the various components of the device (eg, the processor 2010 , the memory 2020 , the input / output interface 2030 , and the communication interface 2040 ).
[0106] It should be noted that although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0107] The electronic device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0108] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the video generation method described in any of the above embodiments.
[0109] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0110] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0111] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0112] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0113] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0114] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A video generation method, comprising: Encoding the image containing the reference image to obtain a feature representation of the reference image; Encoding the video containing the reference action to obtain a feature representation of the reference action; The feature representation of the reference image and the feature representation of the reference action are injected into the latent diffusion network in the video latent diffusion model through at least one timing preservation module in the trained video latent diffusion model; wherein the latent diffusion network includes a plurality of cross attention and denoising units; a timing preservation module is connected in series between at least two of the cross attention and denoising units; and the timing preservation module includes a time converter and a space converter; De-noising the input noise sequence using the latent diffusion network to obtain a denoised latent vector sequence; and The latent vector sequence is decoded to obtain an action video of the target object.
2. The method according to claim 1, wherein: The encoding of the image containing the reference image to obtain the feature representation of the reference image includes: performing semantic encoding and / or pixel encoding on the image containing the reference image to obtain the feature representation of the reference image.
3. The method according to claim 1, wherein: The step of encoding the video containing the reference action to obtain the feature representation of the reference action includes: semantically encoding the video containing the reference action to obtain the feature representation of the reference action.
4. The method according to claim 1, wherein: The step of injecting the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network of the video latent diffusion model through at least one timing preservation module in the video latent diffusion model comprises: Inputting the feature representation of the reference image into the spatial converter of the timing holding module to generate a spatial guidance condition corresponding to the reference image; and The characteristic representation of the reference action is input into the time converter of the timing keeping module to generate a time guiding condition corresponding to the reference action.
5. The method according to claim 4, wherein: The denoising process of the input noise sequence using the implicit diffusion network includes: Using the temporal guidance condition and the spatial guidance condition to perform cross-attention processing on the input noise sequence to obtain predicted noise; and The predicted noise is used to perform denoising on the input noise sequence to obtain a denoised sequence.
6. The method according to claim 5, wherein: The noise sequence is a Gaussian noise sequence; or, the noise sequence is a noise sequence obtained by performing noise processing on an image containing a reference image.
7. The method according to claim 1, further comprising: Extracting the last predetermined number of N image frames of the action video of the target object to replace the first N image frames of the video containing the reference action; wherein N is a positive integer; Executing the video generation method again to obtain the next action video of the target object; and The action video of the target object is spliced with the action video of the next target object to obtain a spliced action video of the target object.
8. The method according to claim 7, wherein: The step of splicing the action video of the target object with the action video of the next target object comprises: Removing the first N image frames of the next target object's action video; and The action video of the target object is directly spliced with the next segment of the action video of the target object with the first N image frames removed to obtain the spliced action video of the target object.
9. The method according to claim 7, further comprising: The spliced action video of the target object is smoothed.
10. The method according to claim 1, wherein: The time converter and the space converter are connected in series.
11. The method according to claim 1, wherein: A timing preservation module is connected in series between any two of the cross attention and denoising units.
12. A video generating device, comprising: An image encoding module, used for encoding an image containing a reference image to obtain a feature representation of the reference image; A video encoding module, used for encoding a video containing a reference action to obtain a feature representation of the reference action; A video latent diffusion model, comprising a latent diffusion network and at least one timing preservation module; wherein the latent diffusion network comprises a plurality of cross attention and denoising units, which are used to denoise an input noise sequence to obtain a denoised latent vector sequence; a timing preservation module is connected in series between at least two of the cross attention and denoising units; the timing preservation module comprises a time converter and a space converter, which are used to inject the feature representation of the reference image and the feature representation of the reference action into the latent diffusion network; as well as The decoder is used to decode the latent vector sequence to obtain the action video of the target object.
13. The video generating device according to claim 12, wherein: The spatial transformer is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto; and the temporal transformer is used to inject the feature representation of the reference image into the cross-attention and denoising unit connected thereto.
14. The video generating device according to claim 13, wherein: The space transformer is implemented by a transformer; and / or The time converter is implemented by a converter.
15. The video generating device according to claim 12, wherein: The time converter and the space converter are connected in series.
16. The video generating device according to claim 12, wherein: A timing preservation module is connected in series between any two of the cross attention and denoising units.
17. The video generating device according to claim 12, wherein: The image encoder includes: a semantic encoder and / or a pixel encoder.
18. The video generating device according to claim 12, further comprising: A video splicing module is used to extract the last predetermined number of N image frames of the action video of the target object and replace the first N image frames of the video containing the reference action; receive the next segment of the action video of the target object output by the decoder; And the action video of the target object is spliced with the action video of the next target object to obtain the spliced action video of the target object; wherein N is a positive integer.
19. The video generating device according to claim 17, further comprising: The smoothing module is used to smooth the action video of the target object after the splicing.
20. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the video generating method according to any one of claims 1 to 11 is implemented.
21. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the video generation method according to any one of claims 1 to 11.
22. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the video generating method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video generation method and server
CN116233491A
Video generation method, and method and device for training video generation model
CN116863003A
Video generation method and device, electronic equipment, storage medium and program product
CN116975357A
Video editing method and device, electronic equipment and storage medium
CN116980541A
Generation of story videos corresponding to user input using generative models
US20230118966A1
Cited By
Video virtual fitting method and device based on diffusion model and program product
CN121330119A
Embryo future development state generation method based on space-time flow attention driven diffusion model
CN121746514A
An embryo future development state generation method based on a space-time flow attention driven diffusion model
CN121746514B
Action planning method and device based on visual language guidance and differentiation diffusion
CN122067159A
Video generation method, system and model
CN122073636A