Video generation method and device, equipment and storage medium

Through the diffusion model, the reference facial marking points are encoded and the target facial marking points sequence is generated in combination with audio features, which solves the problems of poor video time stability and high calculation cost in traditional technology, and achieves high-quality and low-cost video generation.

CN119967259APending Publication Date: 2025-05-09BYTEDANCE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138190.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional facial generation technology has shortcomings in terms of stability and generation speed, resulting in poor stability of generated videos and high computational cost.

Method used

Through a video generation method based on the diffusion model, the reference facial marking points are encoded as feature representations using the trained diffusion model, and the target facial marking points sequence is generated in combination with audio features, and the target video is finally generated.

Benefits of technology

This method can generate continuous and stable sequences of facial marking points, improving the naturalness and expressiveness of the video, while reducing calculation costs and improving inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967259A_ABST
    Figure CN119967259A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and device, equipment and a storage medium. The method comprises the steps of determining at least one first mark point feature representation based on at least one group of reference face mark points corresponding to at least one reference image of a target object; determining a second mark point feature representation sequence by using a trained first diffusion model based on the audio feature representation corresponding to the target audio and the at least one first mark point feature representation; based on decoding of the second mark point feature representation sequence, determining a target face mark point sequence corresponding to the target audio; and generating a target video based on the at least one reference image, the at least one group of reference face mark points and the target face mark point sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to video generation methods, devices, apparatuses, computer-readable storage media, and computer-executable instruction products. Background Art

[0002] With the continuous development of speech-driven face generation technology (Talking Face Generation, TFG), this technology has shown broad potential in application scenarios such as virtual character generation, video conferencing, and intelligent assistants. However, traditional face generation technology still needs to be improved in terms of stability and generation speed. Summary of the invention

[0003] In a first aspect of the present disclosure, a video generation method is provided. The method includes: determining at least one first landmark feature representation based on at least one set of reference facial landmarks corresponding to at least one reference image of a target object; determining a second landmark feature representation sequence based on an audio feature representation corresponding to a target audio and at least one first landmark feature representation using a trained first diffusion model, the second landmark feature representation sequence including a plurality of second landmark feature representations corresponding to a plurality of audio frames in the target audio; determining a target facial landmark sequence corresponding to the target audio based on decoding of the second landmark feature representation sequence, the target facial landmark sequence including a plurality of sets of target facial landmarks corresponding to a plurality of audio frames; and generating a target video based on at least one reference image, at least one set of reference facial landmarks and a target facial landmark sequence, the target audio including a target object and corresponding to the target audio.

[0004] In a second aspect of the present disclosure, a device for video generation is provided. The device includes: a first determination module, configured to determine at least one first landmark feature representation based on at least one group of reference facial landmarks corresponding to at least one reference image of a target object; a second determination module, configured to determine a second landmark feature representation sequence based on an audio feature representation corresponding to a target audio and at least one first landmark feature representation, using a trained first diffusion model, the second landmark feature representation sequence includes a plurality of second landmark feature representations corresponding to a plurality of audio frames in the target audio; a third determination module, configured to determine a target facial landmark sequence corresponding to the target audio based on decoding of the second landmark feature representation sequence, the target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to a plurality of audio frames; and a generation module, configured to generate a target video based on at least one reference image, at least one group of reference facial landmarks and a target facial landmark sequence, the target audio including a target object and corresponding to the target audio.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer executable instruction product is provided, comprising computer executable instructions, wherein when the computer executable instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments according to the present disclosure may be implemented;

[0011] Figure 2 A flowchart showing a process of video generation according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing an example architecture for video generation according to some embodiments of the present disclosure;

[0013] FIG. 4A to FIG. 4C Schematic diagrams of example architectures of video generation according to other embodiments of the present disclosure are respectively shown;

[0014] Figure 5 A schematic diagram showing a training process of a first diffusion model according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A schematic diagram showing a training process of a second diffusion model according to some embodiments of the present disclosure is shown;

[0016] Figure 7A schematic structural block diagram showing an example apparatus for video generation according to some embodiments of the present disclosure; and

[0017] Figure 8 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0020] Herein, unless explicitly stated, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to software or hardware such as electronic devices, applications, servers or storage media that execute operations of the technical solution of the present disclosure based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information is sent to the user in a manner such as a pop-up window, in which the prompt information can be presented in text form. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that after the training is completed, the corresponding output can be generated for a given input. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multi-layer processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0027] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, and typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values ​​are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping of input to output) from the training data. The parameter values ​​of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the application stage, the model can be used to process the actual input based on the parameter values ​​obtained from the training to determine the corresponding output.

[0029] As mentioned above, speech-driven facial generation technology aims to use driving audio to animate the facial expressions and head movements of the target object to create realistic and lip-synced videos. However, it is not easy to achieve realistic facial animation. Early technologies mainly mapped audio signals to lip movements while maintaining the static appearance of the outer lip area. However, the motionless head posture and fixed line of sight limit the naturalness and authenticity of the generated results. Subsequently, related technologies have gradually expanded to the generation of facial expressions and head movements. The driving audio is used to generate facial motion representation, and the facial motion representation is used to guide the generation of head images to achieve the synchronization of facial expressions, head movements and lip movements. In traditional technologies, two-dimensional facial landmarks or three-dimensional deformable model coefficients are used as facial motion representations to guide the generation of head images. However, in these traditional technologies, on the one hand, the generated facial landmarks have high-frequency jitter, resulting in poor temporal stability of the generated video. On the other hand, the video generation model of the transmission has redundant structures, slow inference speed, and high computational cost.

[0030] In view of this, an embodiment of the present disclosure proposes an improved scheme for video generation. In this scheme, based on at least one reference image of a target object, at least one group of reference facial landmarks is determined. Based on the at least one group of reference facial landmarks, at least one first landmark feature representation is determined. Based on the audio feature representation corresponding to the target audio and the at least one first landmark feature representation, a second landmark feature representation sequence is determined using a trained first diffusion model. The second landmark feature representation sequence includes a plurality of second landmark feature representations corresponding to a plurality of audio frames in the target audio. Based on decoding of the second landmark feature representation sequence, a target facial landmark sequence corresponding to the target audio is determined, and the target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to a plurality of audio frames. Afterwards, based on at least one reference image, at least one group of reference facial landmarks and a target facial landmark sequence, a target video is generated, and the target audio includes a target object and corresponds to the target audio.

[0031] In the embodiment of the present disclosure, the reference facial landmarks are encoded into a first landmark feature representation, a second landmark feature representation sequence is generated using a first diffusion model, and a target facial landmark sequence is generated by decoding the second landmark feature representation sequence. In this way, facial landmark jitter can be limited, a continuous and stable facial landmark sequence can be generated, and a continuous, natural and expressive video can be generated.

[0032] Various example implementations of the solution are described in detail below in conjunction with the accompanying drawings.

[0033] Example Environment

[0034] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 In the environment 100 of FIG. 1 , a video generation system 110 may utilize one or more machine learning models 105 to perform video generation. The machine learning model 105 is configured to generate a video 116 based on an image 112 and an audio 114. In some embodiments, the image 112 may include an object (e.g., a person, a cartoon character, an animal, etc.), and the video 116 may represent the object giving the speech contained in the audio 114.

[0035] exist Figure 1 In the embodiment, the video generation system 110 can be implemented on any computing system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may include any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a netbook, a tablet computer, a media computer, a multimedia tablet computer, or any combination of the above devices, including accessories and peripheral devices of these devices or any combination thereof. Servers include but are not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0036] The machine learning model 150 can be different types of models. In some embodiments, one or more machine learning models 150 may include a generative model. In some embodiments, one or more machine learning models 150 may be constructed based on a diffusion model. The diffusion model, also known as a diffusion probability model, is a type of generative model. The model generates data by simulating a diffusion process. This process is inspired by physical processes such as thermal diffusion. The diffusion model includes a forward diffusion process and a reverse diffusion process. The diffusion model generates new data samples by simulating a forward diffusion process that gradually adds noise, and then learns how to reverse this process.

[0037] In the forward diffusion process, noise is gradually added to the data, making the data more and more random through a series of steps until the data resembles pure noise. This process can be viewed as a Markov chain, with Gaussian noise added to the data at each step. The forward diffusion process can be expressed as: where x t is the noise data of step t, α t Used to control the amount of noise added. The forward diffusion process is performed during model training, and the data used to add noise are the training samples.

[0038] In the reverse diffusion process (or reverse denoising process), the model learns how to reverse the steps of adding noise. Starting from pure noise, the diffusion model gradually removes the noise to generate data that matches the training distribution. The reverse diffusion process is usually simulated using a neural network that predicts the noise added at each step: where u θ and σ θ are the learned model parameters. After completing the model training, the model performing the back diffusion process can first start sampling from the noise distribution and use the model to iteratively denoise until the desired data is obtained.

[0039] In the diffusion model, the time step refers to the number of steps in which noise is added during the forward diffusion process. The total number of steps T is usually a preset value, indicating how many steps are required to transform the original data into pure noise. At each time step t, Gaussian noise is added to the data according to a predetermined noise scheme. This process is continuous, and each step depends on the result of the previous step.

[0040] When generating data, the diffusion model inference step refers to the number of steps required to recover from pure noise to the original data during the back diffusion process. The number of inference steps directly affects the quality and speed of generated data. Generally, the more inference steps there are, the higher the quality of the generated data is, but it will also increase the computational cost and time. In practical applications, the generation quality and efficiency can be balanced by adjusting the number of inference steps. In some embodiments, the inference step corresponds to a time step, and each inference step can correspond to one or more time steps. For example, if the total time step of the diffusion model is 1000 steps and the inference step is set to 50 steps, then each inference step can correspond to 20 time steps.

[0041] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and does not imply any limitation on the scope of the present disclosure.

[0042] Example Process

[0043] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings. Figure 2 A flowchart of a process 200 of video generation according to some embodiments of the present disclosure is shown. Part or all of the process 200 may be implemented by the video generation system 110 or may be implemented by other devices, such as other remote devices (terminal devices or service devices) with computing capabilities. In the following, for ease of discussion, the execution of the process 200 is described from the perspective of the video generation system 110, but this is only exemplary.

[0044] At block 210, the video generation system 110 determines at least one first landmark feature representation based on at least one set of reference facial landmarks corresponding to at least one reference image of a target object. The target object may include, but is not limited to, a person, a cartoon character, an animal, a cartoon animal, etc. The at least one reference image includes at least a facial image of the target object, for example, the reference image may include a facial image of a target person.

[0045] In some embodiments, the video generation system 110 may, in response to receiving the input information including the at least one reference image, perform facial landmark recognition on the at least one reference image to obtain the at least one set of reference facial landmarks. Specifically, the video generation system 110 may perform facial landmark recognition on each reference image to obtain a set of reference facial landmarks corresponding to each reference image. If the input information includes one reference image, a set of reference facial landmarks may be obtained. If the input information includes multiple reference images, multiple sets of reference facial landmarks may be obtained.

[0046] Each group of reference facial landmarks may include a predetermined number of predetermined landmarks annotated in the corresponding reference image, and the predetermined landmarks may be located at, for example, the facial contour, eyes, nose, mouth, and other positions of the target object. Each group of reference facial landmarks may indicate the facial contour of the target object in the corresponding reference image, as well as the positions and shapes of facial parts such as eyes, nose, and mouth, thereby indicating the facial movements of the target object. In some examples, each group of reference facial landmarks may be two-dimensional facial landmarks (indicated by two-dimensional coordinates) or three-dimensional facial landmarks (indicated by three-dimensional coordinates). The number of each group of reference facial landmarks may vary depending on the facial landmark recognition algorithm. For example, each group of reference facial landmarks may include 21, 49, 68, 98, or 109 predetermined landmarks. Of course, the above positions and numbers of landmarks are merely exemplary, and the embodiments of the present disclosure are not specifically limited thereto.

[0047] In some embodiments, when determining the at least one set of reference facial landmarks, the video generation system 110 may perform encoding on the at least one set of reference facial landmarks to determine the at least one first landmark feature representation. In some examples, the video generation system 110 may determine at least one first landmark feature representation in a latent feature space corresponding to the first diffusion model based on the encoding of the at least one set of reference facial landmarks. Encoding the reference facial landmarks into the latent feature space can reduce information loss during the encoding process and can provide high-quality facial motion representation, which is beneficial to improving the quality of subsequent generation of target facial landmarks and target videos.

[0048] As an example, Figure 3A schematic diagram of an example architecture 300 for video generation according to some embodiments of the present disclosure is shown. In some embodiments, the video generation system 110 may be deployed with a variational autoencoder (VAE), which includes an encoder (VAE encoder) and a decoder (VAE decoder). The example architecture 300 shows an encoder 304. The encoder 304 may be a VAE encoder. Of course, it is understood that the model structure of the encoder 304 may also be configured as other types as long as it is suitable for performing feature encoding. The video generation system 110 may perform two-dimensional facial key point recognition on the at least one reference facial image to obtain at least one set of reference facial landmarks 302. Afterwards, the video generation system 110 may use the encoder 304 to perform encoding on the at least one set of reference facial landmarks 302 to obtain at least one first landmark feature representation 306 in the latent feature space of the first diffusion model 314, which may also be referred to as a latent feature representation.

[0049] Returning to process 200 , in block 220 , the video generation system 110 determines a second landmark feature representation sequence based on the audio feature representation corresponding to the target audio and the at least one first landmark feature representation using the trained first diffusion model.

[0050] In some embodiments, the input information may further include a target audio. The target audio may be used to drive the video generation system 110 to generate facial movements that are synchronized with the target object expressing the speech content in the target audio. The video generation system 110 may perform feature extraction on the target audio to obtain an audio feature representation corresponding to the target audio. In some examples, the video generation system 110 may divide the target audio into multiple audio frames, perform feature extraction on the multiple audio frames, and obtain an audio feature representation sequence corresponding to the multiple audio frames. The audio feature representation sequence may include multiple audio feature representations corresponding to the multiple audio frames, respectively.

[0051] As an example, combining Figure 3 As shown, the example architecture 300 also shows a feature extractor 310. The feature extractor 310 may include a waveform-to-vector model, such as a Wav2Vec model, a Wav2Vec2.0 model, and the like. The video generation system 110 may use the feature extractor 310 to perform feature extraction on multiple audio frames of the target audio to obtain an audio feature representation sequence. Of course, the feature extractor 310 is not limited to the waveform-to-vector model, and other model structures suitable for performing feature extraction on the target audio may also be used, and the embodiments of the present disclosure are not limited to this.

[0052] In some examples, the video generation system 110 can generate a first model input for a first diffusion model based on the audio feature representation and the at least one first landmark feature representation. Based on the first model input, a first model output is generated using the first diffusion model. The first model output can include a second landmark feature representation sequence. As an example, continue with Figure 3 As shown, at block 312, the video generation model 110 may generate a first model input for a first diffusion model 314 based on the audio feature representation sequence and the at least one first landmark feature representation 306. The first model input is provided to the first diffusion model 314 to obtain a first model output generated by the first diffusion model 314. The first model output may include a second landmark feature signature sequence.

[0053] In some examples, the video generation model 110 may generate a feature vector sequence 324 of a predetermined length based on the at least one first landmark feature representation 306. Specifically, the video generation system 110 may place the at least one first landmark feature representation 306 at the beginning of the feature vector sequence 324, and set at least one noise representation 322 (or placeholder) conforming to a Gaussian distribution after the at least one first landmark feature representation to form a feature vector sequence 324 of a predetermined length. Afterwards, the video generation model 110 may generate a first model input for the first diffusion model 314 based on the feature vector sequence 324 and the audio feature representation sequence.

[0054] The second landmark feature representation sequence includes multiple second landmark feature representations corresponding to multiple audio frames in the target audio. In some examples, the first model input may include a first input feature representation sequence, and the first input feature representation sequence may include multiple first input feature representations corresponding to multiple audio feature representations in the audio feature representation sequence. Each first input feature representation can be generated based on a corresponding first landmark feature representation and a corresponding audio feature representation, or based on a noise representation conforming to a Gaussian distribution and a corresponding audio feature representation. For example, at least one first input feature representation at the front of the first input feature representation sequence can be generated based on the corresponding first landmark feature representation and the audio feature representation, respectively, and several subsequent first input feature representations can be generated based on the noise representation and the corresponding audio feature representation, respectively. The video generation system 110 can generate a second landmark feature representation sequence based on the first input feature representation sequence using a first diffusion model.

[0055] Continuing with process 200, at block 230, the video generation system 110 determines a target facial landmark sequence corresponding to the target audio based on decoding the second landmark feature representation sequence. The target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to the plurality of audio frames, respectively. Each group of target facial landmarks may include a predetermined number of predicted predetermined landmarks. Each group of target facial landmarks may indicate a facial action of the target object that matches the corresponding audio frame. In some examples, each group of target facial landmarks may include two-dimensional facial landmarks or three-dimensional facial landmarks.

[0056] As an example, combining Figure 3 As shown, the example structure 300 also shows a decoder 318. The decoder 318 can be a VAE decoder. The video generation system 110 can use the decoder 318 to decode the second landmark feature representation sequence to determine the target facial landmark sequence 320. In this case, the encoder 304 is used to encode the reference facial landmark from the two-dimensional space to the latent feature space to obtain the first landmark feature representation. In the latent feature space, the first diffusion model 314 is used to perform reasoning from audio to facial motion to obtain a second landmark feature representation sequence indicating the facial motion of the target object. Then, the decoder 318 is used to decode the second landmark feature representation sequence into the two-dimensional space to obtain the target facial landmark sequence. In this way, the facial landmark jitter can be limited, thereby obtaining a continuous and stable facial landmark sequence. It can be understood that the decoder 318 is not limited to the VAE decoder, and other model structures suitable for decoding the target facial landmark sequence 320 can also be used.

[0057] In box 240 of process 200, the video generation system 110 generates a target video based on at least one reference image, at least one set of reference facial landmarks, and a target facial landmark sequence. The target audio includes the target object and corresponds to the target audio. The external features of the face of the target object can be indicated by the at least one reference image, and the facial movements of the target object can be indicated by the at least one set of reference facial landmarks and the target facial landmark sequence. Since the generated target facial landmarks have high continuity and stability, a natural and expressive video can be generated based on the target facial landmark sequence. In some examples, the target video may include multiple video frames corresponding to multiple sets of target facial landmarks in the target facial landmark sequence. That is, the generation of the corresponding video frame can be driven by each set of target facial landmarks, and the facial movements of the target object in each video frame can be accurately controlled to achieve continuous control of the video frame.

[0058] In some embodiments, the video generation system 110 may also be deployed with a second diffusion model. The video generation system 110 may perform encoding on the at least one reference image to determine a first visual feature representation. The video generation system 110 performs encoding on the at least one set of reference facial landmarks and the target facial landmark sequence to determine a third landmark feature representation. The video generation system 110 uses the trained second diffusion model to generate a second visual feature representation sequence based on the first visual feature representation and the third landmark feature representation. Afterwards, the second visual feature representation sequence is decoded to obtain the target video.

[0059] Regarding the generation of the first visual feature representation, in some embodiments, the video generation system 110 may perform encoding on the at least one reference image to determine the first visual feature representation in the latent feature space corresponding to the second diffusion model. In this way, the coherence between video frames can be improved, and the jitter between video frames can be limited to a certain extent.

[0060] Figure 4A A schematic diagram of an example architecture 400A for video generation according to some embodiments of the present disclosure is shown. In some examples, the video generation system 110 may be deployed with a three-dimensional causal variational autoencoder (3D causal VAE), which may include a video VAE encoder (Video VAE Encoder) and a video VAE decoder (VideoVAE Decoder). The example architecture 400A shows an encoder 404 and a decoder 432, where the encoder 404 may adopt a video VAE encoder and the decoder 432 may adopt a video VAE decoder. Of course, the above encoder 404 and decoder 432 are merely exemplary, and other model structures suitable for encoding or decoding videos may also be used.

[0061] The video generation system 110 may use the encoder 404 to perform encoding on the at least one reference image to generate a first visual feature representation 410. Using a three-dimensional causal variational autoencoder to perform encoding on the at least one reference image can compress information in the time dimension and the space dimension, which is conducive to the second diffusion model processing more video frames in the subsequent link, and is conducive to generating an extended duration of the target video.

[0062] Continue to combine Figure 4A As shown, the example architecture 400A further shows an encoder 416. The video generation system 110 can use the encoder 416 to encode the reference facial landmark 412 and the target facial landmark sequence 414 to obtain a third landmark feature representation 436. The third landmark feature representation 436 includes a feature representation 418 corresponding to the reference facial landmark 412 and a feature representation 420 corresponding to the target facial landmark sequence. It should be noted that although Figure 4A 4. The feature representation 418 and the feature representation 420 are shown as separate, but in actual application, the feature representation 418 and the feature representation 420 can be combined. In block 422, the video generation system 110 can combine the first visual feature representation 410 and the third landmark feature representation 436 to obtain a second model input for a second diffusion model.

[0063] In some examples, the video generation system 110 may perform encoding on the at least one reference facial landmark and the target facial landmark sequence to obtain a third landmark feature representation that can be feature aligned with the first visual feature representation in the feature space of the second diffusion model. Afterwards, the video generation system 110 may combine the first visual feature representation and the third landmark feature representation along the dimension of the feature channel to form a second model input for the second diffusion model.

[0064] As an example, Figure 4A As shown, the first visual feature representation 410 and the third landmark feature representation 436 may be configured to have a predetermined size, such as (T+1)×W×H. The video generation system 110 may use the encoder 404 to perform encoding on the at least one reference image to obtain a feature vector 406. The video generation system 110 may add a number of noise representations 408 (or placeholders) that conform to the Gaussian distribution after the feature vector 406 to form a first visual feature representation 410 of (T+1)×W×H. The video generation system 110 may also encode the at least one reference facial landmark and the target facial landmark sequence into a third landmark feature representation 436 of (T+1)×W×H to achieve feature alignment of the first visual feature representation 410 and the third landmark feature representation 436 in the latent feature space of the second diffusion model 426. Afterwards, the video generation system 110 may combine the first visual feature representation 410 and the third landmark feature representation 436 into a second model input of (T+1)×2W×H or (T+1)×W×2H.

[0065] As another example, the landmark point encoder may include a three-dimensional convolution layer and a temporal pooling layer, which may be used to compress the feature dimensions of the reference facial landmark points and the target facial landmark point sequence in the temporal dimension and the spatial dimension, respectively, to achieve feature alignment of the third landmark point feature representation and the first visual feature representation.

[0066] After obtaining the second model input, the video generation system 110 may provide the second model input to the second diffusion model 424 to obtain the second model output output by the second diffusion model 424. The second model output may include a second visual feature representation sequence 430. For example, the second diffusion model 424 may be used to generate a second visual feature representation sequence 430 with a length of T and a feature channel of W×H. Afterwards, the video generation system 110 may use a decoder 432 to decode the second visual feature representation sequence to obtain a target video 434. It is to be understood that the above-mentioned target video generation process is only exemplary. In actual applications, any other appropriate process may also be used to generate the target video, and the embodiments of the present disclosure do not specifically limit this.

[0067] In some embodiments, the input information may further include a task indication, which is used to indicate a task type of the video generation task. For example, the task indication may indicate a single reference image video generation task, a video continuation generation task, a video interpolation task, etc. The video generation system 110 may also generate a target video that meets the task type indicated by the task indication based on the task indication.

[0068] As an example, Figure 4A As shown, the task indication may indicate a single reference image video generation task, the input information may include a reference image 402, and the video generation system 110 may encode the reference image 402 using an encoder 404 to obtain a feature vector 406. The video generation system 110 may add a noise representation 408 after the feature vector 406 to obtain a first visual feature representation 410. Afterwards, the video generation system 110 may generate a target video 434 using the first visual feature representation 410, and the target video 434 does not include the reference image 402.

[0069] In some embodiments, when the video generation task indicated by the task indication is a video continuation generation task or a video interpolation task, the input information may also include at least one group of first video frames. The first video frame here may be understood as a given video frame, and each group of first video frames may include one or more video frames. The video generation system 110 may generate a target video based on the at least one group of first video frames, the task indication, at least one reference image, at least one group of reference facial landmarks, and a target facial landmark sequence. The target video includes at least one group of first video frames and at least one group of second video frames. The second video frame here may be understood as a video frame generated by the video generation system 110 based on the target audio. Each group of second video frames may include one or more video frames.

[0070] In some examples, the task instruction may indicate a video continuation generation task. Based on the task instruction, the video generation system 110 may continue to generate at least one set of second video frames after the at least one set of first video frames. In other examples, the video generation system 110 may also continue to generate at least one set of second video frames before the at least one set of first video frames based on the task instruction. That is, reversely continue to generate a set of second video frames.

[0071] As an example, Figure 4B FIG. 4 is a schematic diagram showing an example architecture 400B for video generation according to some embodiments of the present disclosure. Figure 4B As shown, the input information may also include multiple first video frames 440. The video generation system 110 may perform encoding on the reference image 402 and the multiple first video frames 440 to obtain a feature vector 406 corresponding to the reference image 402 and a feature vector 442 corresponding to the multiple first video frames 440, respectively. The video generation system 110 may also add a number of noise representations 408 after the feature vector 442 to obtain a first visual feature representation 410. The target video 434 generated using the first visual feature representation 410 may include multiple first video frames 440 and a group of second video frames 444 located after the multiple first video frames 440. It should be noted that, for the purpose of simplicity, in Figure 4B The process of encoding the reference facial landmark points and the target facial landmark point sequence is not shown in the figure, but these processes also exist in actual application.

[0072] In some examples, the task indication may indicate a video frame insertion task. In the target video generated based on the task indication, each group of first video frames in at least one group of first video frames is arranged at intervals with each group of second video frames in at least one group of second video frames. The interval arrangement here may include inserting a group of second video frames between two groups of first video frames, inserting a group of first video frames between two groups of second video frames, or each group of first video frames in multiple groups of first video frames is arranged at intervals with each group of second video frames in multiple groups of second video frames.

[0073] As an example, Figure 4CA schematic diagram of an example architecture 400C for video generation according to some embodiments of the present disclosure is shown. Assume that the input information includes a first video frame 446 and a first video frame 448. The task indication may indicate a video interpolation task. The video generation system 110 may perform encoding on a reference image 402 and first video frames 446 and 448 to obtain a feature vector 402 corresponding to the reference image 402, a feature vector 450 corresponding to the first video frame, and a feature vector 452 corresponding to the first video frame 446. The video generation system 110 may insert a number of noise representations 408 (or placeholders) between the feature vector 450 and the feature vector 452 to obtain a first visual feature representation 410. The target video generated using the first visual feature representation 410 includes a first video frame 446 and a first video frame 448, and a group of second video frames 440 located between the first video frames 446 and 448. It should be noted that, for the purpose of brevity, in Figure 4C The process of encoding the reference facial landmark points and the target facial landmark point sequence is not shown in the figure, but these processes also exist in actual application.

[0074] It should be noted that the above-mentioned combination of the first video frame and the second video frame is only exemplary, and in practical applications, the first video frame and the second video frame can be combined in any other appropriate manner according to actual task requirements to form a target video. The embodiments of the present disclosure are not limited in this regard.

[0075] In some embodiments, Figure 4A As shown, the second diffusion model 424 may include a base model 426 and an additional model 428, and the additional model 428 is obtained by performing diffusion model distillation on the trained base model 426. The reasoning step of the additional model 428 is less than the reasoning step of the second diffusion model 424. By setting the additional model 428, the reasoning steps of the second diffusion model 424 can be reduced, the computing cost can be reduced, and the video generation speed can be increased. In some examples, the additional model 428 may include a LoRA (Low-Rank Adaptation) model, and the LoRA model can be used to fine-tune the key layers of the base model 426, such as the Transformer Layer.

[0076] The following will introduce the training process of the first diffusion model. It should be understood that such a training process may be performed by an appropriate training device, which may include but is not limited to the video generation system 110 .

[0077] In some embodiments of the present disclosure, a training device may generate a sample facial landmark sequence and a sample audio feature representation of a sample object based on a sample video containing the sample object. The sample facial landmark sequence includes multiple groups of sample facial landmarks corresponding to multiple video frames in the sample video. The sample audio feature representation is generated by encoding the sample audio corresponding to the sample video. Specifically, the training device may obtain multiple video frames of the sample video, perform facial landmark recognition on the multiple video frames, and obtain multiple groups of sample facial landmarks. The training device may also perform feature extraction on the sample audio corresponding to the sample video to obtain a sample audio feature representation. The sample audio here may be the audio in the sample video, or it may be the audio that matches the sample video.

[0078] In some embodiments of the present disclosure, the training device may determine a first sample landmark feature representation based on a sample facial landmark sequence, wherein the first sample landmark feature representation is generated by encoding the sample facial landmark sequence and masking part of the encoding result. The training device may generate a predicted facial landmark sequence based on the first sample landmark feature representation and the sample audio feature representation using a first diffusion model. Thereafter, the training device may adjust parameters in the first diffusion model based on a first difference between the predicted facial landmark sequence and the sample facial landmark sequence to train the first diffusion model.

[0079] Figure 5 FIG. 5 is a schematic diagram showing a training process 500 of a first diffusion model according to some embodiments of the present disclosure. Figure 5 As shown, the training device can obtain a sample video, and the sample video can be represented as V = {f1, f2, ..., f n}, where n>1, f i The training device can use the landmark predictor to extract sample facial landmarks from each video frame of the sample video 502 to obtain a sample facial landmark sequence 502, where each group of sample facial landmarks can be represented as l i ∈R d ×2, d is the number of facial landmark points in each set of sample facial landmark points.

[0080] The training device can use the trained encoder 304 to encode the sample facial landmark sequence 502 to obtain the sample landmark feature representation 504 in the latent feature space of the first diffusion model 314. The sample landmark feature representation 504 can be expressed as x=E(l), l={l1, l2, ..., l n}, x∈R n×c, where E represents the encoder 304, and c represents the feature channel of the sample landmark feature representation 504. When the sample landmark feature representation 504 is obtained, part of the feature representation 508 may be randomly blocked at the end of the sample landmark feature representation 504 to obtain the sample landmark feature representation 506 (i.e., the first sample landmark feature representation). For example, the training device may randomly block the feature representation corresponding to several video frames at the end of the sample facial landmark sequence 502 to obtain the sample landmark feature representation 506. The training device may also use the trained feature extractor 310 to encode the sample audio 510 to obtain the sample audio feature representation.

[0081] In box 512, the training device can generate a model input for the first diffusion model 314 based on the sample landmark feature representation 506 and the sample audio feature representation a. The model input is provided to the first diffusion model 314 to obtain a model output output by the first diffusion model 314. The model output may include a sample landmark feature representation sequence 514. The training device can use the trained decoder 318 to decode the sample landmark feature representation sequence 514 to obtain a predicted facial landmark sequence 516. Thereafter, the training device can adjust the parameters in the first diffusion model 314 based on the first difference between the predicted facial landmark sequence 514 and the sample facial landmark sequence 502 to train the first diffusion model 314.

[0082] In some examples, the training device may adjust parameters in the first diffusion model 314 based on a loss function as shown below.

[0083]

[0084] Among them, L(θ) represents the loss function; x0 represents the sample noise; x1 represents the original data; M represents the mask; x m =M⊙x1, representing the feature representation of the sample landmark point that is blocked by the mask in the feature representation 504; x t =x m +(1-M)☉(tx1+(1-t)x0), represents the data after processing at time step t; For a given x t , time step t and sample audio feature representation a’s predicted value; θ l represents the parameters of the first diffusion model.

[0085] It should be noted that the above training process and loss function are only exemplary, and any other appropriate training process or any other appropriate loss function may be selected according to actual needs to train the first diffusion model, and the embodiments of the present disclosure are not limited to this. It is understandable that during the first diffusion model training process, the above training process may be repeatedly executed multiple times until the model output of the first diffusion model 314 meets the predetermined end condition.

[0086] The training process of the second diffusion model will be introduced below. It should be understood that such a training process may be performed by an appropriate training device, which may include but is not limited to the video generation system 110 .

[0087] In some embodiments of the present disclosure, a training device may determine a sample visual feature representation based on a sample video containing a sample object, wherein the sample visual feature representation is generated by encoding a sample frame sequence of the sample video and masking part of the encoding result. The training device may determine a sample facial landmark sequence based on the sample frame sequence, wherein the sample facial landmark sequence includes multiple groups of sample facial landmarks. Based on the encoding of the sample facial landmark sequence, a second sample landmark feature representation is determined. Based on the sample visual feature representation and the second sample landmark feature representation, a predicted video is generated using a second diffusion model. Afterwards, the training device may adjust parameters in the second diffusion model based on a second difference between the sample video and the predicted video to train the second diffusion model.

[0088] Figure 6 FIG. 6 is a schematic diagram showing a training process 600 of a second diffusion model according to some embodiments of the present disclosure. Figure 6 As shown, the training device may obtain a sample frame sequence in a sample video, and the sample frame sequence may include multiple video frames 604 in the sample video. The training device may randomly determine a video frame from the multiple video frames as a sample reference image 602. The training device may use the trained video VAE encoder to encode the sample reference image 602 and the sample frame sequence 604 to obtain a sample visual feature representation 606 and a sample visual feature representation 608 in the latent feature space of the second diffusion model 424.

[0089] After obtaining the sample visual feature representation 608, the training device may randomly block at least part of the content in the sample visual feature representation 608. As an example, the training device may block the encoding result 610 corresponding to several video frames in the middle part of the sample frame sequence of the sample visual feature representation 608, and form the sample visual feature representation 612 through the unblocked part. As another example, the training device may block the encoding result 614 corresponding to several video frames in the end part of the sample sequence in the sample visual feature representation 608, and form the sample visual feature representation 616 through the unblocked part. As another example, the training device may select several video frames at intervals from the sample sequence, block the encoding result 618 corresponding to the several video frames selected at intervals in the sample visual feature representation 608, and form the sample visual feature representation 620 through the unblocked part. By randomly blocking at least part of the content in the sample visual feature representation 608 during the training process, the trained second diffusion model 424 can be applied to a variety of video generation tasks, such as video continuation generation tasks, video interpolation tasks, and the like. Of course, the above-mentioned blocking method for the sample visual feature representation 608 is only exemplary, and the embodiments of the present disclosure are not limited to this.

[0090] When the sample reference image 602 and the sample frame sequence 604 are determined, the training device may also extract facial landmarks from each video frame of the sample reference image 602 and the sample frame sequence 604 using, for example, a landmark prediction period to obtain sample reference facial landmarks 622 and a sample facial landmark sequence 624. Afterwards, the training device may use the trained encoder 416 to encode the sample reference facial landmarks 622 and the sample facial landmark sequence 624 to obtain a sample landmark feature representation 626 and a sample landmark feature representation 628 (i.e., a second sample landmark feature representation).

[0091] In block 630, the training device may combine any one of the sample visual feature representations 612, 616, 618, the sample visual representation 606, the sample landmark feature representation 626, and the sample landmark feature representation 628 along the dimension of the feature channel to form a model input of the second diffusion model 424. The model input is provided to the second diffusion model 424 to obtain a model output of the second diffusion model 424. The model output may include a predicted visual feature representation sequence 632. The training device may decode the predicted visual feature representation sequence 632 using the trained decoder 432 to obtain a predicted video frame sequence 634. Thereafter, the training device may adjust the parameters of the second diffusion model 424 based on the second difference between the predicted video frame sequence 634 and the sample frame sequence 604 to train the second diffusion model 424.

[0092] In some examples, the training device may adjust parameters in the second diffusion model 424 based on a loss function as shown below.

[0093] L(θ)=(1-M)☉||v θ (z t , t, z ref , c lm )-(z vf -z0)|| 2 (2)

[0094] Where L(θ) represents the loss function; z ref The sample visual feature representation 606 is represented by encoding the sample reference image 602; vf A sample visual feature representation 608 is formed by encoding the sample frame sequence 604; m Indicates random occlusion z in the time dimension vf The sample visual feature representation obtained, such as one of the sample visual feature representations 612, 616, and 618, is z m =M⊙z vf , M represents the mask; z0 represents the sample noise; z t =z m +(1-M)⊙(tz vf +(1-t)z0);c lm The representation encoder 416 encodes the sample landmark point feature representation 626 and the sample landmark point feature representation 628.

[0095] In some embodiments of the present disclosure, the second diffusion model 424 includes a base model 426 and an additional model 428. After the training device completes training the base model 426 using, for example, the training process 600, the training device may perform diffusion model distillation on the trained base model 426 to obtain the additional model 428.

[0096] In some examples, the additional model 428 may be a LoRA model. The training device may adjust the parameters of the LoRA model based on a loss function as shown below to train the LoRA model.

[0097]

[0098] Among them, L(θ lora ) represents the loss function for training the LoRA model, θ represents the parameters of the basic model 426, and θ lora represents the parameters of the LoRA model; z′ represents the sample visual feature representation at time step t=0; The multi-step sampling prediction result of the basic model 426 is shown in Figure 4. The prediction result of the LoRA model can be calculated using formula (3): And the prediction results of the basic model 426 The training device can adjust the parameters of the LoRA model based on the mean square error (MSE) loss between the two.

[0099] It should be noted that the training process and loss function of the second diffusion model are only exemplary. In practical applications, other training processes or loss functions can be used to train the second diffusion model according to actual needs, and the embodiments of the present disclosure are not limited to this. In addition, during the training process of the second diffusion model, the above process may be repeatedly executed multiple times until the model output of the second diffusion model meets the predetermined end condition.

[0100] In this way, in the embodiment of the present disclosure, the reference facial landmarks are encoded as the first landmark feature representation, the second landmark feature representation sequence is generated using the first diffusion model, and the target facial landmark sequence is generated by decoding the second landmark feature representation sequence. In this way, facial landmark jitter can be limited, a continuous and stable facial landmark sequence can be generated, and a continuous, natural and expressive video can be generated.

[0101] Example devices and equipment

[0102] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 7 Schematic block diagram of an example apparatus 700 for video generation according to some embodiments of the present disclosure is shown. The apparatus 700 may be implemented as or included in the video generation system 110. Each module / component in the apparatus 700 may be implemented by hardware, software, firmware or any combination thereof.

[0103] like Figure 7 As shown, the device 700 includes: a first determination module 710, configured to determine at least one first landmark feature representation based on at least one group of reference facial landmarks corresponding to at least one reference image of the target object; a second determination module 720, configured to determine a second landmark feature representation sequence based on the audio feature representation corresponding to the target audio and at least one first landmark feature representation, using a trained first diffusion model, the second landmark feature representation sequence includes a plurality of second landmark feature representations corresponding to a plurality of audio frames in the target audio; a third determination module 730, configured to determine a target facial landmark sequence corresponding to the target audio based on decoding of the second landmark feature representation sequence, the target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to a plurality of audio frames; and a generation module 740, configured to generate a target video based on at least one reference image, at least one group of reference facial landmarks and a target facial landmark sequence, the target audio including the target object and corresponding to the target audio.

[0104] In some embodiments, the first determination module 710 is further configured to determine at least one first landmark feature representation in a latent feature space corresponding to the first diffusion model based on encoding of at least one set of reference facial landmarks.

[0105] In some embodiments, the generation module 740 is further configured to: generate a target video based on at least one set of first video frames of the target object and the task instruction, the target video including at least one set of first video frames and the generated at least one set of second video frames.

[0106] In some embodiments, the task indication indicates a video continuation generation task, and wherein in the target video, at least one group of first video frames is located before or after at least one group of second video frames, or the task indication indicates a video interpolation task, and wherein in the target video, each group of first video frames in at least one group of first video frames is arranged alternately with each group of second video frames in at least one group of second video frames.

[0107] In some embodiments, the generation module 740 is further configured to: determine a first visual feature representation based on encoding of at least one reference image; determine a third landmark feature representation based on encoding of at least one set of reference facial landmarks and a target facial landmark sequence; generate a second visual feature representation sequence based on the first visual feature representation and the third landmark feature representation using a trained second diffusion model; and determine a target video based on decoding of the second visual feature representation sequence, the target video comprising a plurality of video frames corresponding to a plurality of sets of target facial landmarks in the target facial landmark sequence.

[0108] In some embodiments, the generating module 740 is further configured to determine a first visual feature representation in a latent feature space corresponding to the second diffusion model based on encoding of at least one reference image.

[0109] In some embodiments, a target facial landmark sequence is obtained using a first diffusion model, and the first diffusion model is trained in the following manner: based on a sample video containing a sample object, a sample facial landmark sequence and a sample audio feature representation of the sample object are generated, the sample facial landmark sequence includes multiple groups of sample facial landmarks corresponding to multiple video frames in the sample video, and the sample audio feature representation is generated by encoding the sample audio corresponding to the sample video; based on the sample facial landmark sequence, a first sample landmark feature representation is determined, wherein the first sample landmark feature representation is generated by encoding the sample facial landmark sequence and masking part of the encoding result; based on the first sample landmark feature representation and the sample audio feature representation, a predicted facial landmark sequence is generated using the first diffusion model; and based on a first difference between the predicted facial landmark sequence and the sample facial landmark sequence, parameters in the first diffusion model are adjusted to train the first diffusion model.

[0110] In some embodiments, the target video is obtained using a second diffusion model, and the second diffusion model is trained in the following manner: based on a sample video containing a sample object, a sample visual feature representation is determined, wherein the sample visual feature representation is generated by encoding a sample frame sequence of the sample video and masking part of the encoding result; based on the sample frame sequence, a sample facial landmark sequence is determined, and the sample facial landmark sequence includes multiple groups of sample facial landmarks; based on the encoding of the sample facial landmark sequence, a second sample landmark feature representation is determined; based on the sample visual feature representation and the second sample landmark feature representation, a predicted video is generated using a second diffusion model; and based on a second difference between the sample video and the predicted video, parameters in the second diffusion model are adjusted to train the second diffusion model.

[0111] In some embodiments, the target video is obtained using a second diffusion model, the second diffusion model includes a basic model and an additional model, the additional model is obtained by performing diffusion model distillation on the trained basic model, wherein the reasoning step of the additional model is less than the reasoning step of the second diffusion model.

[0112] The units and / or modules included in the device 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 700 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0113] Figure 8 8 is a block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 8 The electronic device 800 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 8 The electronic device 800 shown may include or be implemented as Figure 1 The video generation system 110, or Figure 7 Device 700.

[0114] like Figure 8 As shown, the electronic device 800 is in the form of a general electronic device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processors 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processor 810 may be an actual or virtual processor and is capable of performing various processes according to executable instructions stored in the memory 820. In a multi-processor system, multiple processors execute computer executable instructions in parallel to improve the parallel processing capabilities of the electronic device 800.

[0115] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable medium, and can include a machine-readable medium, such as a flash drive, a disk, or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 800.

[0116] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 8 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer executable instruction product 825 having one or more executable instruction modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0117] The communication unit 840 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 800 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 800 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0118] The input device 850 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 860 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 800 may also communicate with one or more external devices (not shown) through the communication unit 840 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 800, or communicate with any device that allows the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0119] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer-executable instruction product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0120] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, devices, equipment, and computer-executable instruction products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer-readable executable instructions.

[0121] These computer executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer executable instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0122] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0123] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer executable instruction products according to multiple implementations of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, an executable instruction or a part of an instruction, and a module, an executable instruction or a part of an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0124] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A video generation method, comprising: determining at least one first landmark feature representation based on at least one set of reference facial landmarks corresponding to at least one reference image of the target object; Based on the audio feature representation corresponding to the target audio and the at least one first landmark feature representation, using the trained first diffusion model, determining a second landmark feature representation sequence, wherein the second landmark feature representation sequence includes a plurality of second landmark feature representations corresponding to a plurality of audio frames in the target audio; Determine a target facial landmark sequence corresponding to the target audio based on decoding of the second landmark feature representation sequence, wherein the target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to the plurality of audio frames respectively; as well as A target video is generated based on the at least one reference image, the at least one set of reference facial landmarks, and the target facial landmark sequence, wherein the target audio includes the target object and corresponds to the target audio.

2. The method according to claim 1, wherein determining the at least one first landmark feature representation comprises: Based on the encoding of the at least one set of reference facial landmarks, the at least one first landmark feature representation in a latent feature space corresponding to the first diffusion model is determined.

3. The method according to claim 1, wherein generating the target video comprises: The target video is also generated based on the at least one set of first video frames of the target object and the task instruction, wherein the target video includes the at least one set of first video frames and the generated at least one set of second video frames.

4. The method of claim 3, wherein the task indication indicates a video continuation generation task, and wherein in the target video, the at least one set of first video frames is located before or after the at least one set of second video frames, or The task indication indicates a video frame insertion task, and in the target video, each group of first video frames in the at least one group of first video frames is arranged alternately with each group of second video frames in the at least one group of second video frames.

5. The method according to claim 1, wherein generating the target video comprises: determining a first visual feature representation based on encoding the at least one reference image; determining a third landmark feature representation based on encoding the at least one set of reference facial landmarks and the target facial landmark sequence; Based on the first visual feature representation and the third landmark feature representation, using the trained second diffusion model, generating a second visual feature representation sequence; as well as Based on decoding the second visual feature representation sequence, the target video is determined, where the target video includes a plurality of video frames respectively corresponding to a plurality of groups of target facial landmarks in the target facial landmark sequence.

6. The method of claim 5, wherein determining the first visual feature representation comprises: Based on encoding of the at least one reference image, the first visual feature representation in a latent feature space corresponding to the second diffusion model is determined.

7. The method according to claim 1, wherein the target facial landmark sequence is obtained using a first diffusion model, and the first diffusion model is trained in the following manner: Based on a sample video containing a sample object, generating a sample facial landmark sequence and a sample audio feature representation of the sample object, wherein the sample facial landmark sequence includes a plurality of groups of sample facial landmarks corresponding to a plurality of video frames in the sample video, and the sample audio feature representation is generated by encoding a sample audio corresponding to the sample video; Based on the sample facial landmark sequence, determining a first sample landmark feature representation, wherein the first sample landmark feature representation is generated by encoding the sample facial landmark sequence and masking part of the encoding result; Based on the first sample landmark feature representation and the sample audio feature representation, using the first diffusion model, generating a predicted facial landmark sequence; and Based on a first difference between the predicted facial landmark sequence and the sample facial landmark sequence, parameters in the first diffusion model are adjusted to train the first diffusion model.

8. The method according to claim 1, wherein the target video is obtained using a second diffusion model, and the second diffusion model is trained in the following manner: Determining a sample visual feature representation based on a sample video containing a sample object, wherein the sample visual feature representation is generated by encoding a sample frame sequence of the sample video and masking a portion of the encoding result; Based on the sample frame sequence, determining a sample facial landmark point sequence, the sample facial landmark point sequence comprising a plurality of groups of sample facial landmark points; Determining a second sample landmark feature representation based on encoding the sample facial landmark sequence; Based on the sample visual feature representation and the second sample landmark feature representation, using the second diffusion model, generating a predicted video; as well as Based on a second difference between the sample video and the predicted video, parameters in the second diffusion model are adjusted to train the second diffusion model.

9. The method according to claim 1, wherein the target video is obtained using a second diffusion model, the second diffusion model includes a basic model and an additional model, the additional model is obtained by performing diffusion model distillation on a trained basic model, wherein the reasoning step of the additional model is less than the reasoning step of the second diffusion model.

10. A device for video generation, comprising: A first determination module configured to determine at least one first landmark feature representation based on at least one set of reference facial landmarks corresponding to at least one reference image of the target object; A second determination module is configured to determine a second landmark feature representation sequence using a trained first diffusion model based on the audio feature representation corresponding to the target audio and the at least one first landmark feature representation, wherein the second landmark feature representation sequence includes a plurality of second landmark feature representations respectively corresponding to a plurality of audio frames in the target audio; A third determination module is configured to determine a target facial landmark sequence corresponding to the target audio based on decoding of the second landmark feature representation sequence, wherein the target facial landmark sequence includes a plurality of groups of target facial landmarks corresponding to the plurality of audio frames respectively; as well as A generating module is configured to generate a target video based on the at least one reference image, the at least one set of reference facial landmarks and the target facial landmark sequence, wherein the target audio includes the target object and corresponds to the target audio.

11. An electronic device, comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.

12. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer executable instruction product, comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.