Video generation method and device, equipment and storage medium

By extracting features from the reference image and determining the mixed posture sequence, and performing reverse denoising operations in the denoising network, the problem of inadequate natural and stable digital human video generation in the prior art is solved, and high-quality and coherent human action video generation is achieved.

CN120166239APending Publication Date: 2025-06-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330286.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing digital human video generation technology is difficult to generate human action videos with coherent movements and clear quality, especially when it involves hand and facial movements, the movement performance is not natural and stable enough.

Method used

By extracting the reference features of the target object from the reference image, determining the mixed pose sequence, and adding initial noise to the denoising network, a reverse denoising operation is performed to generate a video of the target object. This method combines appearance coding features, local features and semantic coding features, and uses mixed pose sequences and denoising networks to improve the quality and naturalness of video generation.

Benefits of technology

It realizes the generation of human body movement videos with coherent movements and clear quality, especially in hand and facial movements. The movement performance is more natural and stable, improving the overall quality of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166239A_ABST
    Figure CN120166239A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models, augmented reality and the like, and can be applied to scenes such as digital human. According to the specific implementation scheme, reference features of a target object are extracted from a reference image; determining a mixed attitude sequence; a plurality of target parts in the mixed attitude sequence adopt corresponding description modes to enhance attitude characteristics of the corresponding parts; and adding initial noise into the denoising network, and performing reverse denoising operation on the reference features and the mixed attitude sequence to generate a video of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technical fields such as computer vision, deep learning, large models, and augmented reality, and can be applied to scenarios such as digital humans. Background Art

[0002] With the rapid development of Internet technologies and artificial intelligence, artificial intelligence agent generation technologies provide technical support for the generation of digital humans. The generation of digital humans has been widely applied in application scenarios such as the metaverse, intelligent customer service, and e-commerce. It can not only provide a more intuitive and personalized service experience but also significantly improve user engagement and satisfaction. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a video generation method, including:

[0005] extracting reference features of a target object from a reference image;

[0006] determining a mixed pose sequence; multiple target parts in the mixed pose sequence adopt corresponding description methods to enhance the pose characteristics of the corresponding parts;

[0007] adding initial noise to a denoising network and performing a reverse denoising operation on the reference features and the mixed pose sequence to generate a video of the target object.

[0008] According to another aspect of the present disclosure, there is provided a video generation apparatus, including:

[0009] a first extraction module for extracting reference features of a target object from a reference image;

[0010] a second extraction module for determining a mixed pose sequence; multiple target parts in the mixed pose sequence adopt corresponding description methods to enhance the pose characteristics of the corresponding parts;

[0011] a generation module for adding initial noise to a denoising network and performing a reverse denoising operation on the reference features and the mixed pose sequence to generate a video of the target object.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements any method in the embodiments of the present disclosure when executed by a processor.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings

[0019] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0020] Figure 1 is a schematic flowchart of a video generation method according to an embodiment of the present disclosure;

[0021] Figure 2 is a schematic flowchart of extracting a mixed pose sequence from a driving signal according to an embodiment of the present disclosure;

[0022] Figure 3 is a schematic diagram of each image required to obtain a mixed pose sequence according to an embodiment of the present disclosure;

[0023] Figure 4 is a schematic flowchart of generating a video of a target object according to an embodiment of the present disclosure;

[0024] Figure 5 is a schematic flowchart of processing an input feature based on a denoising network according to an embodiment of the present disclosure;

[0025] Figure 6 is a schematic diagram of the working principle of a denoising module in a denoising network according to an embodiment of the present disclosure;

[0026] Figure 7 is a schematic diagram of the interaction structure between a reference network and a denoising network according to an embodiment of the present disclosure;

[0027] Figure 8 is another schematic diagram of the interaction structure between a reference network and a denoising network according to an embodiment of the present disclosure;

[0028] Figure 9 It is a schematic diagram of the overall architecture of a video generation method according to an embodiment of the present disclosure;

[0029] Figure 10 It is a schematic structural diagram of a video generation device according to an embodiment of the present disclosure;

[0030] Figure 11 It is a block diagram of an electronic device for implementing the video generation method of the embodiments of the present disclosure. Specific embodiments

[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0032] With the continuous development of Internet technology and artificial intelligence, the application of digital humans has gradually expanded from the face to the entire human body. The generated human motion videos can be used in scenarios such as digital human broadcasting and dance video generation.

[0033] Human motion generation methods usually rely on a human reference image, use some driving signals as conditions, and then drive the human reference image to move. Digital human video generation technology aims to generate human motion videos with coherent motion and clear quality.

[0034] To improve the quality of digital human video generation, embodiments of the present disclosure provide a video generation method. As Figure 1 shown, it is a flowchart of the video generation method provided by the embodiments of the present disclosure, including the following:

[0035] S101, extract the reference features of the target object from the reference image.

[0036] In the video task of generating the target object, the reference image is one of the inputs, and the reference image is an image containing the target object. The video generation method provided by the embodiments of the present disclosure can generate coherent actions of the target object based on the driving signal to generate the video of the target object.

[0037] The reference image provides texture information such as the identity and clothing of the target object, which is used to ensure that the target object in the generated video has the same key features as those in the reference image during the video generation process of the target object, so as to ensure the quality of the generated video.

[0038] Among them, the reference features extracted from the reference image are key feature data that can represent the target object. These reference features are usually extracted and encoded by an encoder for use in the subsequent video generation process, so that the target object has relatively stable appearance features in the generated video.

[0039] S102. Determine the mixed pose sequence; corresponding description methods are used for multiple target parts in the mixed pose sequence to enhance the pose characteristics of the corresponding parts.

[0040] During implementation, the mixed pose sequence can be extracted from the drive signal. The drive signal is an input signal used to control the actions of the digital human. The mixed pose sequence extracted from the drive signal is a fusion of various pose data, such as facial actions, body actions, etc., and is used to guide the action generation of the target object in the video.

[0041] Using corresponding description methods for multiple target parts in the mixed pose sequence to enhance the pose characteristics of the corresponding parts can express the action characteristics that the target object needs to show in the generated video from different dimensions or angles, so that in the generated video, the neural network can independently focus on different target parts respectively, making the action performance of the generated target object more accurate, smooth and natural. For example, human key points are used to describe the pose of the body (i.e., the torso and limbs), and 3D HandMesh is used to accurately describe the pose of the hand.

[0042] Therefore, the mixed pose sequence provided in the embodiments of the present disclosure can be understood as independently describing different target parts of the overall pose in the same drive signal using different technical description methods, and the pose used to drive the target object in the drive signal is not limited to the extraction method of a certain technology or model.

[0043] S103. Add initial noise to the denoising network, and perform a reverse denoising operation on the reference features and the mixed pose sequence to generate a video of the target object.

[0044] The initial noise is usually a random Gaussian noise. The initial noise, the reference features, and the mixed pose sequence are input into the denoising network together as conditions. The denoising network can gradually reduce the noise based on these inputs to generate a video of the target object.

[0045] The reverse denoising operation is the reverse process of the diffusion model, opposite to the forward diffusion process. In the forward diffusion process, the original data is gradually added with noise and finally transformed into pure noise data; while in the reverse denoising process, the model starts from pure noise data, gradually removes the noise, and finally restores the original data. Therefore, the denoising network is constructed based on the diffusion model and can recover the image of the target object from the noise to generate a video.

[0046] In the embodiments of the present disclosure, extracting the reference features of the target object from the reference image can accurately describe the unique features of the target object, such as texture information such as identity and clothing, ensuring that the generated video conforms to the expected target character features. Extracting the mixed pose sequence from the driving signal, the driving signal provides the original control information, while the mixed pose sequence processes and adapts this information, and can better guide the digital human motion generation system in the generated video of the target object to generate actions that meet the expectations. The mixed pose sequence uses corresponding description methods for different target parts, which can enhance the characteristics of the corresponding parts and improve the quality of the input, so that during the process of generating the video, the characteristics of important target parts can be inferred and understood, and finally the pose characteristics can be accurately understood, and a video that conforms to the characteristics of the target object and the pose requirements expressed by the driving signal can be gradually generated from the noise, improving the accuracy of the pose of the target object in the generated video, making the pose changes in the video more stable and natural, thereby improving the quality of the generated video.

[0047] In the embodiments of the present disclosure, the multiple target parts in the mixed pose sequence may include at least one of the following: the face, the hand, and the part of the body other than the hand. Since the actions of the body other than the hand have relatively discrete key points compared to the face and are not easily occluded, during implementation, the description in the mixed pose sequence can be preferentially optimized for the face and / or the easily occluded hand. To further ensure the accuracy of other parts of the body, the part of the body other than the hand can also be optimized. By adding special descriptions of the target parts in a differentiated manner in the mixed pose sequence, the understanding of the target part features in the input signal can be improved, ensuring that the actions of the digital human in these key parts in the generated video of the target object are more natural and realistic, thereby ensuring the identity consistency and action stability of the generated video of the target object.

[0048] During implementation, the face in the target part in the mixed pose sequence can be described using face key points; the hand in the target part can be described using a three-dimensional gray scale model.

[0049] A three-dimensional gray scale model generally refers to a model that represents the surface features of a three-dimensional object based on gray values. In the fields of computer vision and graphics, it can be used as an intermediate representation form to describe the characteristics such as the shape of an object.

[0050] During implementation, the hand pose obtained through the hand pose estimation model can be rendered to obtain a three-dimensional grayscale model diagram of the hand. For example, an action video can be used as a driving signal, and it is expected to use the action sequence of the reference object in the action video to drive the target object. That is to say, in the generated video, the target object performs the same expressions and actions as the reference object in the driving video. It can also be understood that the pose of the target object mimics the pose of the reference object. During implementation, the driving signal is input into the HaMer (Hand Mesh Recovery, a hand pose estimation model) network model, and the network model extracts from the action video a hand rendering diagram of the reference object described by the MANO (Mesh And Norms, a hand pose estimation model) based on the hand 3D Mesh parametric model, so as to obtain a three-dimensional grayscale model diagram of the hand.

[0051] In the embodiments of the present disclosure, using a three-dimensional grayscale model diagram for the hand in the target part of the mixed pose sequence can more accurately capture the subtle movements of the hand. Compared with simple two-dimensional images or key point markings, the three-dimensional grayscale model diagram provides richer information, which helps to construct a more realistic hand model, making the generated video of the target object appear more natural and stable when involving hand operations.

[0052] Furthermore, to simplify the understanding and analysis of the hand movements in the target part, and to visually and effectively distinguish between the left hand and the right hand, different colors are used to distinguish the left hand and the right hand of the hand in the target part of the mixed pose sequence.

[0053] During implementation, different colors can be assigned to the left hand and the right hand in the mixed pose sequence. For example, green is used to represent the left hand and pink is used to represent the right hand.

[0054] In the embodiments of the present disclosure, using different colors to distinguish between the left hand and the right hand in the mixed pose sequence can enable the model to better distinguish between the left hand and the right hand, thereby better reconstructing the hand pose of the target object in the generated video and improving the stability and accuracy of the hand region reconstruction in the generated video.

[0055] In the embodiments of the present disclosure, the implementation method for determining the mixed pose sequence can be as Figure 2 shown, including:

[0056] S201, input the driving signal into the hand pose estimation model to obtain a hand rendering diagram of the reference object in the driving signal output by the hand pose estimation model; the hand rendering diagram is used to describe the three-dimensional mesh parameters of the hand of the reference object.

[0057] The hand pose estimation model is a deep learning model specialized for analyzing and estimating the pose of the hand. For example, the HaMer model can achieve high-precision three-dimensional hand reconstruction through a monocular RGB image to obtain hand pose parameters; the MANO model can map the shape parameters of the hand, such as finger lengths, palm widths, and pose parameters, such as joint angles, to a detailed 3D hand mesh.

[0058] Specifically in implementation, the image containing the hand in the drive signal can be input into the hand pose estimation model HaMer, and the hand pose parameters can be extracted based on this model. Then, according to the extracted pose parameters, the hand pose estimation model MANO is used to generate a three-dimensional mesh model of the hand and render it into a hand rendering image, as Figure 3 shown in Figure b in. The hand rendering image provides three-dimensional mesh parameter information of the hand for accurately representing the pose and shape of the hand in subsequent mixed pose sequences.

[0059] S202, Extract the hand mask image of the reference object and the human skeleton parameter image of the reference object from the drive signal.

[0060] The hand mask image refers to a binary image used to distinguish the hand from other parts, as Figure 3 shown in Figure c in. The human skeleton parameter image describes the positions and directions of the joints of the human body and can be used to guide the subsequent generation of the pose of the target object, which may include facial expressions and body movements, as Figure 3 shown in Figure a in.

[0061] The human skeleton parameter image of the reference object includes the facial key points and body key points of the head of the reference object. Among them, the body refers to the main body part of the human body except the head, including the limbs and the torso.

[0062] S203, Based on the hand mask image, fuse the human skeleton parameter image and the hand rendering image to obtain a mixed pose sequence with the hand rendering image as the foreground image.

[0063] In the embodiments of the present disclosure, a limited number of key points are used in the human skeleton parameter image to describe the hand movements, but the ability of these key points to express the details of the hand movements is limited. The hand rendering image can express the hand movements more accurately and in detail. Therefore, the purpose of fusing the hand rendering image as the foreground image with the human skeleton parameter image is mainly to use the hand rendering image to replace the hand part in the human skeleton parameter image and ensure that the fused hand rendering image will not be blocked by other parts. The other parts in the human skeleton parameter image are still described by the human skeleton parameter image.

[0064] Using the hand mask image as a guide, combine the human skeleton parameter image and the hand rendering image to generate a mixed pose sequence, asFigure 3 As shown in Figure d. Among them, the hand rendering image is used as the foreground image, which details the three-dimensional posture of the hand and minimizes the reconstruction loss caused by hand occlusion. The human skeleton parameter image is used as the layer after the foreground, providing the overall posture of the reference object and capable of providing high-quality driving data for the posture generation of the target object.

[0065] During implementation, a mixed posture sequence can be obtained by fusing based on the method described in formula (1):

[0066] I f = I b ×(1 - I m ) + I h × I m (1)

[0067] In formula (1), I f represents the mixed posture image, I b represents the human skeleton parameter image, I m represents the hand mask image, and I h represents the hand rendering image.

[0068] In the embodiments of the present disclosure, by inputting the driving signal into the hand posture estimation model to obtain the hand rendering image of the reference object, the accuracy and robustness of hand posture estimation can be significantly improved. Based on the hand mask image, by fusing the human skeleton parameter image and the hand rendering image to obtain a mixed posture sequence with the hand rendering image as the foreground image, it can provide more explicit conditional guidance for the hand, improving the stability and accuracy of the hand in the generated video.

[0069] In the embodiments of the present disclosure, since the facial key points are relatively dense, when performing certain actions, such as swinging the head left and right, the side face state is likely to occlude the key points, making the key points of different parts of the face prone to confusion. To prevent confusion, for the face in the target part of the mixed posture sequence, different highlighting description methods are used for the key points of different parts of the face.

[0070] During implementation, different colors can be given to the key points of different parts. For example, the distance between the eyes and the eyebrows is relatively close, so the key points around the eyes and the key points between the eyebrows can be distinguished by different colors.

[0071] In the embodiments of the present disclosure, for the key points of different parts of the face in the mixed posture sequence, different highlighting description methods are used for distinction, enabling the model to better distinguish the key points of different parts and reconstruct the facial expressions of the target object. The key points of different parts of the face are distinguished by different colors, providing strong conditional guidance for facial expression reconstruction and capable of improving the stability and accuracy of the facial expressions of the target object in the video.

[0072] In the embodiments of the present disclosure, in order to improve the stability and accuracy of action generation, for the human key points in the mixed pose sequence, the description value used to describe the key points is positively correlated with the confidence of the key points.

[0073] Human key points represent the main joint positions of the human body in the reference image. By extracting these key points, a basic skeleton structure of the human body can be constructed, and based on this, the actions of the digital human, i.e., the target object, in the video of the target object can be guided. During implementation, the description value used to describe the key points can be represented by brightness. When the confidence of the key point is higher, the brightness of the key point is also higher. The human key points can be understood as the key points after the fusion of the warp in the human skeleton parameter map and the hand rendering map.

[0074] In the embodiments of the present disclosure, for the human key points in the mixed pose sequence, the description value used to describe the key points is positively correlated with the confidence of the key points, which helps to ensure the accuracy and reliability of the key points, reduce the errors and uncertainties caused by low-confidence key points, and ensure the coherence and naturalness of the actions of the digital human in the video of the target object.

[0075] In the embodiments of the present disclosure, the reference features of the target object extracted from the reference image include at least one of the following:

[0076] (1) Appearance encoding features;

[0077] Appearance encoding features mainly focus on the overall visual appearance of the target object, such as the appearance of the target object, texture information of clothing, etc.

[0078] During implementation, the reference image can be input into a VAE (Variational Autoencoder Encoder, the encoding part of the variational autoencoder) encoder to obtain appearance encoding features.

[0079] (2) Semantic encoding features;

[0080] Semantic encoding features focus on the content of the image, that is, what each part in the image represents. For example, identifying which regions in the reference image correspond to the head, body, background, etc., and understanding the relationships between them.

[0081] During implementation, the reference image can be input into a semantic feature encoder, such as CLIP (Contrastive Language–Image Pretraining Encoder), to obtain semantic encoding features.

[0082] (3) Local encoding features; the local encoding features include facial features and / or hand features.

[0083] Local coding features focus on the details of specific parts in the reference image, such as facial features and hand features.

[0084] In the embodiments of the present disclosure, appearance coding features help to ensure that the generated video frames are visually consistent with the original reference image, which helps to improve the comprehensibility and coherence of the generated video. Semantic coding features can help the network understand and generate a logical action sequence by performing a high-level understanding of the content of the reference image, making the generated content more reasonable and natural. Local coding features can capture the details of specific parts in the reference image, enabling further optimization of the generated video frames locally and making the generated local postures more accurate and stable.

[0085] In some embodiments, extracting the hand features of the target object from the reference image can be completed based on the following steps:

[0086] Step A1, extract the left-hand local map and the right-hand local map of the target object from the reference image;

[0087] During implementation, computer vision techniques can be used to identify and separate the hand regions in the reference image. To ensure the accuracy of subsequent processing, independent local maps are usually generated for each hand, namely the left-hand local map and the right-hand local map.

[0088] Step A2, input the left-hand local map and the right-hand local map into the hand encoder to obtain hand features.

[0089] During implementation, the extracted left-hand local map and right-hand local map can be input through corresponding channels to combine them into a complete hand map. By separately distinguishing the left and right hands through different channels, corresponding conditional guidance can be established for each hand, improving the signal quality of the left and right hands. Further input into the hand encoder to obtain hand features, which can include the shape, texture, joint positions, etc. of the hand, thus facilitating the neural network model to accurately understand the hand posture.

[0090] To achieve efficient hand feature coding, the hand encoder can adopt a lightweight convolutional layer design and be trained together with the backbone network to learn hand feature coding.

[0091] In the embodiments of the present disclosure, extracting the left and right hand local maps of the target object from the reference image and inputting them into the hand encoder to obtain hand features, focusing on the hand region, can capture the details of the hand more precisely, thereby providing more accurate hand feature information for the subsequent generation process and making the generated digital human hand movements more natural and realistic.

[0092] Correspondingly, extracting the facial features of the target object from the reference image can be completed based on the following steps:

[0093] Step B1: Extract the facial local map of the target object from the reference image;

[0094] During implementation, a facial detection model can be used to accurately extract the facial local map containing the facial position from the reference image.

[0095] Step B2: Input the facial local map into the facial encoder to obtain facial features.

[0096] Based on the obtained facial local map, input it into the facial encoder to obtain facial features, such as the shape of the eyes, the contour of the nose, the expression of the mouth, etc.

[0097] The facial encoder can be constructed using a relatively mature neural network model capable of extracting facial expressions.

[0098] In some other embodiments, similar to the hand encoder, the facial encoder can also adopt a lightweight convolutional layer design. Or an existing mature facial encoder can be used to accurately extract facial features for reconstructing the facial expression of the target object.

[0099] In the embodiments of the present disclosure, the facial local map of the target object is extracted from the reference image and input into the facial encoder to obtain facial features. By focusing on feature extraction in the facial region, facial expression details can be captured more precisely, thereby providing more accurate facial feature information for subsequent generation of the video of the target object, and ensuring identity consistency and expression richness.

[0100] In the embodiments of the present disclosure, an initial noise is added to the denoising network, and a reverse denoising operation is performed on the reference feature and the mixed pose sequence to generate the video of the target object. The method, as Figure 4 shown, includes:

[0101] S401: Input the appearance encoding feature in the reference feature into the reference network to obtain an intermediate feature.

[0102] The reference network is a neural network architecture used for digital human motion generation tasks.

[0103] In the embodiments of the present disclosure, the main function of the reference network is to extract intermediate features related to the target object from the reference image and pass these features to the subsequent denoising network.

[0104] S402: Input the intermediate feature, local feature, initial noise, and semantic encoding feature into the denoising network, and the denoising network processes them to obtain the action feature of the target object.

[0105] During implementation, the intermediate feature, local feature, initial noise, and semantic encoding feature can be input into the denoising network, and the following operations are performed by multiple denoising modules in the denoising network:

[0106] S4021, process the intermediate features, the noise features corresponding to the initial noise, and the mixed pose sequence based on the spatial attention layer in the denoising module to obtain the first output feature.

[0107] S4022, process the first output feature and the semantic encoding feature based on the first cross-attention layer in the denoising module to obtain the second output feature.

[0108] S4023, process the second output feature and the local feature based on the second cross-attention layer in the denoising module to obtain the third output feature.

[0109] S4024, process the third output feature based on the temporal attention layer in the denoising module to obtain the fourth output feature. Among them, the fourth output feature output by the temporal attention layer of the last denoising module of the denoising network is the action feature of the target object.

[0110] During implementation, as Figure 5 shown, the intermediate features, the local features, the initial noise, the semantic encoding features, and the mixed pose sequence can be input into denoising module 1 of the denoising network, and the fourth encoded feature is output after being processed by the attention layers of each level; then, the fourth encoded feature output by denoising module 1, the intermediate features, the semantic encoding features, and the local features are input into denoising module 2 to obtain a new fourth output feature; and so on. The output feature of each denoising module will be used as one of the inputs of the next module, and finally, the action feature of the target object is output by denoising module n.

[0111] To more clearly show the processing flow of the denoising network, the processing flow of denoising module 1 is taken as an example for detailed description below. As Figure 6 shown, by taking the obtained intermediate features, the noise features corresponding to the initial noise, and the mixed pose sequence as inputs and inputting them into denoising module 1 of the denoising network, the first output feature is obtained through the spatial attention layer 601. Then, the obtained first feature and the semantic encoding feature are input into the first cross-attention layer 602, and the first feature is further optimized based on the semantic information of the semantic encoding feature to obtain the second output feature. Next, the obtained second feature and the local feature are input into the second cross-attention layer 603, and the local feature is incorporated into the second feature to obtain the third output feature. Finally, the third output feature is input into the temporal attention layer 604 to obtain the fourth output feature.

[0112] It should be noted that the denoising network and the reference network in the embodiments of the present disclosure have corresponding structures. Each denoising module in the denoising network corresponds one-to-one with the feature construction module in the reference network.

[0113] In a possible implementation manner, referring to Figure 7As shown, the reference network includes a plurality of feature construction modules connected in series in sequence. Each feature construction module includes a spatial attention layer, a cross-attention layer, and a temporal attention layer connected in series in sequence. Among them, the output of the spatial attention layer of the reference network is used as intermediate features and input to the corresponding denoising module in the denoising network. For example, the intermediate feature 1 output by the spatial attention layer in the first feature construction module of the reference network is input to the spatial attention layer of the first denoising module in the denoising network. The intermediate feature 2 output by the spatial attention layer in the second feature construction module of the reference network is input to the spatial attention layer of the second denoising module in the denoising network. And so on, the intermediate feature n output by the spatial attention layer in the nth feature construction module of the reference network is input to the spatial attention layer of the nth denoising module in the denoising network.

[0114] Referring to Figure 7 As shown, the semantic encoding features are input to the denoising module in the denoising network and at the same time input to the corresponding feature construction module in the reference network. So that the reference network can continuously optimize the intermediate features in the subsequent feature construction modules and improve the accuracy and stability of the pose of the reconstructed target object.

[0115] Figure 7 Shows different structures of the feature construction module and the denoising module. In some other embodiments, the feature construction module and the denoising module may have the same network structure. As Figure 8 described, like the denoising module, the feature construction module can have a spatial attention layer, a first cross-attention layer, a second cross-attention layer, and a temporal attention layer. The difference is that the model parameters of the two modules are independent.

[0116] In Figure 8 , the processing method of the intermediate features and the processing method of the semantic encoding features are both the same as Figure 7 . The difference is that Figure 8 the feature construction module and the corresponding denoising module share local features. That is, as Figure 8 shown, the local features are not only input to the denoising module for processing, but also input to the second cross-attention layer of the corresponding feature construction module for processing, so as to improve the quality of the intermediate features.

[0117] In the embodiments of the present disclosure, the spatial attention layer can automatically learn the importance weights of the input features at different spatial positions, thereby enhancing the model's ability to capture and express the actions of key parts. The cross-attention layer enables the model to better understand the correlation and dependence between features through the interaction between different features. Among them, the second cross-attention layer further fuses local features, can adopt more local information guidance, and improves the accuracy of local information construction; the temporal attention layer can ensure that the generated actions are coherent and natural in the time dimension, avoiding abrupt or inconsistent action transitions. Thus, through the attention layers at each level, each input feature can be gradually refined and optimized, so that the finally obtained action features can more accurately reconstruct the action intention and details of the target object.

[0118] Further, in the case where the local features include facial features and hand features, in the denoising module, the facial features are processed based on the facial attention sub-layer in the second cross-attention layer, and the hand features are processed based on the hand attention sub-layer.

[0119] As Figure 6 shown, the facial attention sub-layer 6031 is specifically designed to process facial features, aiming to enhance the expressiveness of details such as facial expressions and postures. Based on the input second output feature and the facial features, the cross-attention mechanism is used to dynamically adjust the importance weights of the facial features, so that the model can more accurately capture the subtle changes in the face.

[0120] The hand attention sub-layer 6032 focuses on the processing of hand features, aiming to capture and enhance details such as finger positions and gesture forms. Based on the input second output feature and the facial features, the cross-attention mechanism is used to dynamically adjust the importance weights of the hand features to ensure the smoothness and naturalness of hand actions.

[0121] It should be noted that in the denoising module, the facial features are processed based on the facial attention sub-layer in the second cross-attention layer, and the hand features are processed based on the hand attention sub-layer, and there is no order of processing between the two.

[0122] In the embodiments of the present disclosure, by using the facial attention sub-layer and the hand attention sub-layer respectively, the performance of the face and hands can be independently and accurately optimized without affecting the features of other body parts.

[0123] S403. Based on the action decoder, parse the action features to obtain the video of the target object.

[0124] That is, the abstract action features are restored into specific video frames through the action decoder to obtain the video of the target object.

[0125] In the embodiments of the present disclosure, by combining appearance coding features, local features, initial noise, and semantic coding features, the denoising network can generate richer and more realistic action features; the appearance coding features provide the overall appearance information of the target object; the local features focus on the detailed parts; the initial noise introduces randomness to increase diversity; and the semantic coding features provide the semantic information of actions and postures. The fusion of such multi-source information makes the generated action features more comprehensive and accurate, thereby improving the quality of the generated video and ensuring the consistency and coherence between different frames of the generated video.

[0126] In the embodiments of the present disclosure, the human key points in the mixed pose sequence are aligned with the human key points of the target object.

[0127] The human key points are used to describe the posture of the reference object in the mixed pose sequence. The human key points of the target object refer to the posture of the target object extracted from the reference image. Aligning the human key points in the mixed pose sequence with the human key points of the target object means matching and adjusting the human key points of the reference object in the mixed pose sequence with the human key points of the target object to ensure their consistency in terms of body proportion, action amplitude, etc.

[0128] For example, if the target object is taller and thinner, and the human key points in the mixed pose sequence come from a shorter and fatter person, the alignment can adjust the positions of the key points so that the actions of the generated digital human are more in line with the body characteristics of the target object.

[0129] In the embodiments of the present disclosure, by aligning the human key points in the mixed pose sequence with the human key points of the target object, it can be ensured that the generated digital human is consistent with the target object in terms of body proportion, and the coherence and consistency of the actions can be ensured.

[0130] In summary, the overall process of the video generation method provided in the embodiments of the present disclosure is as Figure 9 shown:

[0131] S901, taking the reference image as one of the inputs, obtaining the appearance coding features through the VAE encoder, obtaining the semantic coding features through the CLIP encoder, and inputting the appearance coding features into the reference network to obtain intermediate features.

[0132] S902, extracting the human key points from the driving video to obtain the human skeleton parameter map, and after extracting the hand rendering map and performing fusion processing, obtaining the mixed pose sequence.

[0133] Among them, the higher the confidence level of a point, the higher the brightness of the key point. The hand rendering diagram is extracted by the hand pose estimation model HaMer, which uses the hand 3D Mesh parametric model MANO. Through hand pose estimation, it can finally be rendered into a three-dimensional grayscale diagram of the hand. Among them, in order to distinguish between the left and right hands, different colors are used to distinguish the left and right hands of the hand. After that, the human skeleton parameter diagram and the hand rendering diagram are fused to obtain a mixed pose sequence with the hand rendering diagram as the foreground diagram.

[0134] By adding a local attention mechanism, specifically, in S903, for the facial local area of the target object, a facial encoder is used to increase the facial local feature embedding. In S904, for the hand of the target object, the corresponding hand area is cropped and input into the hand encoder to extract the hand features of the target object.

[0135] Among them, the hand features of the target object extracted refer to the features that can represent the hand pose extracted from the hand local area of the target object. These features include but are not limited to the shape of the hand, the pose of the fingers, the movement trend of the hand, and the texture of the hand, etc.

[0136] Correspondingly, facial features refer to the features that can represent information such as the identity, expression, and pose of the target object extracted from the facial local area of the target object. These features cover the details of the facial features, such as the degree of opening and closing of the eyes, the shape and position of the mouth, as well as the overall contour of the face, skin texture, etc.

[0137] Based on the content described above, in the feature extraction process, the extraction of facial features and hand features is not limited to being the input of the denoising network. It can also be used as the input of the reference network like the semantic encoding features. For example, inputting the facial features into the reference network can enable the reference network to fully consider the details of the facial features when extracting features; inputting the hand features into the reference network can provide the reference network with detailed feature descriptions about the hand, so that the reference network can more accurately capture the feature details of the hand. Thus, inputting the facial features and hand features into the reference network can provide the reference network with richer information and diverse feature perspectives, thereby enhancing the overall feature expression ability and task processing ability of the model, making the generated video of the target object more natural and realistic.

[0138] S905, input the intermediate features, local features, initial noise, semantic encoding features, and pose sequence into the denoising network, and input the local features into the reference network. The denoising network processes them to obtain the action features of the target object.

[0139] S906, based on the action decoder, parse the action features to obtain the video of the target object.

[0140] In Figure 9Based on the network structure shown, in the embodiments of the present disclosure, the denoising network, the reference network, the modality encoding network for obtaining the semantic features of the target object, and the hand encoding network for extracting hand features are optimized through training with training samples.

[0141] In the embodiments of the present disclosure, by training and optimizing the denoising network, its network parameters can be optimized, and the robustness to noise and the ability to restore details can be improved, so that the generated video is more realistic. By optimizing the training of the reference network, the reference network can more effectively extract and utilize the features in the reference image, thereby improving the accuracy and coherence of the generation result. By training and optimizing the modality encoding network for obtaining the semantic features of the target object, the fusion strategy can be optimized, and the quality and guidance of the generated pose image can be improved. By optimizing the training of the hand encoding network for extracting hand features, the key features of the hand can be better captured and effectively encoded into feature vectors that can be used for generation.

[0142] In the embodiments of the present disclosure, the loss function required for training and optimizing the network includes at least one of the following losses:

[0143] (1) The pixel loss within the entire image range between the image frames in the generated video and the corresponding ground truth images;

[0144] That is, calculate the pixel difference between each frame image in the generated video and the corresponding real image within the entire image range. During implementation, this loss can be calculated based on the MSE (Mean Squared Error) loss.

[0145] (2) The first local pixel loss between the face in the image frame of the generated video and the face of the ground truth image;

[0146] That is, specifically used to measure the pixel difference between the face part of the image frame in the generated video and the face of the real image. During implementation, this pixel difference can still be calculated based on the MSE loss.

[0147] (3) The second local pixel loss between the hand in the image frame of the generated video and the hand of the ground truth image.

[0148] That is, specifically calculate the pixel difference for the hand regions in the generated image and the real image.

[0149] In the embodiments of the present disclosure, the pixel loss within the full image range between the image frames in the generated video and the corresponding ground truth images can ensure that the entire generated image is as close as possible to the real image as a whole, maintaining global visual consistency. The first local pixel loss between the face in the image frame of the generated video and the face in the ground truth image can help the model better capture and restore the subtle changes of the face, improving the quality of the face region in the generated video. The second local pixel loss between the hand in the image frame of the generated video and the hand in the ground truth image can optimize details such as the posture and joint positions of the hand, ensuring the naturalness and accuracy of the hand movement and making the hand region clearer and more realistic.

[0150] As Figure 9 shown, the training loss function will be constrained in the latent space. In addition to the normal MSE loss (global loss), a Mask Loss (local loss) will be additionally added, that is, for the hand and face regions, an MSE Loss constraint for an additional local region will be added to increase the supervision of the face and hand.

[0151] In summary, the video generation method provided by the embodiments of the present disclosure can improve the stability and effect of hand and body movements through a hybrid pose condition guidance, introduce face feature embeddings to enhance the face generation quality, and generate a video effect of stable and temporally consistent human body movements based on the temporal diffusion model.

[0152] Based on the same technical concept, the embodiments of the present disclosure also provide a video generation device 1000, as Figure 10 shown, including:

[0153] A first extraction module 1001 for extracting the reference features of the target object from the reference image;

[0154] A second extraction module 1002 for determining the hybrid pose sequence; multiple target parts in the hybrid pose sequence adopt corresponding description methods to enhance the pose characteristics of the corresponding parts;

[0155] A generation module 1003 for adding initial noise to the denoising network and performing a reverse denoising operation on the reference features and the hybrid pose sequence to generate a video of the target object.

[0156] In some embodiments, multiple target parts in the hybrid pose sequence include at least one of the following: face, hand, and the part of the body excluding the hand.

[0157] In some embodiments, the hand in the target part of the hybrid pose sequence adopts a three-dimensional gray-scale map.

[0158] In some embodiments, the left hand and the right hand of the hand in the target part in the mixed pose sequence are distinguished by different colors.

[0159] In some embodiments, the second extraction module includes:

[0160] An output unit for inputting a driving signal into a hand pose estimation model to obtain a hand rendering diagram of a reference object in the driving signal output by the hand pose estimation model; the hand rendering diagram is used to describe the three-dimensional mesh parameters of the hand of the reference object; and,

[0161] An extraction unit for extracting a hand mask diagram of the reference object and a human body skeleton parameter diagram of the reference object from the driving signal;

[0162] A fusion unit for fusing the human body skeleton parameter diagram and the hand rendering diagram based on the hand mask diagram to obtain a mixed pose sequence with the hand rendering diagram as the foreground diagram.

[0163] In some embodiments, for the face in the target part in the mixed pose sequence, key points of different parts of the face are described in different prominent ways.

[0164] In some embodiments, for the human body key points in the mixed pose sequence, the description value for describing the key points is positively correlated with the confidence of the key points.

[0165] In some embodiments, the reference features of the target object include at least one of the following: appearance coding feature, semantic coding feature, local coding feature;

[0166] The local coding feature includes a face feature and / or a hand feature.

[0167] In some embodiments, the first extraction module includes:

[0168] Extract the left hand local diagram and the right hand local diagram of the target object from the reference image;

[0169] Input the left hand local diagram and the right hand local diagram into a hand encoder to obtain hand features.

[0170] In some embodiments, extracting the face feature of the target object from the reference image includes:

[0171] A face acquisition unit for extracting a face local diagram of the target object from the reference image;

[0172] A face coding unit for inputting the face local diagram into a face encoder to obtain a face feature.

[0173] In some embodiments, the generation module includes:

[0174] A first processing unit for inputting the appearance encoding feature in the reference feature into a reference network to obtain an intermediate feature;

[0175] A second processing unit for inputting the intermediate feature, the local feature, the initial noise, and the semantic encoding feature into a denoising network, and processing by the denoising network to obtain the action feature of the target object;

[0176] A decoding unit for parsing the action feature based on an action decoder to obtain the video of the target object.

[0177] In some embodiments, the second processing unit is specifically configured to:

[0178] Input the intermediate feature, the local feature, the initial noise, and the semantic encoding feature into a denoising network, and the following operations are performed by multiple denoising modules in the denoising network:

[0179] Based on the spatial attention layer in the denoising module, process the intermediate feature, the noise feature corresponding to the initial noise, and the mixed pose sequence to obtain a first output feature;

[0180] Based on the first cross-attention layer in the denoising module, process the first output feature and the semantic encoding feature to obtain a second output feature;

[0181] Based on the second cross-attention layer in the denoising module, process the second output feature and the local feature to obtain a third output feature;

[0182] Based on the temporal attention layer in the denoising module, process the third output feature to obtain a fourth output feature;

[0183] Wherein, the fourth output feature output by the temporal attention layer of the last denoising module of the denoising network is the action feature of the target object.

[0184] In some embodiments, the second processing unit is specifically configured to, when the local feature includes a face feature and a hand feature, process the face feature based on the face attention sub-layer in the second cross-attention layer in the denoising module, and process the hand feature based on the hand attention sub-layer.

[0185] In some embodiments, the denoising network, the reference network, the modality encoding network for obtaining the semantic feature of the target object, and the hand encoding network for extracting the hand feature are optimized by training with training samples.

[0186] In some embodiments, the loss function required for training and optimizing the network includes at least one of the following losses:

[0187] The pixel loss within the full image range between the image frames in the generated video and the corresponding ground truth images;

[0188] Generate a first local pixel loss between the face in the image frame of the generated video and the face in the ground truth image;

[0189] Generate a second local pixel loss between the hand in the image frame of the generated video and the hand in the ground truth image.

[0190] In some embodiments, the human key points in the mixed pose sequence are aligned with the human key points of the target object.

[0191] For the specific functions and examples of each module and sub-module of the device according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.

[0192] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0193] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0194] Figure 11 FIG. shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0195] As Figure 11 shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1102 or the computer program loaded from the storage unit 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.

[0196] Multiple components in device 1100 are connected to I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0197] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the video generation method in any other suitable manner (e.g., by means of firmware).

[0198] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0199] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0200] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0201] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0202] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0203] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0204] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0205] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video generation method, comprising: Extracting reference features of the target object from the reference image; Determining a mixed posture sequence; wherein a plurality of target parts in the mixed posture sequence adopt corresponding description methods to enhance posture characteristics of corresponding parts; Initial noise is added to the denoising network, and a reverse denoising operation is performed on the reference features and the mixed pose sequence to generate a video of the target object.

2. The method according to claim 1, wherein: The multiple target parts in the mixed posture sequence include at least one of the following: face, hands, and parts of the body except hands.

3. The method according to claim 2, wherein: The hand in the target part in the mixed posture sequence adopts a three-dimensional gray model.

4. The method according to claim 2, wherein: The left hand and the right hand of the hand in the target part in the mixed gesture sequence are distinguished by using different colors.

5. The method according to any one of claims 2 to 4, wherein: The step of determining a mixed posture sequence comprises: Inputting the driving signal into a hand posture estimation model to obtain a hand rendering of a reference object in the driving signal output by the hand posture estimation model; the hand rendering is used to describe the three-dimensional mesh parameters of the hand of the reference object; and, Extracting a hand mask image of the reference object and a human skeleton parameter image of the reference object from the driving signal; Based on the hand mask image, the human skeleton parameter image and the hand rendering image are fused to obtain the mixed posture sequence with the hand rendering image as the foreground image.

6. The method according to any one of claims 2 to 5, wherein: For the face in the target part in the mixed posture sequence, key points of different parts of the face are described in different prominent ways.

7. The method according to any one of claims 1 to 6, wherein: For the human body key points in the mixed posture sequence, the description value used to describe the key points is positively correlated with the confidence of the key points.

8. The method according to any one of claims 1 to 7, wherein: The reference feature of the target object includes at least one of the following: an appearance coding feature, a semantic coding feature, and a local coding feature; The local coding features include facial features and / or hand features.

9. The method according to claim 8, wherein: Extracting the hand feature of the target object from the reference image includes: Extracting a left-hand partial image and a right-hand partial image of the target object from the reference image; The left hand partial image and the right hand partial image are input into a hand encoder to obtain the hand features.

10. The method according to claim 8, wherein: Extracting facial features of the target object from the reference image includes: Extracting a facial local map of the target object from the reference image; The facial local map is input into a facial encoder to obtain the facial features.

11. The method according to any one of claims 8 to 10, wherein: The adding of initial noise into the denoising network and performing a reverse denoising operation on the reference features and the mixed posture sequence to generate a video of the target object comprises: Inputting the appearance coding features in the reference features into the reference network to obtain intermediate features; Inputting the intermediate features, the local features, the initial noise and the semantic coding features into the denoising network, and processing the denoising network to obtain the action features of the target object; The motion feature is parsed based on a motion decoder to obtain a video of the target object.

12. The method according to claim 11, wherein: The step of inputting the intermediate features, the local features, the initial noise and the semantic coding features into the denoising network, and obtaining the action features of the target object by processing the denoising network, comprises: The intermediate features, the local features, the initial noise and the semantic coding features are input into the denoising network, and the multiple denoising modules in the denoising network perform the following operations: Processing the intermediate features, the noise features corresponding to the initial noise, and the mixed posture sequence based on the spatial attention layer in the denoising module to obtain a first output feature; Processing the first output feature and the semantic encoding feature based on a first cross attention layer in the denoising module to obtain a second output feature; Processing the second output feature and the local feature based on a second cross attention layer in the denoising module to obtain a third output feature; Processing the third output feature based on the temporal attention layer in the denoising module to obtain a fourth output feature; Among them, the fourth output feature of the temporal attention layer output of the last denoising module of the denoising network is the action feature of the target object.

13. The method according to claim 12, wherein: In the case where the local features include the facial features and the hand features, the denoising module processes the facial features based on the facial attention sublayer in the second cross-attention layer, and processes the hand features based on the hand attention sublayer.

14. According to the method of claim 12, the denoising network, the reference network, the modal encoding network for obtaining the semantic features of the target object, and the hand encoding network for extracting the hand features are obtained through training optimization of training samples.

15. The method according to claim 14, wherein: The loss function required for training and optimizing the network includes at least one of the following losses: Generate full-image pixel loss between image frames in the video and the corresponding ground-truth images; generating a first local pixel loss between a face in an image frame in a video and a face in the ground truth image; A second local pixel loss between a hand in an image frame in the video and a hand in the ground truth image is generated.

16. The method according to any one of claims 1 to 15, wherein: The human body key points in the mixed pose sequence are aligned with the human body key points of the target object.

17. A video generating device, comprising: A first extraction module is used to extract reference features of the target object from the reference image; The second extraction module is used to determine a mixed posture sequence; a plurality of target parts in the mixed posture sequence adopt corresponding description methods to enhance the posture characteristics of the corresponding parts; A generation module is used to add initial noise to the denoising network and perform a reverse denoising operation on the reference features and the mixed posture sequence to generate a video of the target object.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 16.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-16.

20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 16.