Method, apparatus, device, storage medium and program product for video generation

By generating 3D pose representations and pose sequence based on source images and audio of the target object, this technology solves the problem that facial movements cannot convey emotions and thoughts in existing technologies, and achieves personalized and natural video generation, which is suitable for applications such as virtual characters and intelligent assistants.

CN119653200BActive Publication Date: 2025-10-28BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411604048.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-10-28
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

In existing technologies for voice-driven image generation, facial movements cannot accurately convey emotions and thoughts, affecting video quality and lacking generalization ability.

Method used

The three-dimensional pose representation is determined based on the source image of the target object, a baseline pose action sequence is generated by combining the target audio, and then adjusted to the pose action sequence corresponding to the target object. Finally, the target video is generated to ensure that the action of the target object in the video can reflect the audio content and personalized appearance features.

Benefits of technology

The generated videos can reflect the personalized appearance features and natural movements of the target object, enhancing the expressiveness and generalizability of the videos, and are suitable for scenarios such as virtual character generation and intelligent assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653200B_ABST
    Figure CN119653200B_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a method, apparatus, device, storage medium, and program product for video generation. The method includes: determining a three-dimensional pose representation of the target object based on a source image of the target object; determining a reference pose action sequence corresponding to the target audio based on target audio; adjusting the reference pose action sequence based on the three-dimensional pose representation to obtain a target pose action sequence corresponding to the target object; and generating a target video of the target object based on the source image of the target object and the target pose action sequence, the target video representing the target object performing pose actions corresponding to the target pose action sequence while speaking the target audio. This improves the quality of the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, devices, storage media, and program products for video generation. Background Technology

[0002] With the continuous development of voice-driven image generation technology, it has shown broad potential in applications such as virtual character generation, video conferencing, and intelligent assistants. Traditional technologies typically focus on image generation related to facial movements. However, facial movements alone often fail to convey accurate emotions and thoughts, thus affecting video quality. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for video generation is provided. The method may include: determining a three-dimensional pose representation of the target object based on a source image of the target object; determining a reference pose sequence corresponding to the target audio based on target audio; adjusting the reference pose sequence based on the three-dimensional pose representation to obtain a target pose sequence corresponding to the target object; and generating a target video of the target object based on the source image of the target object and the target pose sequence, the target video representing the target object performing pose actions corresponding to the target pose sequence while speaking the target audio.

[0004] In a second aspect of this disclosure, an apparatus for video generation is provided. The apparatus may include: a three-dimensional pose representation determination module configured to determine a three-dimensional pose representation of a target object based on a source image of the target object; a reference pose action sequence determination module configured to determine a reference pose action sequence corresponding to a target audio based on target audio; a target pose action sequence determination module configured to adjust the reference pose action sequence based on the three-dimensional pose representation to obtain a target pose action sequence corresponding to the target object; and a video generation module configured to generate a target video of the target object based on the source image of the target object and the target pose action sequence, wherein the target video represents the target object performing pose actions corresponding to the target pose action sequence while speaking the target audio.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method of the first aspect.

[0008] It should be understood that the description in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A flowchart of a method for video generation according to some embodiments of the present disclosure is shown;

[0012] Figure 3A A schematic diagram of a method for video generation according to some embodiments of the present disclosure is shown;

[0013] Figure 3B An example schematic diagram of the rendering model according to some embodiments of the present disclosure is shown;

[0014] Figure 4A Example diagrams of video coding models according to some embodiments of the present disclosure are shown;

[0015] Figure 4B A schematic diagram illustrating the pre-training principle of a video coding model according to some embodiments of the present disclosure is shown;

[0016] Figure 5A A schematic diagram illustrating the training of a rendering model according to some embodiments of the present disclosure is shown;

[0017] Figure 5B A schematic diagram illustrating the training of a weight determination model according to some embodiments of the present disclosure is shown;

[0018] Figure 6 A schematic structural block diagram of an apparatus for video generation according to some embodiments of the present disclosure is shown; and

[0019] Figure 7 A block diagram of an electronic device that can implement one or more embodiments of the present disclosure is shown. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0022] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.

[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, environment 100 may include electronic device 110.

[0030] In this example environment 100, electronic device 110 can acquire input information 102. Input information 102 includes at least target audio 114 and a source image 113 of the target object. As an example, the target object can include humans, animals, cartoon characters, and virtual characters, etc. Electronic device 110 can generate a three-dimensional pose representation corresponding to the target object using a target model 115 based on the source image 113 of the target object. Furthermore, electronic device 110 can also generate a reference pose sequence 307 synchronized with the target audio 114 using the target model 115 based on the target audio 114. The reference pose sequence 307 corresponds to the target audio and is used to represent gestures, postures, etc., corresponding to the target audio. Based on the reference pose sequence 307 and the three-dimensional pose representation corresponding to the target object, a pose sequence corresponding to the target object can be obtained. Mapping the source image 113 of the target object to the pose sequence corresponding to the target object yields the target video 104 of the target object. This target video 104 can synchronously exhibit the pose movements corresponding to the reference pose sequence 307 with the target audio 114. Figure 1 The example shown is only one target model 115. In reality, multiple different target models 115 may be used in collaboration to complete the video generation.

[0031] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server-side equipment (not shown) can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server-side equipment can, for example, provide background services for the applications of electronic device 110.

[0032] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0033] In dialogues or speeches, speakers often use communal gestures to express their thoughts and feelings. These gestures help bridge communication gaps, enhance intimacy, and increase the credibility of speech. Therefore, in speech-driven image generation technology, endowing virtual objects with communal gestures is significant for human-computer interaction and also helps enhance the realism of virtual reality environments. Consequently, some related technologies explore the generation of communal gestures by generating gesture sequences aligned with speech. However, while these generated gesture sequences can produce actions, they fail to reflect the speaker's personalized appearance, thus limiting their potential applications. Some studies have overcome this limitation by synthesizing gesture sequences and videos of specific speakers, but these studies are limited to specific speakers and lack generalization.

[0034] In embodiments of this disclosure, an improved scheme for video generation is proposed. In this scheme, an electronic device determines a three-dimensional pose representation of a target object based on a source image of the target object. Based on target audio, a reference pose action sequence corresponding to the target audio is determined. The reference pose action sequence is adjusted based on the three-dimensional pose representation to obtain a pose action sequence corresponding to the target object. Based on the source image of the target object and the target pose action sequence, a target video of the target object is generated, the video representing the target object performing pose actions corresponding to the target pose action sequence while speaking the target audio.

[0035] Through the above process, the source image is used to determine the 3D pose representation of the target object, enabling the generated video to reflect the personalized appearance characteristics of the target object and avoiding the limitations of generating only a pose sequence without appearance information. Secondly, by generating a synchronized baseline pose sequence based on the target audio, it is ensured that the target object's actions in the video reflect the thoughts and emotions in the target audio, enhancing the naturalness and expressiveness of the generated video. Furthermore, the improved scheme demonstrates good generalization; it can generate corresponding personalized videos by providing source images of different target objects, adapting to the generation needs of various speakers. Therefore, this scheme shows broader potential and application value in scenarios such as virtual character generation and intelligent assistants, effectively improving the quality and applicability of video generation.

[0036] Figure 2 An example flow diagram of a method 200 for video generation according to some embodiments of the present disclosure is shown. For ease of discussion, reference will be made to... Figure 1 The process 200 is described in the context of the environment. In environment 100, video generation can be completed by electronic device 110, but some of the operations can be performed by requesting a server device (not shown) (such as determining the 3D pose representation, determining the reference pose action sequence, video generation, or the training process of part of the model can be implemented at the server device).

[0037] Electronic device 110 can acquire input information for generating video as a trigger condition for video generation. The input information may include at least a source image 113 containing the target object and target audio 114. As an example, the target object may include humans, animals, cartoon characters, and virtual characters, etc. In box 201, electronic device 110 determines the three-dimensional pose representation of the target object based on the source image of the target object.

[0038] Figure 3A A schematic diagram 300A illustrating a method 300 for video generation according to some embodiments of the present disclosure is shown. The source image 113 containing the target object can be a still image or a video image obtained by extracting frames from a video containing the target object.

[0039] The source image 113 containing the target object can be used to determine the three-dimensional pose representation 304 of the target object. For example, the source image 113 containing the target object can be processed using a three-dimensional pose generation model 320 to obtain the three-dimensional pose representation 304 of the target object. The three-dimensional pose generation model 320 can determine the three-dimensional pose representation 304 of the target object by detecting and locating the skeletal keypoints of the target object in the source image 113 and mapping the skeletal keypoints into three-dimensional space. As an example, the three-dimensional pose generation model 320 can correspond to an Expressive Human Pose and Shape Predictor (EHPS).

[0040] In box 202, the electronic device 110 can first extract features from the target audio 114 to obtain a feature representation corresponding to the audio. Then, this feature representation is processed to generate a baseline pose sequence 307.

[0041] Each reference posture in the reference posture sequence 307 is a standardized gesture, corresponding to the content of the target audio 114. A reference posture is a general way of representing posture; it defines the direction of movement and range of motion of the target object in three-dimensional space, without being limited by specific body shape, proportions, or other characteristics.

[0042] The versatility of this posture representation method allows it to be applied to different target objects and can be adjusted in subsequent steps to adapt to the specific posture characteristics of different objects. As an example, each reference posture action in the reference posture action sequence 307 can be in three-dimensional or two-dimensional form. In this application example, the reference posture actions in the reference posture action sequence 307 are illustrated and described in three-dimensional form.

[0043] In box 203, the electronic device 110 adjusts the reference posture action sequence 307 based on the three-dimensional posture representation 304 to obtain the target posture action sequence 308 corresponding to the target object. As previously mentioned, the reference posture actions in the reference posture action sequence 307 are not limited by specific body shape, proportions, or other features. Therefore, in order to better represent the target object, the reference posture action sequence 307 can be adjusted based on the three-dimensional posture representation 304.

[0044] The 3D pose representation 304 of the target object can indicate the corresponding skeletal proportions and body features of the target object. Using the 3D pose representation 304 of the target object, each reference pose action in the reference pose action sequence 307 is adapted and adjusted to generate a target pose action sequence 308 corresponding to the skeletal proportions and body features of the target object. As an example, the adaptation adjustment may include adjusting the lengths and angles between keypoints in the reference pose actions.

[0045] This adjustment ensures that the final generated target posture and movement sequence 308 not only expresses the posture and movement corresponding to the speech, but also matches the body shape and proportions of the target object, ensuring that the generated video has personalized appearance features and natural movement performance.

[0046] In box 204, electronic device 110 generates target video 350 of target object based on source image 113 of target object and target posture action sequence 308. Target video 350 represents the target object performing a posture action corresponding to target posture action sequence 308 while speaking target audio 114.

[0047] Electronic device 110 can use source image 113 of the target object to determine a visual feature representation of the target object. The visual feature representation can indicate the appearance characteristics of the target object, and may include information such as facial features, body shape, and clothing. These appearance features will be used to ensure that the generated target video retains the personalized visual effect of the target object.

[0048] Electronic device 110 combines the target posture and action sequence 308 with the visual feature representation of the target object, and generates dynamic video frames corresponding to the posture and action through a rendering process. In the generated target video 350, the target object not only speaks synchronously according to the target audio 114, but also exhibits natural and personalized performance accompanied by gestures and actions associated with the audio content of the target audio 114.

[0049] The aforementioned process tightly integrates the target subject's facial expressions, body movements, and speech, resulting in a more vivid and realistic effect in the target video, meeting the needs of applications such as virtual character generation and intelligent assistants. Furthermore, the process demonstrates good generalization; it can generate corresponding personalized target videos by providing source images of different target subjects, adapting to the generation needs of various speakers. Therefore, this solution shows broader potential and application value in scenarios such as virtual character generation and intelligent assistants, effectively improving the quality and applicability of video generation.

[0050] In determining the target pose sequence 308 corresponding to the target audio 114, the electronic device 110 can first determine the feature representation of the target audio 114. Based on the feature representation of the target audio 114, a reference pose sequence 307 corresponding to the target audio 114 is determined using a motion generation model 311. In some embodiments, the motion generation model 311 can be implemented based on a diffusion model, and therefore can also be called a diffusion motion generation model. Of course, other types of model structures can also be selected to construct the motion generation model, as long as the model can construct pose actions from audio features.

[0051] Electronic device 110 can process target audio 114 using audio encoder 310 to obtain feature representation of target audio 114. Feature representation of target audio 114 can indicate semantic information, speech rhythm and emotional changes corresponding to target audio 114, and thus can be used as reference information for generating reference gesture sequence 307.

[0052] By processing the feature representation of the target audio 114 using the action generation model 311, a baseline posture action sequence 307 corresponding to the target audio 114 can be generated. The baseline posture action sequence 307 is a standardized posture action representation. Since the baseline posture action sequence 307 is generated based on the feature representation of the target audio 114, it can reflect the posture actions related to the content expressed by the target audio 114. The baseline posture actions in the baseline posture action sequence 307 are standardized and do not yet reflect the personalized characteristics of the target object, such as the target object's skeletal proportions and body shape. The baseline posture action sequence 307 may include multiple baseline posture actions, which correspond to various audio units in the target audio (e.g., audio frames as audio units). Each baseline posture action corresponds to a specific segment in the audio, reflecting the general limb movements that should be performed in that audio segment.

[0053] Through the above process, the audio coding model can capture the rhythm and emotional features of the target audio 114, thereby enabling the reference posture action sequence 307 to be highly consistent with the content and emotional expression of the target audio 114.

[0054] As previously mentioned, the baseline posture sequence 307 is standardized. To obtain the posture action corresponding to a specific target object, the baseline posture sequence 307 needs to be adjusted using the target object's three-dimensional posture representation 304 to obtain the target posture sequence corresponding to the target object, thus reflecting the visual characteristics of the target object's posture. Based on this, the electronic device 110 can determine the three-dimensional spatial coordinates corresponding to each key point in the target object's three-dimensional posture representation 304. Based on the three-dimensional spatial coordinates corresponding to each key point, the angle and length representations between each key point are determined. Using the angle and length representations between each key point, the baseline posture actions included in the baseline posture sequence 307 are adjusted to obtain the target posture sequence 308.

[0055] Each keypoint in the 3D pose representation 304 of the target object corresponds to a joint in the target object's skeletal system, such as the joints in the shoulder, elbow, wrist, and fingers. Based on the 3D pose representation 304 of the target object, the electronic device 110 can determine the coordinate positions of each keypoint in 3D space. Based on the coordinate positions of these keypoints in 3D space, the electronic device 110 can calculate the angle and length representations between each keypoint. These angle and length representations reflect the skeletal proportions and body structure of the target object, such as the presentation state of the fingers, the extension angle of the arms, the extension state of the fingers, and the length of the legs. With this information, the spatial relationships between different parts of the target object can be accurately determined.

[0056] Based on the angle and length representations between key points in the 3D pose representation 304 of the target object, each reference pose movement in the reference pose sequence 307 can be adjusted. Specifically, each reference pose movement in the reference pose sequence 307 will be scaled and adjusted (e.g., length adjustment, body shape adjustment) to obtain the target pose sequence 308 that conforms to the skeletal proportions and body features of the target object. Each pose movement in the target pose sequence 308 will be consistent with the personalized visual features of the target object, thereby ensuring that in the generated target video 350, the target object's pose movements are synchronized with the audio content of the target audio 114 and visually matched with the target object's body features.

[0057] As previously mentioned, each reference pose in the reference pose sequence 307 can be in three-dimensional form. After adjusting each reference pose in the reference pose sequence 307 using the three-dimensional pose representation 304, a three-dimensional pose sequence can be obtained. Based on this, the three-dimensional pose sequence can be converted into a two-dimensional representation (two-dimensional target pose sequence 308) using camera parameters to adapt to the two-dimensional characteristics of the final video generation. Camera parameters are used to determine the viewpoint and projection method in three-dimensional space, thereby ensuring that the generated two-dimensional target pose sequence 308 can accurately reflect the action and characteristics of the target object in three-dimensional space.

[0058] Through this adjustment method, the generated target video 350 can reflect the real appearance and natural movements of the target object, making the target object's performance in the target video 350 more realistic and believable, and enhancing the video's personalization and visual effects.

[0059] After adjusting the reference pose sequence 307 based on the 3D pose representation 304 to obtain the target pose sequence 308 corresponding to the target object, the target pose sequence 308 can be rendered based on the source image 113 of the target object to generate the target video 350 of the target object. As an example, the electronic device 110 can determine the visual feature representation of the target object based on the source image 113 of the target object. Based on the visual feature representation of the target object, each pose action in the target pose sequence 308 is rendered to obtain the target video 350 of the target object.

[0060] The electronic device 110 can extract visual feature representations of the target object based on the source image 113 containing the target object. These visual feature representations may include, but are not limited to, information such as the target object's facial features, body appearance, clothing color, and texture. With the help of visual feature representations, it can be ensured that the subsequently generated target video 350 visually accurately reflects the personalized characteristics of the target object.

[0061] Electronic device 110 renders each pose / action in target pose / action sequence 308 based on the visual feature representation of the target object. During rendering, electronic device 110 combines the visual features of the target object with each pose / action in target pose / action sequence 308 to generate corresponding image frames, ensuring that the target object can exhibit actions corresponding to target pose / action sequence 308 in each image frame. In this way, the target object presented in target video 350 not only moves according to specified pose / actions but also retains its appearance features and dynamic details. To enhance the visual effect of the video, techniques such as lighting processing and texture detail enhancement can also be applied during rendering to improve the realism and expressiveness of target video 350. Finally, electronic device 110 can combine these rendered image frames into a continuous target video 350, allowing the target object to exhibit gestures and actions corresponding to target audio 114 in target video 350.

[0062] In some embodiments of this disclosure, the source images of the target object include multiple source images 113, each with a different presentation angle of the target object. In this case, taking a given pose action in the target pose action sequence 308 as an example, the generation process of the target video 350 is described. For a given pose action in the target pose action sequence 308, the electronic device 110 determines weight parameters corresponding to each of the multiple source images based on the action differences and the given pose action, wherein the action differences are determined separately based on the given pose action and the source pose actions corresponding to the multiple source images 113. Based on the multiple source images 113 and their respective weight parameters, a visual feature representation of the target object is determined. Based on the visual feature representation of the target object, the given pose action is rendered to obtain a video segment corresponding to the given pose action, the video segment being part of the target video of the target object.

[0063] Figure 3B A schematic diagram 300B illustrating the rendering process of a rendering model 330 according to some embodiments of the present disclosure is shown. For multiple source images 113, the source pose action corresponding to the target object in each source image 113 can be determined first, and the pose action feature representation 113-a of the source pose action can be obtained. Taking four source images as an example, they correspond to the first to the fourth source images respectively. As an example, the multiple source images can respectively present the front view, back view, and side view (left side and right side) of the target object.

[0064] For a given posture action 308-1 in the target posture action sequence 308, the electronic device 110 can extract the action difference feature representation 362 corresponding to each source image from the given posture action and the source posture actions corresponding to multiple source images. The action difference feature representation indicates the action difference. Based on the action difference feature representation corresponding to each source image, the weight parameters corresponding to each source image are determined.

[0065] For example, a pose action feature representation 308-a can be determined for a given pose action 308-1, and a pose action feature representation 113-a can be determined for the source pose action. As an example, the pose action feature representation can indicate information such as the coordinates of key points. Taking the first source image as an example, the action difference determination model 360 can compare the action differences between the given pose action 308-1 and the source pose action in the first source image based on the pose action feature representation. Based on the action differences, the action difference feature representation 362 of the first source image can be determined. The action differences can indicate information such as differences in pose angles and joint positions. Similarly, the action difference determination model 360 can obtain the action difference feature representations 362 for other source images.

[0066] Based on the action difference feature representation 362 of each source image, the electronic device 110 uses a weight determination model 364 to evaluate the importance of each source image to determine the weight of each source image in generating a visual feature representation of the target object. For example, if the action difference between a given pose action 308-1 and the source pose action in the first source image is large (e.g., the given pose action requires the left side of the face, but the source pose action in the first source image corresponds to the right side of the face), then the weight determination model 364 can determine that the first source image reflects less of the given pose action, and therefore the weight parameter of the first source image is low. As another example, if the action difference between a given pose action 308-1 and the source pose action in the second source image is small, then the weight determination model 364 can determine that the weight parameter of the second source image is high.

[0067] Weight parameters indicate the relative importance of each source image in generating the visual feature representation of the target object. A higher weight parameter indicates that the source image better reflects the target object's specific pose or viewpoint. Therefore, the degree of reference can be understood as a measure of the contribution of each source image to the final synthesized visual features. Specifically, the degree of reference can be further refined to the pixel level. The information contained in different pixels of each source image has different importance in generating the visual features of the target object. Therefore, the electronic device 110 generates pixel-level weighted representations based on the importance of each pixel in multiple source images. These pixel-level weights determine the contribution of a certain pixel in a source image to the visual features of the target object during the synthesis of the visual feature representation.

[0068] For example, when the target object needs to turn to the left, the left side information in a source image may be more important than other parts, and the pixel weight of the left side region will be higher. Conversely, the back of the target object in the same source image will have a lower pixel weight because this part is not the focus of the pose movement. Under such pixel-level weight control, the image decoder 369 can ensure that it can accurately select the feature information that best reflects the specific pose and perspective of the target object during the determination of the target object's visual features.

[0069] By processing each source image 113 using image encoder 366, the feature encoding representation of source image 113 can be obtained. By processing a given posture action 308-1 using posture encoder 368, the action feature encoding representation of the given posture action 308-1 can be obtained.

[0070] The feature encoding representation of source image 113, the action feature encoding representation of given pose 308-1, and the weight of each source image are used as input information for image decoder 369. Image decoder 369 can obtain video segment 350-1 corresponding to the given pose 308-1. Similarly, when rendering poses, electronic device 110 renders each pose in the target pose sequence 308 frame by frame. Each pose corresponds to a video segment. By rendering each pose sequentially, the complete target video 350 is finally generated.

[0071] Through the above process, based on each source image and its corresponding weight parameters, the visual feature representation of the target object is determined to ensure that the visual features can fully combine the source image information from different angles, thereby making up for the deficiencies brought about by a single image perspective and obtaining a more complete visual feature representation of the target object.

[0072] Combination Figure 3A As shown, generating the baseline pose action sequence 307 involves an audio coding model 310 and an action generation model 311. The audio coding model 310 can be trained through pre-training.

[0073] After pre-training, the audio coding model 310 and the action generation model 311 are jointly trained through fine-tuning. The fine-tuning process includes processing (first or second) video samples, extracting audio samples and baseline action sequence samples from the video samples. The audio coding model 310 processes the audio samples to extract audio features. The action generation model 311 generates corresponding action sequence prediction results based on the audio features. Then, the action sequence prediction results are compared with the baseline action sequence samples, and a loss function is calculated. This loss function is used to jointly fine-tune the parameters of the audio coding model 310 and the action generation model 311.

[0074] The following section first introduces the tuning process of action generation model 311. Action generation model 311 is used to generate action sequence x given an audio feature sequence a. Action generation model 311 generates baseline gesture actions through progressive denoising. This generation process starts with random noise and gradually guides the audio features into gesture actions that reflect the target audio content by progressively applying a diffusion step. The training objective of action generation model 311 is to minimize the mean square error in the noise space. The training objective of action generation model 311 can be expressed as follows:

[0075]

[0076] ∈ can represent noise, conforming to the distribution of random noise, i.e., ∈ ~ N(0, I). θ This can represent a denoising network. t can represent a random time step, x... tLet represent the noisy gesture sequence at time step t. Let c represent the input obtained by processing the audio sequence using the audio coding model 310. The training objective is to minimize the noise ∈ such that the generated gesture sequence is consistent with the distribution of the real gesture sequence.

[0077] The pre-training process of the audio coding model 310 will be described in detail below. The electronic device 110 extracts audio samples and baseline action sequence samples from a first video sample containing a first object sample. Multiple positive samples are constructed, each including a first audio segment and a first baseline action sequence. The first audio segment is segmented from the audio sample, and the first baseline action sequence is segmented from the baseline action sequence sample. The first audio segment and the first baseline action sequence are temporally aligned. Multiple negative samples are constructed, each including a second audio segment and a second baseline action sequence. The second audio segment is segmented from the audio sample, and the second baseline action sequence is segmented from the baseline action sequence sample. The second audio segment and the second baseline action sequence are temporally misaligned. Using multiple positive and negative samples, the audio coding model is pre-trained through contrastive learning.

[0078] Electronic device 110 can extract audio samples and reference action sequence samples from a first video sample containing a first object sample. These audio samples and reference action sequence samples reflect the synchronization relationship between the audio content and gestures of the first object sample during speech.

[0079] Figure 4A A schematic diagram 400A illustrating the principle of an audio coding model 310 according to some embodiments of the present disclosure is shown. The audio coding model 310 includes a convolutional layer 405 for processing audio samples 401 to obtain first audio features. Furthermore, the audio coding model 310 also includes an encoding layer 402 and a feature extraction layer 404. The encoding layer 402 processes the audio samples 401 to obtain an encoded representation of the audio samples 401. The feature extraction layer 404 processes the encoded representation of the audio samples 401 to obtain second audio features. By fusing the first and second audio features, an audio feature representation of the audio sample 401 can be obtained.

[0080] In some embodiments, the encoding layer 402 may employ a Hubert feature encoder, which encodes the audio sample 401 into a series of semantic tags. For example, the audio sample 401, after processing by the Hubert feature encoder, can be represented as a series of numbers such as "32 67 354...". In some embodiments, the feature extraction layer 404 may be a context encoder (Context Transformer) that performs further feature extraction on the output of the encoding layer 402.

[0081] Figure 4B A schematic diagram 400B illustrating the pre-training principle of an audio coding model 310 according to some embodiments of the present disclosure is shown. The pre-training process of the audio coding model 310 may include training the coding layer 402 and the feature extraction layer 404. Before training, the electronic device 110 may first construct positive and negative samples for training. For example, the audio sample 401 and the baseline action sequence sample 403 may first be segmented into multiple audio segments a1, a2, ..., a n and corresponding multiple baseline action sequence fragments x1, x2, ..., x n Audio clip a1 and the reference action sequence clip x1 are of the same length and aligned on the timeline. Similarly, audio clip a2 and the reference action sequence clip a2 are of the same length and aligned on the timeline, and so on down to audio clip a1. n and benchmark action sequence fragment x n They are of equal length and aligned on the timeline.

[0082] For each audio segment, a corresponding audio feature representation (T1, T2, ..., T) can be obtained. n Similarly, by processing each baseline motion sequence segment using the motion encoder 407, the corresponding motion feature representation (S1, S2, ..., S...) can be obtained. n The determination of audio feature representation and action feature representation can be expressed as follows:

[0083]

[0084] LN can represent layer normalization, L speech and L gesture These can be used to represent linear transformations. speech and f speech These can represent the processing steps of the semantic encoder (coding layer 402 and feature extraction layer 404) and the action encoder 407, respectively. 'a' can represent an audio sample, where a = {a1, a2, ..., a...}. n x can represent a baseline action sequence sample x = {x1, x2, ..., x}. n}

[0085] Multiple positive samples can be audio feature representations and motion feature representations aligned along the time axis. For example, in similarity matrix 410, a feature pair consisting of the audio feature representation T1 corresponding to the first audio segment and the motion feature representation S1 corresponding to the first baseline motion sequence segment can correspond to the first positive sample. Similarly, a feature pair consisting of the audio feature representation T2 corresponding to the second audio segment and the motion feature representation S2 corresponding to the second baseline motion sequence segment can correspond to the second positive sample. And so on, the audio feature representation T of the nth audio segment... nThe motion feature representation S corresponding to the nth baseline motion sequence segment n The feature pairs formed can correspond to the nth positive sample.

[0086] Multiple negative samples can be audio feature representations and motion feature representations that are not aligned on the time axis. For example, in similarity matrix 410, all feature pairs other than positive samples can be used as negative samples. For instance, the audio feature representation T1 corresponding to the first audio segment and the motion feature representation S2 corresponding to the second baseline motion sequence segment can form a negative sample.

[0087] During the pre-training phase, a contrastive learning approach can be used to train the audio coding model 310, ensuring that it learns the close correspondence between audio and action. The training objective of the audio coding model 310 can be expressed as follows:

[0088] L CL =0.5×(l speech (τ·C ges )+l gesture (τ·C ges (4)

[0089] τ can represent the parameter that needs to be adjusted. C ges This can represent a similarity matrix of 410. speech This can represent the contrast loss on audio features calculated based on the similarity between audio feature representations and action feature representations. gesture This can represent the contrastive loss for action features calculated based on the similarity between audio feature representations and action feature representations. The goal of contrastive learning is to increase the similarity values ​​in the similarity matrix for identical (aligned) audio feature representations and action feature representations (positive samples) on the time axis, while decreasing the similarity between dissimilar (misaligned) audio feature representations and action feature representations (negative samples) on the time axis.

[0090] The pre-trained audio coding model 310 can more accurately capture the correlation between audio and gestures, making the generation of baseline gesture sequences more natural and synchronized with the audio content. At the same time, this contrastive learning-based training method enables the audio coding model 310 to have stronger generalization ability when faced with diverse speech inputs, thus generating expressive gestures that match the speech in various scenarios.

[0091] The generation of the target video also involves a rendering model 330. Taking the training process performed via electronic device 110 as an example, the training process of the rendering model 330 is described. Electronic device 110 selects a source image sample 511 and a target image sample 512 containing the second object sample from a second video sample containing the second object sample. Based on the source image sample, a visual feature representation of the second object sample is determined. Based on the second video sample, pose / action sequence samples associated with the second object sample are determined. Based on the visual feature representation, a given action sample from the pose / action sequence samples is rendered using the rendering model to obtain a predicted image of the second object sample, where the given action sample is extracted from the target image sample. Based on the difference between the predicted image and the target image sample, the parameters of the rendering model are adjusted.

[0092] Figure 5A A schematic diagram 500A illustrating the training principle of a rendering model 330 according to some embodiments of the present disclosure is shown. First, the training process of the rendering model 330 is described using a single source image sample as an example.

[0093] Electronic device 110 selects a source image sample 511 and a target image sample 512 containing the second object sample from a second video sample containing the second object sample. The source image sample 511 is used to determine the visual feature representation of the second object sample.

[0094] Furthermore, the electronic device 110 can also determine pose and motion sequence samples related to the second object sample from the second video sample containing the second object sample. A given motion sample 506 in the pose and motion sequence sample is extracted from the target image sample 512.

[0095] The target image sample 512 corresponds to the given action sample 506 in the pose action sequence sample, that is, the given action sample 506 is extracted from the target image sample 512.

[0096] By processing the source image sample 511 using the image encoder 366, a feature encoding representation of the source image sample 511 can be obtained. By processing the given action sample 506 using the pose encoder 368, a motion feature encoding representation of the given action sample 506 can be obtained. Using the feature encoding representations of the source image sample 511 and the motion feature encoding representations of the given action sample 506 as input information to the image decoder 369, the image decoder 369 can obtain the prediction image 520. Based on the difference between the prediction image 520 and the target image sample 512, the parameters of the rendering model 330 can be adjusted.

[0097] The training objective for rendering model 330 can be represented as follows:

[0098]

[0099] i can represent the number of layers in the loss-aware model, V i Let represent the output feature map of the i-th layer of the loss-aware model. Let j represent the number of downsampling iterations. j This can be interpreted as the target image sample 512 undergoing j downsampling operations before being input into the loss-aware model. This can be represented as the predicted image 520 undergoing j downsampling iterations before being input into the loss-aware model. The loss-aware model is used to calculate the feature differences between the generated predicted image 520 and the target image sample 512 at different levels. Based on the feature differences between the predicted image 520 and the target image sample 512 at different levels, the electronic device 110 can continuously adjust the parameters of the rendering model 330 until the training cutoff conditions are met (such as the feature differences meeting the requirements, or the training time reaching a given duration, etc.).

[0100] It should be noted that, since the current embodiment uses a single source image sample to train the rendering model 330 as an example, therefore Figure 5A The weight determination model 503 in the model is represented by a dashed line (i.e., the training process of a single source image sample on the rendering model 330 may not involve the weight determination model 503) and does not involve the training process related to the weight determination model 503.

[0101] To improve the training performance of the rendering model 330, multiple source image samples 511 containing object samples can be used, and the object samples in each source image sample 511 are presented from different angles. Based on this, combined with Figure 5A The weight determination model 364 in the text introduces the training process of the rendering model 330.

[0102] For a given source image sample in each source image sample 511, the difference determination model 360 processes the target pose action feature representation 504 and the source pose action feature representation 502 to obtain the action difference sample feature representation 562 corresponding to each source image. The target pose action feature representation 504 is determined based on the target image sample 512. The source pose action feature representation 502 is determined based on the source image sample 511. Based on the action difference sample feature representation 562 corresponding to each source image sample 511, the weight determination model 364 determines the weight parameters corresponding to each source image sample 511.

[0103] Figure 5B A schematic diagram 500B illustrating the principle of a weight determination model 364 according to some embodiments of the present disclosure is shown. Combined with... Figure 5A and Figure 5BAs shown, after performing a linear transformation 531 on the source pose action feature representation 502 and the target pose action feature representation 504, the weight determination model 364 inputs the linear transformation result as a query (Q), key (K), and value (V) into the cross-attention calculation model 532 to determine the similarity between the source pose action feature representation 502 and the target pose action feature representation 504, thereby determining the weight of each source image sample 511 (pixel-by-pixel). The calculation result can be called the keypoint fusion feature representation 533. The above process fuses the feature information of the source pose action feature representation 502 and the target pose action feature representation 504. For example, the importance of each source image sample 511 in each region can be determined based on the distance and relative position between corresponding keypoints. If a keypoint in a source image sample 511 is very close to the corresponding keypoint in the target image sample 512, it indicates that the pose matching degree of the source image sample in that region is high, and the weight of the corresponding pixel position should also be high.

[0104] On the other hand, the weight determination model 364 performs a linear transformation 531 on the action difference sample feature representation 562. The result of the linear transformation, along with the keypoint fusion feature representation 533, is used as the query (Q), key (K), and value (V), and is then input into the cross-attention calculation model 532 to obtain the visual-keypoint fusion feature 534. In this process, the action difference sample feature representation 562 and the keypoint fusion feature representation 533 interact, further refining the weight parameters corresponding to each source image sample.

[0105] By combining the visual-keypoint fusion features 534 with the target pose action feature representation 504 to complete the cross-calculation 535, a cross-attention map is obtained, which indicates the parameters of each source image sample 511. Based on each source image sample and its corresponding weight parameters, the visual feature representation of the object sample is determined.

[0106] To improve the generalization effect of model training, masking can be performed on each source image sample 511 before determining the weights. The masking process for the source image samples is as follows:

[0107] f i =w i (I i )⊙M i (6)

[0108] w i This can represent the i-th source image sample (I) determined by the action difference determination model 501. i The feature representation of the action difference samples and the given action sample. M i It can represent an occlusion mask.

[0109] For each masked source image sample 511, an aggregation process can also be included. The aggregation process is represented as follows:

[0110] S=∑ i f i ⊙W i (7)

[0111] W i This can represent the weight of the i-th source image sample. By aggregation, we can obtain the visual feature representation of the (masked) object sample.

[0112] By introducing a weighted model 364, the rendering model 330 can better utilize information from multiple source image samples during training. Compared to the case where only a single source image can be relied upon without introducing the weighted model 364, the rendering model 330 can perform pixel-by-pixel weighted fusion of pose and action features from different source image samples after introducing the weighted model 364, thereby selecting the most suitable features according to the needs of action differences. This pixel-by-pixel weighting method ensures that the rendering model 330 can comprehensively extract effective visual information from multiple angles and poses when generating prediction images, improving the performance of the generated images in terms of detail and overall consistency, and ultimately significantly improving the training effect of the rendering model 330 and the quality and naturalness of the generated images.

[0113] In some embodiments of this disclosure, the reference gesture sequence corresponding to the target audio is a gesture sequence including hand gestures. Compared to full-body gestures, hand gestures generally express the relevance to the speech content more directly, especially during communication, where gestures can more effectively convey specific information and emotions. Therefore, hand-based gesture sequences can better synchronize with speech, enhancing the clarity and expressiveness of communication. Furthermore, the rendering process for hand gestures requires less computation than that for full-body gestures, helping to improve the efficiency of video generation.

[0114] Figure 6 A schematic structural block diagram of an apparatus 600 for video generation according to some embodiments of the present disclosure is shown. The apparatus 600 may be implemented in or included in an electronic device 110, for example. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0115] like Figure 6As shown, the device 600 may include a 3D pose representation determination module 601, configured to determine the 3D pose representation of the target object based on a source image of the target object. A reference pose action sequence determination module 602 is configured to determine a reference pose action sequence corresponding to the target audio based on the target audio. A target pose action sequence determination module 603 is configured to adjust the reference pose action sequence based on the 3D pose representation to obtain a target pose action sequence corresponding to the target object. A video generation module 604 is configured to generate a target video of the target object based on the source image of the target object and the target pose action sequence, wherein the target video represents the target object performing pose actions corresponding to the target pose action sequence while speaking the target audio.

[0116] In some embodiments of this disclosure, the reference posture action sequence determination module 602 may be configured to: determine the feature representation of the target audio; and, based on the feature representation of the target audio, determine the reference posture action sequence corresponding to the target audio using an action generation model.

[0117] In some embodiments of this disclosure, the target posture action sequence determination module 603 can be configured to: determine the three-dimensional spatial coordinates corresponding to each key point in the three-dimensional posture representation of the target object; determine the angle and length representations between each key point based on the three-dimensional spatial coordinates corresponding to each key point; and adjust each reference posture action included in the reference posture action sequence using the angle and length representations between each key point.

[0118] In some embodiments of this disclosure, the video generation module 604 may be configured to: determine the visual feature representation of the target object based on a source image of the target object; and render each pose action in the target pose action sequence based on the visual feature representation of the target object to obtain a target video of the target object.

[0119] In some embodiments of this disclosure, the source images of the target object include multiple source images, each with a different presentation angle of the target object. In this case, the video generation module 604 can be configured to: for a given pose action in a target pose action sequence, determine weight parameters corresponding to each of the multiple source images based on the action differences and the given pose action, where the action differences indicate the differences between the given pose action and the source pose actions corresponding to the multiple source images; determine a visual feature representation of the target object based on the multiple source images and their respective weight parameters; and render the given pose action based on the visual feature representation of the target object to obtain a video segment corresponding to the given pose action, where the video segment is part of the target video of the target object.

[0120] In some embodiments of this disclosure, the video generation module 604 may be specifically configured to: extract motion difference feature representations corresponding to each source image from a given pose action and source pose actions corresponding to multiple source images, wherein the motion difference feature representations indicate motion differences; and determine weight parameters corresponding to each source image based on the motion difference feature representations corresponding to each source image.

[0121] In some embodiments of this disclosure, the baseline pose sequence is determined using an action generation model, which includes an audio coding model and an action generation model. The audio coding model is trained through pre-training, and the audio coding model and the action generation model are jointly trained through fine-tuning.

[0122] In some embodiments of this disclosure, a model training module is also included. The model training module can be configured to: extract audio samples and baseline action sequence samples from a first video sample containing a first object sample; construct a plurality of positive samples, each positive sample including a first audio segment and a first baseline action sequence segment, wherein the first audio segment is segmented from the audio sample and the first baseline action sequence segment is segmented from the baseline action sequence sample, and the first audio segment and the first baseline action sequence segment are temporally aligned; construct a plurality of negative samples, each negative sample including a second audio segment and a second baseline action sequence segment, wherein the second audio segment is segmented from the audio sample and the second baseline action sequence segment is segmented from the baseline action sequence sample, and the second audio segment and the second baseline action sequence segment are temporally misaligned; and pre-train an audio coding model using the plurality of positive samples and the plurality of negative samples through contrastive learning.

[0123] In some embodiments of this disclosure, the target video of the target object is generated by a rendering model, and the model training module can be configured to: select source image samples and target image samples containing the second object sample from a second video sample containing the second object sample; determine the visual feature representation of the second object sample based on the source image sample; determine pose / action sequence samples associated with the second object sample based on the second video sample; render a given action sample in the pose / action sequence samples using the rendering model based on the visual feature representation to obtain a predicted image of the second object sample, wherein the given action sample is extracted from the target image sample; and adjust the parameters of the rendering model based on the difference between the predicted image and the target image sample.

[0124] In some embodiments of this disclosure, the source image samples include multiple source image samples, each containing an object sample presented from a different angle. The model training module can be configured to: determine weight parameters corresponding to each of the multiple source image samples based on action difference samples and a given action sample; the action difference is determined based on the given action sample and the source pose actions corresponding to each of the multiple source image samples. Based on each source image sample and its corresponding weight parameters, the visual feature representation of the object sample is determined.

[0125] In some embodiments of this disclosure, the reference posture sequence corresponding to the target audio is a sequence of actions including hand postures.

[0126] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The illustrated electronic device 700 may include or be implemented as Figure 1 Electronic devices 110 or Figure 6 Device 600.

[0127] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.

[0128] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 700.

[0129] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0130] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0131] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0132] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0133] According to an exemplary implementation of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform... Figure 2 The methods provided in the various optional modes are already available, so they will not be elaborated upon here.

[0134] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0135] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0136] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0138] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for video generation, comprising: Based on the source image of the target object, determine the three-dimensional pose representation of the target object; Based on the target audio, determine the reference posture action sequence corresponding to the target audio; The reference pose sequence is adjusted based on the three-dimensional pose representation to obtain the target pose sequence corresponding to the target object; as well as Based on the source image of the target object and the target posture action sequence, a target video of the target object is generated. The target video represents the target object performing a posture action corresponding to the target posture action sequence while speaking the target audio.

2. The method according to claim 1, wherein determining the target posture action sequence corresponding to the target audio comprises: Determine the feature representation of the target audio; as well as Based on the feature representation of the target audio, a reference posture action sequence corresponding to the target audio is determined using an action generation model.

3. The method according to claim 1, wherein adjusting the reference pose action sequence based on the three-dimensional pose representation comprises: Determine the three-dimensional spatial coordinates corresponding to each key point in the three-dimensional pose representation of the target object; Based on the three-dimensional spatial coordinates corresponding to each key point, the angle and length representations between the key points are determined. as well as By utilizing the angle and length representations between the key points, the reference posture actions included in the reference posture action sequence are adjusted.

4. The method according to claim 1, wherein generating the target video of the target object comprises: Based on the source image of the target object, determine the visual feature representation of the target object; as well as Based on the visual feature representation of the target object, each posture action in the target posture action sequence is rendered to obtain the target video of the target object.

5. The method according to claim 1, wherein the source image of the target object comprises a plurality of source images, the presentation angle of the target object in each source image is different from that of the others, and generating the target video of the target object comprises: For a given posture action in the target posture action sequence Based on the action difference and the given pose action, the weight parameters corresponding to each of the plurality of source images are determined, wherein the action difference indicates the difference between the given pose action and the respective source pose actions corresponding to the plurality of source images; Based on the multiple source images and the weight parameters corresponding to each of the multiple source images, the visual feature representation of the target object is determined; as well as The given pose action is rendered based on the visual feature representation of the target object to obtain a video segment corresponding to the given pose action, and the video segment is part of the target video of the target object.

6. The method according to claim 5, wherein determining the weight parameters corresponding to each of the plurality of source images includes: From the given pose action and the source pose actions corresponding to the plurality of source images, extract the action difference feature representation corresponding to each source image, wherein the action difference feature representation indicates the action difference; and Based on the action difference feature representation corresponding to each source image, the weight parameters corresponding to each source image are determined.

7. The method according to claim 1, wherein the reference pose sequence is determined using a motion generation model, the motion generation model comprising an audio coding model and a motion generation model. The audio coding model is trained in a pre-training manner, and the audio coding model and the action generation model are jointly trained in a fine-tuning manner.

8. The method of claim 7, wherein the audio coding model is pre-trained in the following manner: Extract audio samples and baseline motion sequence samples from a first video sample containing a first object sample; Multiple positive samples are constructed, each positive sample including a first audio segment and a first reference action sequence segment. The first audio segment is obtained by segmenting from the audio sample, and the first reference action sequence segment is obtained by segmenting from the reference action sequence sample. The first audio segment and the first reference action sequence segment are time-aligned. Multiple negative samples are constructed, each negative sample including a second audio segment and a second reference action sequence segment. The second audio segment is obtained by segmenting from the audio sample, and the second reference action sequence segment is obtained by segmenting from the reference action sequence sample. The second audio segment and the second reference action sequence segment are not aligned in time. as well as The audio coding model is pre-trained using the multiple positive samples and the multiple negative samples through contrastive learning.

9. The method of claim 1, wherein the target video of the target object is generated by a rendering model, and the rendering model is trained in the following manner: Select a source image sample and a target image sample containing the second object sample from the second video sample containing the second object sample; Based on the source image samples, determine the visual feature representation of the second object sample; Based on the second video sample, determine the pose and motion sequence sample related to the second object sample; Based on the visual feature representation, the rendering model is used to render a given action sample in the pose action sequence sample to obtain a predicted image of the second object sample, wherein the given action sample is extracted from the target image sample; as well as The parameters of the rendering model are adjusted based on the differences between the predicted image and the target image sample.

10. The method of claim 9, wherein the source image samples comprise a plurality of source image samples, the object samples in each source image sample have different presentation angles relative to each other, and determining the visual feature representation of the object samples comprises: Based on the action difference samples and the given action sample, the weight parameters corresponding to each of the plurality of source image samples are determined, wherein the action difference is determined based on the source pose action corresponding to each of the given action sample and the plurality of source image samples. as well as Based on each source image sample and its corresponding weight parameters, the visual feature representation of the object sample is determined.

11. The method of claim 1, wherein the reference posture action sequence corresponding to the target audio is an action sequence including hand posture.

12. An apparatus for video generation, comprising: The 3D pose representation determination module is configured to determine the 3D pose representation of the target object based on the source image of the target object; The reference posture action sequence determination module is configured to determine the reference posture action sequence corresponding to the target audio based on the target audio. The target posture action sequence determination module is configured to adjust the reference posture action sequence based on the three-dimensional posture representation to obtain the target posture action sequence corresponding to the target object; as well as The video generation module is configured to generate a target video of the target object based on the source image of the target object and the target posture action sequence. The target video represents the target object performing a posture action corresponding to the target posture action sequence while speaking the target audio.

13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 11.

15. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Training method of virtual image action generation model and action generation method and device

    CN114972590A

  • Method and device for generating video from voice

    CN115550744A