Human body video generation method and apparatus based on multi-modal control
Patent Information
- Application Number
- PCT/CN2026/076107
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2026-01-30
- Publication Date
- 2026-09-03
Smart Images

Figure CN2026076107_03092026_PF_FP_ABST
Abstract
Description
A method and apparatus for generating human body videos based on multimodal control
[0001] This application claims priority to Chinese Patent Application No. 2025102382051, filed on February 28, 2025, entitled "A Method and Apparatus for Generating Human Body Video Based on Multimodal Control", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, specifically to a method and apparatus for generating human body videos based on multimodal control. Background Technology
[0003] With the development of video generation technology, using a given audio signal to drive a target speaker in a reference image to produce synchronized, full-body motion video has become an important research direction. Existing methods mainly rely on audio input and predefined motion templates to drive the target speaker in the reference image to generate video.
[0004] However, limited by template constraints, most existing methods can only generate animations of the target speaker's head or upper body. Furthermore, these methods still have significant limitations in generating natural and fluid full-body animations, especially in scenarios requiring the expression of complex emotions and fine movements, where the results are often unsatisfactory. Therefore, developing a video generation system capable of flexibly generating full-body animations is of great significance. Summary of the Invention
[0005] This specification provides a human body video generation scheme based on multimodal control, which can generate human body videos through multimodal driving signals and achieve fine control over human body movements in the generated videos.
[0006] In a first aspect, embodiments of this specification provide a method for generating human body video based on multimodal control, comprising: acquiring text prompt information, audio driving signals, and a reference image of a target speaker, wherein the text prompt information includes text information for providing action prompts to the target speaker, and the audio driving signals are audio information containing speech content; generating motion posture representation information of the target speaker based on the text prompt information and the audio driving signals, wherein the motion posture representation information is used to represent the motion posture of the target speaker; and generating a speaking video of the target speaker based on the reference image, the audio driving signals, and the motion posture representation information, wherein the speaking video includes the body movements of the target speaker when expressing the speech content.
[0007] In some embodiments, generating motion posture representation information of the target speaker based on the text prompt information and the audio driving signal includes: generating a three-dimensional 3D posture sequence of the target speaker based on the text prompt information and the audio driving signal, wherein the 3D posture sequence is a time series representing the 3D positions of key points of the human body; and converting the 3D posture sequence to obtain motion posture representation information of the target speaker, wherein the motion posture representation information is a time series representing the two-dimensional 2D positions of key points of the human body.
[0008] In some embodiments, the 3D pose sequence is a motion code sequence used to characterize the 3D position of key points on the human body. The step of converting the motion pose representation information of the target speaker based on the 3D pose sequence includes: determining the motion pose representation information that matches the 3D pose sequence based on a pre-established relational database, wherein the relational database contains the mapping relationship between the motion code corresponding to the key points on the human body and the 2D pose.
[0009] In some embodiments, the 3D pose sequence is a motion code sequence used to characterize the 3D positions of key points on the human body. Generating the 3D pose sequence of the target speaker based on the text prompt information and the audio driving signal includes: extracting fusion features based on the audio driving signal and the text prompt information; determining the code probability distribution corresponding to each moment in the audio driving signal based on the fusion features, wherein the code probability distribution represents the probability that the target motion code at that moment is any of the motion codes in the codebook, and the codebook contains multiple motion codes; and determining the motion code sequence of the target speaker based on the code probability distribution corresponding to each moment in the audio driving signal.
[0010] In some embodiments, generating motion posture representation information of the target speaker based on the text prompt information and the audio driving signal includes: extracting features from the text prompt information by the text branch of the motion generator to obtain text features; extracting features from the audio driving signal by the audio branch of the motion generator to obtain audio features; and generating motion posture representation information of the target speaker based on the fusion result of the text features and the audio features.
[0011] In some embodiments, the motion generator is trained as follows: masked motion pose representation labels and text cue samples are input into the text branch of the motion generator for feature extraction to obtain text feature samples; a first motion pose prediction result is generated based on the text feature samples; the text branch of the motion generator is trained based on the motion pose representation labels and the first motion pose prediction result; fixed text cue samples are input into the text branch of the motion generator for feature extraction to obtain fixed text features; audio-driven samples are input into the audio branch of the motion generator for feature extraction to obtain audio feature samples; a second motion pose prediction result is generated based on the fusion result of the fixed text features and the audio feature samples; and the audio branch of the motion generator is trained based on the motion pose representation labels and the second motion pose prediction result.
[0012] In some embodiments, generating a speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: determining facial motion features based on the audio driving signal; determining limb motion features based on the motion posture representation information; and generating a speaking video of the target speaker based on the reference image, the facial motion features, and the limb motion features.
[0013] In some embodiments, determining facial motion features based on the audio driving signal includes: determining the facial motion features based on the audio driving signal, facial expression tags, and / or blink frequency signals.
[0014] In some embodiments, generating a speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: extracting the identity features of the target speaker based on the facial image of the target speaker in the reference image; and generating a speaking video of the target speaker based on the reference image, the audio driving signal, the identity features, and the motion posture representation information.
[0015] In some embodiments, generating the speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: generating the speaking video of the target speaker using a multimodal diffusion model based on the reference image, the audio driving signal, and the motion posture representation information; the multimodal diffusion model is trained as follows: inputting a first sample image and a first motion posture representation label into the first multimodal diffusion model to generate a first predicted video; training the first multimodal diffusion model based on the difference between the first predicted video and the first sample video to obtain a second multimodal diffusion model; inputting an audio driving sample, a second sample image, and a second motion posture representation label into the second multimodal diffusion model to generate a second predicted video; and training the second multimodal diffusion model based on the difference between the second predicted video and the second sample video to obtain the multimodal diffusion model.
[0016] Secondly, embodiments of this specification provide a human body video generation device based on multimodal control, comprising: a data acquisition unit configured to acquire text prompt information, an audio driving signal, and a reference image of a target speaker, wherein the text prompt information includes text information for providing action prompts to the target speaker, and the audio driving signal is audio information containing speech content; a motion generation unit configured to generate motion posture representation information of the target speaker based on the text prompt information and the audio driving signal, wherein the motion posture representation information is used to represent the motion posture of the target speaker; and a video generation unit configured to generate a speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information, wherein the speaking video includes the body movements of the target speaker when expressing the speech content.
[0017] Thirdly, embodiments of this specification provide a computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any of the implementations in the first aspect.
[0018] In the solutions provided by the above embodiments of this specification, by combining audio driving signals and text prompts to jointly drive the target speaker in the reference image to generate video, the full body movements of the target speaker can be flexibly controlled, achieving fine control over the generated video. Moreover, it is not restricted by the actions specified by the template and the body range displayed in the reference image, and can generate full body movements from any reference image, no longer limited to the head or upper body, thereby improving the diversity and expressiveness of the generated video. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the various embodiments disclosed in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only a few embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 is a schematic diagram of the human body video generation process in an embodiment of this specification;
[0021] Figure 2 is a flowchart of the human body video generation method based on multimodal control in the embodiments of this specification;
[0022] Figure 3 is the network structure of the motion generator in the embodiment of this specification;
[0023] Figure 4 is a flowchart of the training steps of the motion generator in the embodiments of this specification;
[0024] Figure 5 shows the network structure of the audio branch and dual-input branch in the motion generator in the embodiments of this specification;
[0025] Figure 6 is a flowchart of the training steps of the multimodal diffusion model in the embodiments of this specification;
[0026] Figure 7 is a schematic diagram of the human body video generation device based on multimodal control in the embodiments of this specification. Detailed Implementation
[0027] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0028] As previously mentioned, current speaker video generation typically involves providing an audio driving signal and a reference image of the speaker, with the goal of generating a video of the speaker matching the audio content, ensuring that the speaker's lip movements are synchronized with the audio. The generated video primarily focuses on the head and shoulder area, particularly the dynamic changes in lips and facial expressions. To enhance speaker expressiveness, subsequent research has extended animation generation to the upper body, allowing for the display of gestures and body language. These methods perform well in formal settings (such as news broadcasts or speeches) because, in these scenarios, generated body movements are usually limited to the upper body and primarily focus on gestures, with little involvement of the lower body. However, this approach has limitations. For instance, in talk shows or stand-up comedy performances, full-body movements are indispensable for actors to engage the audience with compelling language. Furthermore, since humans rarely remain still while speaking, to make the generated video more realistic, even using a half-body image as a reference, it is necessary to generate full-body movements to more accurately recreate the natural posture during speech. For example, even if only the upper body of the speaker is shown in the video, the swaying of the lower body can still affect the movement of the upper body.
[0029] While audio driving signals can precisely determine the movement of a speaker's lips, it is difficult to link audio with specific body movements. That is, it is difficult to achieve precise control of the whole body movement based solely on audio. Considering this problem, embodiments of this specification propose at least one human body video generation method based on multimodal control. In addition to audio driving signals, text prompts are introduced as control signals to jointly drive the target speaker in the reference image to generate video. While controlling the speaker's whole body movement in sync with the given audio driving signals, text prompts are used to flexibly control the specific movements of the target speaker's whole body, achieving fine control over the generated video. Moreover, it is not constrained by the actions specified by the template or the body range displayed in the reference image. It can generate whole body movements from reference images of any size (such as 1 / 4 or 1 / 2 size upper body images or full-body portraits), enabling the target speaker to convey richer information through body language, thereby improving the diversity and expressiveness of the generated video.
[0030] The process of generating human body videos based on multimodal control is described below with reference to Figure 1. Figure 1 is a schematic flowchart of human body video generation in an embodiment of this specification. The human body video generation method based on multimodal control in this embodiment can be executed by any device, platform, or device cluster with computing and processing capabilities, including steps S201-S203 as shown below.
[0031] As shown in Figure 2, in step S201, text prompt information, audio driving signal and reference image of target speaker are acquired.
[0032] The text prompts include text information used to prompt the target speaker to perform actions. For example, the text prompts may contain text indicating a specified action, such as "The target speaker is spinning around." The text prompts may also include the specified action and the time information when the action occurs, such as "The target speaker waves their right hand to greet at the beginning of the video and waves both hands to say goodbye at the end," or "The target speaker waves both hands at the 10th second." The text prompts may also include a description of the emotion conveyed by the action, such as "The target speaker is very excited." The audio driving signal is audio information containing speech content, such as a recording of a human speech or an audio recording of singing. This embodiment does not limit the specific content of the text prompts and audio driving signals.
[0033] In practice, the target speaker can refer to any object capable of expressing speaking actions. Specifically, it can be a real person, a virtual character or digital human, or even another species. The reference image of the target speaker includes the target speaker's facial area. This embodiment does not limit the body area covered by the reference image. For example, it can include the target speaker's head and shoulders, the area above the waist, or the entire body.
[0034] This embodiment does not limit the method of acquiring text prompts, audio driving signals, and reference images of the target speaker. For example, it can acquire text prompts, audio driving signals, and reference images that are specified, edited, or uploaded by the user.
[0035] Next, in step S202, motion posture representation information of the target speaker is generated based on the text prompt information and the audio driving signal.
[0036] The motion posture representation information is used to represent the motion posture of the target speaker. This motion posture can be synchronized with the audio driving signal, for example, it can be a time series of the same length as the audio driving signal. The motion posture representation information can include the target speaker's body movement information in space at different times, including specific movements of the limbs and head movements. Specifically, it can include the coordinate positions of joints and body parts in 2D (2-Dimensional) plane or 3D (3-Dimensional) space, the angles of each joint, the degree of rotation or bending between joints, and the combination of relative positions between various body parts.
[0037] When generating motion posture information of the target speaker, the motion posture of the target speaker can be determined based on the text prompts for the target speaker's actions and the speech rhythm and content in the audio driving signal. For example, faster speech is usually associated with more excited or tense postures, which may be accompanied by more frequent body movements, such as waving or quickening of pace. A pleasant tone may be accompanied by relaxed body movements, while an angry tone may be accompanied by body tension.
[0038] For example, motion posture representation information can be a 2D posture sequence, specifically a time series representing the positions of human key points in 2D space, or a 3D posture sequence, specifically a time series representing the positions of human key points in 3D space. Here, human key points are used to represent the position of each joint or body part, and combinations of human key point positions can describe different human motion postures.
[0039] It should be noted that the selection and setting of human body key points can be determined according to actual application needs. For example, when computing resources are limited, fewer human body key points can be used, while when computing resources are sufficient, more human body key points can be used to make the generated animation more refined. In this embodiment, human body key points representing the whole body movement are used so that the generated speaking video includes the target speaker's whole body movement. It is understood that in other embodiments, the body parts to be represented can be flexibly selected according to different needs, and human body key points representing partial body movements, such as key points of the upper body, can be used to generate animation or video effects that better meet actual needs.
[0040] In practice, a motion generator can be used to process text prompts and audio driving signals to generate motion posture representation information of the target speaker. This embodiment does not limit the network structure of the specific motion generator used.
[0041] In one example, the motion generator can have two input branches: a text branch and an audio branch. The text branch of the motion generator extracts features from the text prompts to obtain text features; the audio branch of the motion generator extracts features from the audio driving signals to obtain audio features; and the motion posture representation information of the target speaker is generated based on the fusion result of the text features and audio features.
[0042] For example, the motion generator can employ a two-input-branch transformer network structure to effectively capture long-range dependencies in textual cues and audio driving signals through a self-attention mechanism. Based on the combined information from the text and audio, it generates the motion posture of the target speaker. When fusing text and audio features, the fusion can be performed layer by layer after extracting the audio and text features at each transformer layer of the audio and text branches. Alternatively, the audio and text features output from the last layer of the audio and text branches can be fused to obtain a fused result. The motion posture representation information is then generated using the fused result. Combining text and audio information helps generate more realistic and consistent motion performance.
[0043] The training process of the motion generator will be explained below with reference to the network structure of the motion generator shown in Figure 3. As one implementation method, the motion generator can be trained through the following steps S401-S407:
[0044] As shown in Figure 4, in step S401, the masked motion tokens and text prompt samples are input into the text branch of the motion generator for feature extraction to obtain text feature samples.
[0045] In this embodiment, the text branch is trained first, then the weight parameters in the trained text branch are frozen, and the audio branch is trained. The motion pose representation label is a sequence of multiple motion poses. The masking process can be achieved by replacing some motion poses in the sequence with special masks to randomly mask the motion poses in the sequence.
[0046] For example, for a given sample video, the speaker's motion posture can be extracted from each frame of the sample video to obtain motion posture representation labels. Where, m t Let T represent the speaker's motion posture in frame t. The sequence length is T, and the motion posture of some frames is randomly occluded. Semantic information is extracted from the target speaker's actions in the sample video to obtain a text prompt sample. For example, if the speaker in the sample video is dancing, the text prompt sample would be "The speaker is dancing".
[0047] For the text branch, the training objective is to predict the occluded motion pose based on a given text cue sample, thereby enabling the text branch to learn the mapping from text to motion pose. Specifically, CLIP (Contrastive Language-Image Pretraining) can be used to extract initial features from the text cue samples and input them into the text branch to obtain text feature samples.
[0048] In step S402, a first motion pose prediction result is generated based on the text feature samples.
[0049] Then, the text branch predicts the occluded motion pose based on text feature samples to obtain the first motion pose prediction result.
[0050] In step S403, the text branch of the motion generator is trained based on the motion pose representation label and the first motion pose prediction result.
[0051] Specifically, the loss function can be calculated using the motion pose representation label and the first motion pose prediction result, and the text branch can be trained with the goal of minimizing the loss function.
[0052] For example, the training objective could be to maximize the log-likelihood of the first motion pose prediction. To minimize the loss function, It is expressed as follows:
[0053] Among them, [mask] t This indicates whether the t-th motion pose is masked. If it is masked, the value is 0; otherwise, it is 1, to ensure that the model only needs to predict the masked motion pose, rather than the entire sequence.
[0054] The loss function can be expressed as follows:
[0055] When the loss function reaches the preset minimum value, or the training reaches the preset number of iterations, the training of the text branch ends and the training of the audio branch begins.
[0056] Next, in step S404, the fixed text prompt is input into the text branch of the motion generator for feature extraction to obtain fixed text features.
[0057] In this stage, the model is trained based on audio input and corresponding text to predict motion poses. Here, a fixed text input (e.g., "a person is giving a speech") can be given as the control signal for the text modality.
[0058] In step S405, the audio driving sample is input into the audio branch of the motion generator for feature extraction to obtain audio feature samples.
[0059] Audio-driven samples can be obtained by extracting audio from sample videos. First, the audio-driven samples can be processed by wav2vector (waveform to vector) to obtain audio tokens (initial features of audio samples). Then, the audio tokens are input into the audio branch for feature extraction to obtain audio feature samples.
[0060] In step S406, a second motion pose prediction result is generated based on the fusion result of fixed text features and audio feature samples.
[0061] For example, after extracting the audio feature samples and fixed text features from the self-attention layer of each transformer in the audio and text branches, the features are fused layer by layer, and the second motion pose prediction result is obtained based on the fusion result of the last layer.
[0062] In step S407, the audio branch of the motion generator is trained based on the motion pose representation label and the second motion pose prediction result.
[0063] Specifically, the loss function can be calculated using the motion pose representation label and the second motion pose prediction result, and the audio branch can be trained with the goal of minimizing the loss function. During this process, the weight parameters of the text branch are frozen, that is, the gradient parameters of the text branch are not updated during backpropagation.
[0064] For example, the loss function can be the negative log-likelihood function, as shown below:
[0065] Where, m i a represents the i-th motion posture. i This represents the corresponding audio token, text is a fixed text prompt, and p φ (m i |a i (,text) is a given a i When combined with text, the predicted motion pose m is obtained. i The probability is represented by the parameter φ, which indicates the probability distribution. The smaller the negative log-likelihood, the closer the model's prediction is to the actual value. The model is trained by minimizing this loss, so that the model's predicted action is as close as possible to the actual action.
[0066] In the training process described above, training the text branch first and then freezing it means that the weight parameters of the text branch do not need to be trained repeatedly. This can greatly reduce the computational complexity. Moreover, in multimodal learning tasks, by keeping the parameters of the text modality fixed, the model can better focus on learning other modalities, avoid overfitting, and accelerate the training process.
[0067] As an implementation approach, considering that 3D pose is highly controllable and can provide more precise information for controlling the movement of human skeletons and joints, making it easier for users to control and edit human movements, and it is also easy to semantically associate with audio driving signals and text prompts to better model human movement, while 2D pose contains less detailed information about human movements, such as lacking key angle information, making it difficult to generate 2D poses to describe fine human movements using audio driving signals and text prompts. In addition, considering that the correspondence between 2D pose and image pixels is more explicit, it can be directly used to guide the generation of details for each pixel during rendering, ensuring that the rendered image looks more natural and accurate, this step can generate a 3D pose sequence of the target speaker based on text prompts and audio driving signals, and then convert the 3D pose sequence to obtain the motion pose representation information of the target speaker.
[0068] As shown in Figure 1, a motion generator can generate a 3D pose sequence using text prompts and audio driving signals. This sequence is then converted into a 2D pose sequence by a motion transducer, allowing the 2D poses to guide pixel generation during the rendering phase. The text prompts can be processed using CLIP to obtain initial text features before being input into the motion generator, and the audio driving signals can be processed using wav2vector to obtain initial audio features before being input into the motion generator. In this example, the motion generator is a neural network that outputs a 3D pose sequence.
[0069] The 3D pose sequence is a time series representing the 3D positions of key points on the human body. For example, the 3D pose (also called motion tokens) at each time point in the 3D pose sequence can be represented using the 3D position parameters of key points on the human body. These position parameters include the specific coordinates of each joint in three-dimensional space. For example, key points on the human body can include joints such as the head, shoulders, elbows, knees, and ankles. The 3D position of each joint can be represented by three coordinates: x, y, and z, thus forming a complete 3D pose. The combination of 3D position parameters of different joints can accurately describe every movement of the target speaker in space. Through these position parameters, details of the target speaker's rotation, movement, bending, and other movements can be obtained. For example, SMPL (Skinned Multi-Person Linear Model) can be used to model the shape and pose of the human body. By representing the 3D geometry of the human body through linear blending skin, various poses of the human body can be represented. These 3D poses are controlled by a set of joint parameters, which can be denoted as... Where J is the number of joints, and each joint is represented by three parameters that indicate its coordinates on the x, y, and z axes.
[0070] Next, after obtaining the 3D pose sequence, it is converted into a time series representing the 2D positions of human key points (i.e., a 2D pose sequence) for video generation. The 2D motion pose at each time point can be represented by the 2D positions of human key points.
[0071] This embodiment does not limit the specific conversion method used from 3D pose to 2D pose. For example, the coordinate information of key human body points in 3D space at each time point can be projected onto a 2D plane to obtain the corresponding 2D motion pose. Specifically, 3D pose includes the x, y, and z coordinates of each key human body point in 3D space, while 2D pose converts these 3D coordinates into 2D coordinates (x and y coordinates). In this conversion process, the 3D motion pose at each time point contains a series of three-dimensional positions of key human body points. After projection or mapping, these positions are converted into corresponding 2D coordinates. In this way, a 2D pose sequence with time as the dimension can be obtained, where the 2D coordinates (x, y) of each key human body point reflect its position in the image plane. Such a 2D pose sequence can better guide pixel generation during the image rendering stage. At the same time, since its position is obtained by converting precise 3D positions, it can ensure the accuracy of human body movements.
[0072] As one implementation approach, to reduce computational complexity, motion codes can be used as a vectorized discrete representation of the original human 3D motion. In this case, the 3D pose sequence is a sequence of motion code sequences used to characterize the 3D positions of key points on the human body; that is, the 3D pose at each time point is represented by the corresponding motion code. Assume the human 3D motion at frame t (or time point t) is represented by m... t It means that m t Containing key information such as the position, rotation, and velocity of each joint in the human body, a continuous time-step 3D motion sequence of the human body can be represented as follows: It is understandable that for human 3D motion sequences, the changes in motion are usually continuous, which will result in very complex data contained in the human 3D motion sequence and a very large amount of subsequent computation. This embodiment maps the continuous human 3D motion to a finite discrete space (codebook). In this way, complex motion data can be represented in a finite, discrete code space, thereby simplifying data processing and reducing computational complexity.
[0073] Considering that directly projecting 3D poses into 2D poses can result in stiff and unnatural human movements, this embodiment can also construct a relational library to link motion codes in the codebook with 2D poses, making the converted 2D poses more realistic. In the motion converter, motion pose representation information matching the 3D pose sequence can be determined based on this pre-established relational library.
[0074] The relational database contains the mapping relationship between motion codes corresponding to human key points and 2D poses. By querying the relational database for the 2D pose corresponding to each motion code in the motion code sequence, motion pose representation information composed of multiple 2D poses can be obtained.
[0075] The generation of motion code sequences will be explained first, followed by the process of establishing the relational database.
[0076] First, a VQ-VAE (Vector Quantized Variational AutoEncoder) can be trained as a 3D Human Motion Tokenizer to learn and quantize the discrete representation of human 3D motion, thereby mapping continuous human 3D motion to discrete motion codes.
[0077] Specifically, VQ-VAE uses a codebook of size K that can be learned. An autoencoder is used to reconstruct the original 3D human motion sequence, where each motion code... Corresponding to a discretized 3D human motion, d c This represents the dimension of the motion code. Given an original 3D human motion sequence... Encoder ε maps it to a latent feature sequence in l is the temporal downsampling rate of ε. For each latent feature z i The quantized motion code sequence is obtained by selecting the closest motion code in codebook C and quantizing it. As shown below:
[0078] code sequence Reconstructing the 3D motion sequence of the human body
[0079] For example, a codebook of size 512 can be constructed through the above process. The codebook contains 512 motion codes, and different motion codes can be combined into a sequence of motion codes to express the continuous movement of the human body.
[0080] This embodiment does not limit the method of training VQ-VAE. For example, VQ-VAE can be trained by combining motion reconstruction loss with the latent embedding loss of the quantization layer, and the loss function is as follows:
[0081] in, This represents the motion reconstruction loss, which uses the L1 norm to measure the difference between the input original human 3D motion sequence M and the output reconstructed by the model. The differences between them This represents the potential embedding loss, which uses the L2 norm to measure Z and The difference between them is that sg[·] indicates stopping the gradient operation, β is the weighting coefficient, and the codebook can be updated by the exponential moving average method.
[0082] After the 3D human motion segmenter is trained, the obtained codebook can be used to generate motion code sequences. For example, the process of generating motion code sequences may include the following steps:
[0083] First, fusion features are extracted based on audio driving signals and text prompts.
[0084] For example, a motion generator with the network structure shown in Figure 5 can be used to extract fused features. Figure 5 shows the network structure in the audio branch for converting audio to motion (left figure) and the network structure in the two input branches driven by multimodal control signals (right figure).
[0085] To facilitate understanding, the process of extracting audio features from the audio branch and predicting 3D pose sequences will be explained first, followed by the process of extracting fused features through a network structure with two input branches.
[0086] In practice, the audio branch transformer can be used. Predicting motion code sequences based on audio driving signals. Specifically, the audio driving signal can be processed to obtain initial audio tokens, which are then input into... In the encoder, long-range dependencies and contextual relationships in the audio sequence are captured. The output of each transformer self-attention layer is collected and represented as... in Let S represent the output of the s-th layer, where S can be set to 8. Then, the audio features of the last layer's output are processed through a simple linear layer. Probability of converting to audio code Where T and K represent the time length of the human 3D motion sequence and the size of the codebook, respectively (for example, when the codebook contains 512 motion codes, K is 512). audio The code probability distribution of the motion code at each time t is shown below:
[0087] For each time t, a matching motion code is selected from the codebook based on the probability distribution of the code. For example, the code with the highest probability in formula (3) is selected. The corresponding c k ), to obtain the motion code sequence
[0088] It should be noted that if motion code is not used as the vectorized discrete representation of the original human 3D motion, the previously trained VQ-VAE decoder can also be used. Convert motion code sequences into human 3D motion sequences As a 3D pose sequence.
[0089] The following describes the two-input branch network structure, which can be viewed as an audio branch network structure with an added text branch (transformer). This is used for predicting motion code sequences based on text. The text branch, which converts text into motion, has the same network structure as the audio branch. Specifically, CLIP can be used to extract text prompts to obtain initial text features, and then the text prompts are input... In the encoder, the processing procedure is the same as for the audio branch; the output of each transformer self-attention layer is collected and represented as... in The output of the s-th layer is shown in Figure 5. encoder and The text features and audio features output from each transformer self-attention layer in the encoder are fused together to obtain the final fused features in the last layer.
[0090] Then, based on the fusion features, the code probability distribution corresponding to each time t in the audio driving signal is determined. This code probability distribution can be found in formula (6), which represents the probability that the target motion code at time t is one of the motion codes in the codebook. Based on the code probability distribution corresponding to each time in the audio driving signal, the motion code sequence of the target speaker is determined. The process of determining the matching motion code sequence based on the code probability distribution is the same as that in the aforementioned audio branch and will not be repeated here.
[0091] In this way, the fusion features of text prompts and audio driving signals can be extracted through a two-input-branch network structure, and discrete motion code sequences can be generated. By processing audio and text information simultaneously, information from different modalities can be effectively fused. This multimodal learning helps improve the model's understanding and generation capabilities for complex tasks, especially when it comes to the common representation of audio and text, it can better capture the relationship between the two and generate matching motion code sequences.
[0092] After obtaining the motion code sequence, the next step is to map them to the 2D pose sequence using a relational database. This database can be obtained by extracting the 2D pose and motion code from the template video separately and establishing the mapping relationship between them.
[0093] Specifically, SMPL-X (extended SMPL) data and 2D poses can be extracted from template videos. Then, a previously trained 3D human motion marker is used to convert the SMPL-X data (i.e., the original human 3D motion) into motion codes. A mapping relationship is established between the motion codes and 2D poses corresponding to each video frame and saved in a relational database. In this way, by extracting the 2D poses of the human body in the real world from the template video and aligning them with the motion codes in the codebook, a large number of motion code-2D pose pairs can be created, ensuring that the 2D poses obtained through motion code conversion can generate natural and smooth motion.
[0094] Next, in step S203, a speaking video of the target speaker is generated based on the reference image, audio driving signal and motion posture representation information. The speaking video includes the body movements of the target speaker when expressing speech content.
[0095] After obtaining the motion posture representation information indicating full-body movements, the motion representation can be transformed into an explicit posture sequence. This sequence, combined with audio-driven signals and reference images, generates the final video with accurate lip-sync, rich gestures, and full-body motion. The reference images primarily guide the generation of the target speaker's image, while the audio-driven signals primarily guide the target speaker's lip movements. Furthermore, both the reference images and audio-driven signals can further guide the generation of the target speaker's facial expressions. For example, facial expressions (such as smiling, frowning, and mouth movements) can be adjusted based on the rhythm and emotional changes in the audio.
[0096] For example, when the reference image is an image of the upper body of the target speaker, the generated speaking video can be a video of the upper body of the target speaker. In addition to including the target speaker's lip movements, gestures, and facial expressions synchronized with the audio driving signal, the speaking video also includes the target speaker's body movements driven by the text and audio. Because this embodiment considers the positions of key points of the entire human body when generating motion posture representation information, the generated target speaker will have more natural body movements during speech, such as tilting or swaying of the upper body caused by lower body movement.
[0097] Specifically, this step can use deep learning models to generate speaking videos, such as generative adversarial networks, VAEs (variational autoencoders), or video stream generation models. The reference image, audio driving signals, and motion pose representation information are input into the model, and the output is the speaking video of the target speaker.
[0098] In this embodiment, a multimodal diffusion model can be used to generate spoken video. It generates video data by gradually noise-enhancing the input data and recovering the noise. For example, referring to Figure 1, the multimodal diffusion model includes at least: a VAE encoder for generating noise latent variables based on a reference image; a U-Net (a convolutional neural network for image segmentation) encoder for encoding the input reference image features, motion pose representation information, and audio driving signals to obtain a feature vector, wherein the reference image features can be obtained by feature extraction from the reference image using CLIP; a U-Net decoder for decoding the feature vector; and a VAE decoder for decoding the features output by the U-Net decoder to generate the spoken video. It is understood that in other embodiments, the multimodal diffusion model may also employ other types of network structures.
[0099] As one approach, to avoid the influence of the location information of key facial points in the motion posture representation information on facial expressions, this step can determine facial motion features based on audio driving signals; determine limb motion features based on motion posture representation information; and generate a speaking video of the target speaker based on the reference image, facial motion features, and limb motion features.
[0100] Among them, facial motion features refer to the motion characteristics of the facial region, while limb motion features refer to the motion characteristics of other body regions besides the face. In this way, facial motion features can be generated without being affected by the position of facial key points in the motion posture representation information. As a result, lip movements and facial expressions in the generated video can be driven by audio-driven signals, making the generated facial expressions more natural.
[0101] For example, in the posture condition part of the conditional synergy effect shown in Figure 1, if the motion posture representation information obtained in the previous step is a full-body posture sequence including the head, the full-body posture sequence can be redefined as a composite representation. Specifically, the posture sequence of the body below the neck can be retained and the face can be covered with a head mask, the center of which can be located at the midpoint of the face above the neck.
[0102] Then, this composite representation is mapped through PoseNet to obtain limb motion features, which are then added element-wise to the output of the first convolutional layer of U-Net. This allows the influence of limb motion features to be initially superimposed on the noisy latent variables to guide the diffusion at each step, thus playing a role in each diffusion step for more accurate human motion generation. Next, an additional cross-attention layer (layer 4 in Figure 1) is introduced after the original cross-attention layer in each U-Net block (layer 3 in Figure 1, i.e., the layer where the reference image features are input), specifically for inputting facial motion features to generate facial-related dynamics.
[0103] As one implementation method, in order to further ensure the consistency of the identity of the speaker in the generated video with the target speaker in the reference image, this embodiment can extract the identity features of the target speaker based on the facial image of the target speaker in the reference image; then, based on the reference image, audio driving signal, identity features and motion posture representation information, the speaking video of the target speaker is generated.
[0104] Identity features include facial feature information of the target speaker, used to identify the target speaker's identity. For example, it may include information such as the relative positions of facial organs, skin texture features, and the three-dimensional structure of the face.
[0105] Specifically, while extracting reference image features from the reference image, a pre-trained face recognition network is used to detect facial regions in the reference image and extract identity features. These identity features are then input into a newly introduced additional cross-attention layer, which interacts with facial motion features to accurately generate facial movements while maintaining consistency in the identity of the target speaker.
[0106] In other examples, in addition to audio driving signals, other control signals can be used to drive facial expressions, such as determining facial motion features based on audio driving signals, expression labels, and / or blink frequency signals.
[0107] In this example, in addition to the audio driving signal, conditional signals such as facial expression tags and blink frequency signals related to facial movement are also used as additional control signals to drive facial expressions.
[0108] Facial expression labels typically refer to tags used to categorize and label facial expressions. These tags represent specific facial expressions such as smiling, anger, sadness, and surprise, helping to accurately adjust and reproduce specific facial expressions, making the generated video look more natural and realistic. Facial expression labels can be detected using an emotion detector. For example, as shown in the audio condition section of Figure 1, a facial image can be extracted from a reference image of the target speaker and input into the emotion detector to obtain facial expression labels. Alternatively, facial images or videos of other speakers can be input into the emotion detector to obtain a sequence of facial expression labels.
[0109] Blink frequency refers to the frequency or speed of eye blinking, typically used to describe the number of blinks over a period of time. It can be extracted from video of another speaker or set manually. Adjusting the blink frequency can make the facial expressions of the target speaker more natural and realistic.
[0110] As shown in Figure 1, the audio conditional part of the conditional synergy effect can directly input the above-mentioned conditional signals related to facial movements as additional control signals into the newly added cross-attention layer in U-Net (the 4th layer of the encoder in Figure 1). These signals interact with identity features and accurately generate facial movements such as lip movements, facial expressions, and eye blinks while maintaining the consistency of the target speaker's identity.
[0111] The training method of the multimodal diffusion model is described below. It can be understood that in this embodiment, the motion generator and the multimodal diffusion model can be trained separately. In other embodiments, an end-to-end training method can also be used to train the motion generator and the multimodal diffusion model at the same time.
[0112] Considering that the input signal of the multimodal diffusion model comes from a variety of different modalities, in order to ensure the quality and stability of the generated video, this embodiment trains the multimodal diffusion model through a two-stage training method. Each stage focuses on the task of a different modality, so that the model can gradually learn the representation and generation capabilities of different modalities, as shown in Figure 6. The training process specifically includes steps S601-S604.
[0113] In step S601, the first sample image and the first motion pose representation label are input into the first multimodal diffusion model to generate the first predicted video.
[0114] First, based on the input of a first sample image of the visual modality and a first motion pose representation label, video is generated to learn the motion of the whole body.
[0115] The first sample video is a video of a speaker expressing speech content, the first sample image contains at least an image of the speaker's facial area, the first sample image can be a frame in the first sample video, and the first motion pose representation label is a sequence of multiple motion poses that can be extracted from the first sample video.
[0116] Taking Figure 1 as an example, the first motion pose representation label can be input into the U-Net of the first multimodal diffusion model after the limb motion features are extracted by PoseNet. The first sample image can be input into the VAE encoder of the first multimodal diffusion model. The image features of the first sample image extracted by CLIP can also be input into the U-Net. The speaker's identity features in the first sample image extracted by the face recognition network can also be input into the U-Net. The VAE decoder of the first multimodal diffusion model outputs the first predicted video.
[0117] In step S602, the first multimodal diffusion model is trained based on the difference between the first predicted video and the first sample video to obtain the second multimodal diffusion model.
[0118] For example, the cross-entropy loss function can be used to calculate the difference between each frame of the first predicted video and the first sample video. With the goal of minimizing the loss function, the network parameters of the first multimodal diffusion model are gradually adjusted in backpropagation. When the preset training objective is reached, such as when the loss function is less than a certain value, the second multimodal diffusion model is reached. At this point, the model has initially learned how to generate full-body motion.
[0119] In step S603, the audio driving sample, the second sample image, and the second motion pose representation label are input into the second multimodal diffusion model to generate the second predicted video.
[0120] It should be noted that the descriptions of the first sample image and the second sample image, the first sample video and the second sample video are used to distinguish the two-stage training process. In practical applications, the two can be images / videos from the same training set or images / videos from different training sets.
[0121] In addition to the visual modality, this step introduces audio-driven samples from the audio modality as input for video generation, in order to learn audio-driven facial motion generation in addition to learning the motion of the whole body.
[0122] Taking Figure 1 as an example, audio-driven samples can be processed by Wav2vector and then input into the U-Net of the second multimodal diffusion model. They can also be combined with facial expression labels and blink frequency signals and input into the U-Net. The VAE decoder of the second multimodal diffusion model outputs the second predicted video.
[0123] The second sample video is a video of the speaker expressing speech content. The second sample image contains at least an image of the speaker's facial region and can be a frame from the second sample video. The second motion pose representation label is a sequence of multiple motion poses that can be extracted from the second sample video. The audio-driven sample can be obtained by extracting audio from the second sample video, and the facial expression label and blink frequency signal can also be extracted from the speaker's face in the second sample video.
[0124] In step S604, the second multimodal diffusion model is trained based on the difference between the second predicted video and the second sample video to obtain the multimodal diffusion model.
[0125] Similarly, the cross-entropy loss function can be used to calculate the difference between each frame of the second predicted video and the second sample video. With the goal of minimizing the loss function, the network parameters of the second multimodal diffusion model are gradually adjusted in backpropagation. When the preset training objective is reached, such as when the loss function is less than a certain value, the multimodal diffusion model is trained and completed. At this point, the model has learned how to generate full-body movements and fine facial movements.
[0126] In other embodiments, in the first and / or second phases of the training process described above, sample videos of different scales can be further used. These videos cover different body ranges, such as the full body range of a standing posture, the upper body range of a close-up sitting posture, and the head range of a speaking posture, so that the model can have better generation performance under reference images of different scales.
[0127] In the solutions provided by the above embodiments of this specification, by combining audio driving signals and text prompts to jointly drive the target speaker in the reference image to generate video, the full body movements of the target speaker can be flexibly controlled, achieving fine control over the generated video. Moreover, it is not restricted by the actions specified by the template and the body range displayed in the reference image, and can generate full body movements from any reference image, no longer limited to the head or upper body, thereby improving the diversity and expressiveness of the generated video.
[0128] Figure 7 is a schematic diagram of the human body video generation device based on multimodal control in the embodiments of this specification. This device can be applied to any device, platform, or device cluster with computing and processing capabilities. The device includes:
[0129] The data acquisition unit 701 is configured to acquire text prompt information, audio driving signal and reference image of target speaker. The text prompt information includes text information for prompting action to target speaker, and the audio driving signal is audio information containing speech content.
[0130] The motion generation unit 702 is configured to generate motion posture representation information of the target speaker based on text prompt information and audio driving signals. The motion posture representation information is used to represent the motion posture of the target speaker.
[0131] The video generation unit 703 is configured to generate a speaking video of the target speaker based on a reference image, an audio driving signal, and motion posture representation information. The speaking video includes the body movements of the target speaker as they express speech content.
[0132] In one embodiment, the motion generation unit 702 is specifically configured to generate a three-dimensional 3D posture sequence of the target speaker based on text prompt information and audio driving signals. The 3D posture sequence is a time series representing the 3D positions of key points of the human body. Based on the 3D posture sequence, motion posture representation information of the target speaker is obtained. The motion posture representation information is a time series representing the two-dimensional 2D positions of key points of the human body.
[0133] In one implementation, the 3D pose sequence is a sequence of motion codes used to characterize the 3D positions of key points on the human body. When the motion generation unit 702 converts the motion pose representation information of the target speaker based on the 3D pose sequence, it is specifically configured to determine the motion pose representation information that matches the 3D pose sequence based on a pre-established relational database. The relational database contains the mapping relationship between the motion codes corresponding to the key points on the human body and the 2D poses.
[0134] In one implementation, the 3D pose sequence is a motion code sequence used to characterize the 3D positions of key points on the human body. The motion generation unit 702 generates a 3D pose sequence of the target speaker based on text prompts and audio driving signals. Specifically, it is configured to extract fusion features based on the audio driving signals and text prompts; determine the code probability distribution corresponding to each moment in the audio driving signal based on the fusion features, where the code probability distribution represents the probability that the target motion code at that moment is any of the motion codes in the codebook, and the codebook contains multiple motion codes; and determine the motion code sequence of the target speaker based on the code probability distribution corresponding to each moment in the audio driving signal.
[0135] In one embodiment, the motion generation unit 702 is specifically configured to extract features from the text prompt information by the text branch of the motion generator to obtain text features; extract features from the audio driving signal by the audio branch of the motion generator to obtain audio features; and generate motion posture representation information of the target speaker based on the fusion result of the text features and audio features.
[0136] In one implementation, the motion generator is trained as follows: Masked motion pose representation labels and text prompt samples are input into the text branch of the motion generator for feature extraction to obtain text feature samples; a first motion pose prediction result is generated based on the text feature samples; the text branch of the motion generator is trained based on the motion pose representation labels and the first motion pose prediction result; fixed text prompts are input into the text branch of the motion generator for feature extraction to obtain fixed text features; audio-driven samples are input into the audio branch of the motion generator for feature extraction to obtain audio feature samples; a second motion pose prediction result is generated based on the fusion result of the fixed text features and the audio feature samples; and the audio branch of the motion generator is trained based on the motion pose representation labels and the second motion pose prediction result.
[0137] In one embodiment, the video generation unit 703 is specifically configured to determine facial motion features based on audio driving signals; determine limb motion features based on motion posture representation information; and generate a speaking video of the target speaker based on a reference image, facial motion features, and limb motion features.
[0138] In one embodiment, the video generation unit 703, when determining facial motion features based on an audio driving signal, is specifically configured to determine facial motion features based on the audio driving signal, facial expression tags, and / or blink frequency signals.
[0139] In one embodiment, the video generation unit 703 is specifically configured to extract the identity features of the target speaker based on the facial image of the target speaker in the reference image; and generate a speaking video of the target speaker based on the reference image, audio driving signal, identity features and motion posture representation information.
[0140] In one embodiment, the video generation unit 703 is specifically configured to generate a speaking video of a target speaker based on a reference image, an audio driving signal, and motion posture representation information using a multimodal diffusion model. The multimodal diffusion model is trained as follows: a first sample image and a first motion posture representation label are input into the first multimodal diffusion model to generate a first predicted video; the first multimodal diffusion model is trained based on the difference between the first predicted video and the first sample video to obtain a second multimodal diffusion model; an audio driving sample, a second sample image, and a second motion posture representation label are input into the second multimodal diffusion model to generate a second predicted video; the second multimodal diffusion model is trained based on the difference between the second predicted video and the second sample video to obtain a multimodal diffusion model.
[0141] This specification also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, it causes the computer to perform the method described in FIG2.
[0142] This specification also provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in FIG2.
[0143] This specification also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the steps of the method described in FIG2.
[0144] Those skilled in the art will recognize that the functions described in the various embodiments disclosed in this specification in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0145] In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0146] The above detailed embodiments further illustrate the purpose, technical solutions, and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above are merely specific implementations of the multiple embodiments disclosed in this specification and are not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of the multiple embodiments disclosed in this specification should be included within the protection scope of the multiple embodiments disclosed in this specification.
Claims
1. A method for generating human body videos based on multimodal control, the method comprising: The system acquires text prompt information, audio driving signals, and a reference image of the target speaker. The text prompt information includes text information for prompting the target speaker to perform actions, and the audio driving signals are audio information containing speech content. Based on the text prompt information and the audio driving signal, motion posture representation information of the target speaker is generated, and the motion posture representation information is used to represent the motion posture of the target speaker; Based on the reference image, the audio driving signal, and the motion posture representation information, a speaking video of the target speaker is generated, the speaking video including the body movements of the target speaker when expressing the speech content.
2. The method according to claim 1, wherein, The step of generating motion posture representation information of the target speaker based on the text prompt information and the audio driving signal includes: Based on the text prompt information and the audio driving signal, a three-dimensional 3D posture sequence of the target speaker is generated, wherein the 3D posture sequence is a time series representing the 3D position of key points of the human body. Based on the 3D pose sequence, the motion pose representation information of the target speaker is obtained, which is a time series representing the two-dimensional 2D position of the key points of the human body.
3. The method according to claim 2, wherein, The 3D pose sequence is a motion code sequence used to represent the 3D positions of key points on the human body. The process of converting the 3D pose sequence to obtain the motion pose representation information of the target speaker includes: Based on a pre-established relational database, motion posture representation information matching the 3D posture sequence is determined. The relational database contains the mapping relationship between motion codes corresponding to human key points and 2D postures.
4. The method according to claim 2, wherein, The 3D pose sequence is a motion code sequence used to represent the 3D positions of key points on the human body. Generating the 3D pose sequence of the target speaker based on the text prompt information and the audio driving signal includes: Based on the audio driving signal and the text prompt information, the fusion features are extracted; Based on the fusion features, the code probability distribution corresponding to each moment in the audio driving signal is determined. The code probability distribution is used to represent the probability that the target motion code corresponding to that moment is any of the motion codes in the codebook. The codebook contains multiple motion codes. Based on the code probability distribution corresponding to each moment in the audio driving signal, the motion code sequence of the target speaker is determined.
5. The method according to claim 1, wherein, The step of generating motion posture representation information of the target speaker based on the text prompt information and the audio driving signal includes: The text prompt information is feature extracted from the text branch of the motion generator to obtain text features; The audio features are obtained by extracting features from the audio driving signal using the audio branch of the motion generator; Based on the fusion result of the text features and the audio features, motion posture representation information of the target speaker is generated.
6. The method according to claim 5, wherein, The motion generator is trained in the following way: The masked motion pose representation labels and text prompt samples are input into the text branch of the motion generator for feature extraction to obtain text feature samples. Based on the text feature samples, a first motion pose prediction result is generated; Based on the motion pose representation label and the first motion pose prediction result, the text branch of the motion generator is trained; The fixed text prompt is input into the text branch of the motion generator for feature extraction to obtain the fixed text features; The audio driving sample is input into the audio branch of the motion generator for feature extraction to obtain the audio feature sample. Based on the fusion result of the fixed text features and the audio feature samples, a second motion pose prediction result is generated; The audio branch of the motion generator is trained based on the motion posture representation label and the second motion posture prediction result.
7. The method according to claim 1, wherein, The step of generating the speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: Based on the audio driving signal, facial motion features are determined; Based on the motion posture representation information, limb movement characteristics are determined; Based on the reference image, the facial motion features, and the limb motion features, a speaking video of the target speaker is generated.
8. The method according to claim 7, wherein, The step of determining facial motion features based on the audio driving signal includes: The facial motion features are determined based on the audio driving signal, facial expression tags, and / or blink frequency signals.
9. The method according to claim 1, wherein, The step of generating the speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: Based on the facial image of the target speaker in the reference image, the identity features of the target speaker are extracted; Based on the reference image, the audio driving signal, the identity features, and the motion posture representation information, a speaking video of the target speaker is generated.
10. The method according to claim 1, wherein, The step of generating the speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information includes: The target speaker's speaking video is generated using a multimodal diffusion model based on the reference image, the audio driving signal, and the motion pose representation information; the multimodal diffusion model is trained in the following manner: The first sample image and the first motion pose representation label are input into the first multimodal diffusion model to generate the first predicted video; Based on the difference between the first predicted video and the first sample video, the first multimodal diffusion model is trained to obtain the second multimodal diffusion model; The audio-driven samples, the second sample image, and the second motion pose representation label are input into the second multimodal diffusion model to generate the second predicted video. Based on the difference between the second predicted video and the second sample video, the second multimodal diffusion model is trained to obtain the multimodal diffusion model.
11. A human body video generation device based on multimodal control, the device comprising: The data acquisition unit is configured to acquire text prompt information, audio driving signals, and a reference image of the target speaker. The text prompt information includes text information for prompting the target speaker to perform actions, and the audio driving signal is audio information containing speech content. The motion generation unit is configured to generate motion posture representation information of the target speaker based on the text prompt information and the audio driving signal, wherein the motion posture representation information is used to represent the motion posture of the target speaker; The video generation unit is configured to generate a speaking video of the target speaker based on the reference image, the audio driving signal, and the motion posture representation information, wherein the speaking video includes the body movements of the target speaker as they express the speech content.
12. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the steps of the method according to any one of claims 1-10.