Human body video generation method and device based on multi-modal control

By combining audio driving signals and text prompt information to generate multimodal driving signals, the problem of difficulty in generating natural and smooth full-body moving videos in the prior art is solved, and the fine control and diversity of generated videos are achieved.

CN120034707APending Publication Date: 2025-05-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510238205.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to generate natural and smooth full-body motion videos, especially in scenes where complex emotions and fine movements are required, the effect is not satisfactory.

Method used

By combining audio driving signals and text prompt information, a multimodal driving signal is generated to achieve fine control of human body video generation, including generating videos that move the whole body.

Benefits of technology

The fine control of the generated video is achieved, and the full body movement can be generated from any reference image, no longer limited to the head or upper body, thereby improving the diversity and expressiveness of the generated video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034707A_ABST
    Figure CN120034707A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a human body video generation method and device based on multi-modal control. The method comprises the steps that text prompt information, an audio driving signal and a reference image of a target speaker are acquired, the text prompt information comprises text information used for carrying out action prompt on the target speaker, and the audio driving signal is audio information comprising voice content; on the basis of the text prompt information and the audio driving signal, motion posture representation information of the target speaker is generated, and the motion posture representation information is used for representing the motion posture of the target speaker; and based on the reference image, the audio driving signal and the motion posture representation information, generating a speaking video of the target speaker, the speaking video including body motion when the target speaker expresses the voice content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of artificial intelligence technology, and specifically, to a method and device for generating human body videos based on multimodal control. Background Art

[0002] With the development of video generation technology, given an audio signal, driving the target speaker in the reference image to speak in order to generate a full-body motion video synchronized with the audio has become an important research direction. Existing methods mainly rely on audio input and predefined action templates to drive the target speaker in the reference image to generate a video.

[0003] However, due to the limitation of templates, most existing methods can only generate animations of the target speaker's head or upper body, and these methods still have obvious limitations in generating natural and smooth full-body movements, especially in scenes that need to express complex emotions and fine movements. Therefore, it is of great significance to develop a video generation system that can flexibly generate full-body movements. Summary of the invention

[0004] The embodiments of this specification provide a human body video generation solution based on multi-modal control, which can generate human body videos through multi-modal driving signals and achieve fine control of human body movements in the generated video.

[0005] In a first aspect, an embodiment of the present specification provides a method for generating a human body video based on multimodal control, comprising: obtaining text prompt information, an audio drive signal, and a reference image of a target speaker, wherein the text prompt information includes text information for providing action prompts to the target speaker, and the audio drive signal is audio information including voice content; based on the text prompt information and the audio drive signal, generating motion posture representation information of the target speaker, wherein the motion posture representation information is used to represent the motion posture of the target speaker; based on the reference image, the audio drive signal, and the motion posture representation information, generating a speaking video of the target speaker, wherein the speaking video includes body movements of the target speaker when expressing the voice content.

[0006] In some embodiments, the generating the motion posture representation information of the target speaker based on the text prompt information and the audio drive signal includes: generating a three-dimensional 3D posture sequence of the target speaker based on the text prompt information and the audio drive signal, the 3D posture sequence being a time series representing the 3D positions of key points of the human body; and converting the motion posture representation information of the target speaker based on the 3D posture sequence, the motion posture representation information being a time series representing the two-dimensional 2D positions of the key points of the human body.

[0007] In some embodiments, the 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of a human body, and the conversion based on the 3D posture sequence to obtain the motion posture representation information of the target speaker includes: determining the motion posture representation information matching the 3D posture sequence based on a pre-established relationship library, wherein the relationship library contains a mapping relationship between the motion code corresponding to the key points of the human body and the 2D posture.

[0008] In some embodiments, the 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of a human body, and generating the 3D posture sequence of the target speaker based on the text prompt information and the audio drive signal includes: extracting fusion features based on the audio drive signal and the text prompt information; determining the code probability distribution corresponding to each moment in the audio drive signal based on the fusion features, the code probability distribution being used to represent the probability that the target motion code corresponding to the moment is each motion code in a code book, and the code book contains multiple motion codes; determining the motion code sequence of the target speaker based on the code probability distribution corresponding to each moment in the audio drive signal.

[0009] In some embodiments, the generating of the motion posture representation information of the target speaker based on the text prompt information and the audio drive signal includes: performing feature extraction on the text prompt information by a text branch of the motion generator to obtain text features; performing feature extraction on the audio drive signal by an audio branch of the motion generator to obtain audio features; and generating the motion posture representation information of the target speaker based on a fusion result of the text features and the audio features.

[0010] In some embodiments, the motion generator is trained in the following manner: inputting masked motion posture representation labels and text prompt samples into the text branch of the motion generator for feature extraction to obtain text feature samples; generating a first motion posture prediction result based on the text feature samples; training the text branch of the motion generator based on the motion posture representation label and the first motion posture prediction result; inputting fixed text prompts into the text branch of the motion generator for feature extraction to obtain fixed text features; inputting audio drive samples into the audio branch of the motion generator for feature extraction to obtain audio feature samples; generating a second motion posture prediction result based on the fusion result of the fixed text features and the audio feature samples; and training the audio branch of the motion generator based on the motion posture representation label and the second motion posture prediction result.

[0011] In some embodiments, generating the speaking video of the target speaker based on the reference image, the audio drive signal and the motion posture representation information includes: determining facial motion features based on the audio drive signal; determining limb motion features based on the motion posture representation information; generating the speaking video of the target speaker based on the reference image, the facial motion features and the limb motion features.

[0012] In some embodiments, determining facial movement features based on the audio drive signal includes: determining the facial movement features based on the audio drive signal, facial expression tags and / or blinking frequency signals.

[0013] In some embodiments, generating a speaking video of the target speaker based on the reference image, the audio drive signal, and the motion posture representation information includes: extracting identity features of the target speaker based on the facial image of the target speaker in the reference image; and generating a speaking video of the target speaker based on the reference image, the audio drive signal, the identity features, and the motion posture representation information.

[0014] In some embodiments, the generating of the speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information comprises: generating the speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information by a multimodal diffusion model; the multimodal diffusion model is trained in the following manner: inputting a first sample image and a first motion posture representation label into a first multimodal diffusion model to generate a first predicted video; training the first multimodal diffusion model based on the difference between the first predicted video and the first sample video to obtain a second multimodal diffusion model; inputting an audio driving sample, a second sample image and a second motion posture representation label into the second multimodal diffusion model to generate a second predicted video; training the second multimodal diffusion model based on the difference between the second predicted video and the second sample video to obtain the multimodal diffusion model.

[0015] In a second aspect, an embodiment of the present specification provides a human body video generation device based on multimodal control, comprising: a data acquisition unit, configured to acquire text prompt information, an audio drive signal, and a reference image of a target speaker, wherein the text prompt information includes text information for providing action prompts to the target speaker, and the audio drive signal is audio information including voice content; a motion generation unit, configured to generate motion posture representation information of the target speaker based on the text prompt information and the audio drive signal, wherein the motion posture representation information is used to represent the motion posture of the target speaker; and a video generation unit, configured to generate a speaking video of the target speaker based on the reference image, the audio drive signal, and the motion posture representation information, wherein the speaking video includes body movements of the target speaker when expressing the voice content.

[0016] In a third aspect, an embodiment of the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any implementation manner in the first aspect is implemented.

[0017] In the solution provided in the above-mentioned embodiments of the present specification, by combining the audio driving signal and the text prompt information, the target speaker in the reference image is jointly driven to generate a video, and the whole body movement of the target speaker can be flexibly controlled, thereby achieving fine control over the generated video. Moreover, it is not restricted by the movements specified by the template and the body range displayed in the reference image, and the whole body movement can be generated from any reference image, and is no longer limited to the head or upper body, thereby improving the diversity and expressiveness of the generated video. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 This is a schematic diagram of the process of generating a human body video in an embodiment of this specification;

[0020] Figure 2 is a flow chart of a method for generating a human body video based on multimodal control in an embodiment of this specification;

[0021] Figure 3 is the network structure of the motion generator in the embodiment of this specification;

[0022] Figure 4is a flow chart of the training steps of the motion generator in an embodiment of this specification;

[0023] Figure 5 It is the network structure of the audio branch and the dual input branch in the motion generator in the embodiment of this specification;

[0024] Figure 6 is a flowchart of the training steps of the multimodal diffusion model in the embodiments of this specification;

[0025] Figure 7 It is a structural schematic diagram of a human body video generation device based on multimodal control in an embodiment of this specification. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0027] As mentioned above, in the current field of speaker video generation, given an audio driving signal and a reference image of the speaker, the goal is to generate a speaker speech video that matches the audio content, so that the speaker's mouth shape is synchronized with the audio. The generated video mainly focuses on the head and shoulder area, especially the dynamic changes of the lips and facial expressions. In order to improve the expressiveness of the speaker, subsequent research has extended the animation generation to the upper body, so that the speaker's gestures and body language can be displayed. These methods perform well in formal occasions (such as news broadcasts or speeches) because in these scenes, the generated body movements are usually limited to the upper body and mainly focus on gestures, with almost no involvement of the lower body. However, this approach has certain limitations. For example, in occasions such as talk shows or crosstalk performances, in order to impress the audience with infectious language, the actor's full body movements are indispensable. In addition, since humans rarely remain still when speaking, in order to make the generated video more realistic, even if a half-body image is used as a reference, it is necessary to generate full-body movements to more realistically restore the natural posture when speaking. For example, although only the upper body of the speaker is presented in the video, the shaking of the lower body will still affect the movement of the upper body.

[0028] Although the audio driving signal can accurately determine the movement of the speaker's lips, it is difficult to link the audio with specific body movements, that is, it is difficult to achieve precise control of whole-body movements based on audio alone. Taking this problem into consideration, the embodiments of this specification propose at least one method for generating human body videos based on multimodal control. In addition to the audio driving signal, text prompt information is also introduced as a control signal to jointly drive the target speaker in the reference image for video generation. While controlling the speaker's whole-body movement to be synchronized with the given audio driving signal, the text prompt is used to flexibly control the specific movements of the target speaker's whole body, thereby achieving fine control of the generated video. Moreover, it is not restricted by the movements specified by the template and the body range displayed in the reference image, and can generate whole-body movements from reference images of any size (such as 1 / 4, 1 / 2 size upper body images or full-body portraits), so that the target speaker can convey richer information through body language, thereby improving the diversity and expressiveness of the generated video.

[0029] Next, combine Figure 1 , introduces the process of human video generation based on multimodal control. Figure 1 The method for generating a human body video based on multimodal control in this embodiment can be executed by any device, platform or device cluster with computing and processing capabilities, including steps S201-S203 as shown below.

[0030] like Figure 2 As shown, in step S201, text prompt information, an audio driving signal, and a reference image of a target speaker are obtained.

[0031] Among them, the text prompt information includes text information for prompting the target speaker to perform actions. For example, the text prompt information can be text containing a specified action, and exemplarily, it can be "the target speaker turns in circles"; the text prompt information can also include the specified action and the time information of the action. For example, it can be "the target speaker waves his right hand to say hello at the beginning of the video, and waves his hands goodbye at the end of the video", or "the target speaker waves his hands at the 10th second". The text prompt information can also include a description of the emotion conveyed by the action, and exemplarily, it can be "the target speaker is very excited". The audio drive signal is audio information containing voice content. For example, the audio drive signal can be a recording of a human speech, or it can be an audio of singing. This embodiment does not limit the specific content of the text prompt information and the audio drive signal.

[0032] In practice, the target speaker may refer to any object that can express speaking actions, and may be a real person, a virtual character or a digital person, or other species. The reference image of the target speaker includes the face area of ​​the target speaker. This embodiment does not limit the body range covered by the reference image. For example, it may include the head and shoulder area of ​​the target speaker, the area above the waist, or the area of ​​the whole body.

[0033] This embodiment does not limit the method for obtaining text prompt information, audio drive signal and reference image of the target speaker. For example, the text prompt information, audio drive signal and reference image specified, edited or uploaded by the user may be obtained.

[0034] Next, in step S202, the movement gesture representation information of the target speaker is generated based on the text prompt information and the audio drive signal.

[0035] The motion posture representation information is used to represent the motion posture of the target speaker, and the motion posture can be synchronized with the audio driving signal, for example, it can be a time series of the same length as the audio driving signal. The motion posture representation information can include the movement information of the target speaker's body in space at different times, can include the specific movements of the limbs, can also include the movement of the head, and can specifically include the coordinate positions of the joints and various parts of the body in a 2D (2-Dimensional) plane or a 3D (3-Dimensional) space, the angles of each joint, the degree of rotation or bending between the joints, and the combination of relative positions between various body parts.

[0036] When generating the target speaker's motion posture representation information, the target speaker's motion posture can be determined based on the text prompt information of the target speaker's motion prompt and the speech rhythm and content in the audio driving signal. For example, faster speech is usually associated with a more excited or nervous posture, which may be accompanied by more frequent body movements, such as waving and quickening pace. A pleasant tone may be accompanied by relaxed body movements, while an angry tone may be accompanied by body tension.

[0037] Exemplarily, the motion posture representation information can be a 2D posture sequence, specifically a time series that represents the positions of key points of the human body in 2D space, or a 3D posture sequence, specifically a time series that represents the positions of key points of the human body in 3D space, wherein the key points of the human body are used to represent the position of each joint or body part, and the position combination of the key points of the human body can describe different human motion postures.

[0038] It should be noted that the selection and setting of human key points can be determined according to actual application requirements. For example, when computing resources are limited, a smaller number of human key points can be used, and when computing resources are sufficient, a larger number of human key points can be used to make the generated movements more refined. In this embodiment, human key points that represent the whole body movements of the human body are used so that the generated speaking video contains the whole body movements of the target speaker. It can be understood that in other embodiments, the body parts to be represented can be flexibly selected according to different requirements, and human key points that represent partial human body movements, such as key points of the upper body, can be used to generate animations or video effects that better meet actual needs.

[0039] In practice, a motion generator may be used to process the text prompt information and the audio driving signal to generate the motion gesture representation information of the target speaker. This embodiment does not limit the network structure of the specific motion generator used.

[0040] In an example, the motion generator may have two input branches, namely a text branch (Textbrunch) and an audio branch (Audio brunch). The text branch of the motion generator performs feature extraction on the text prompt information to obtain text features; the audio branch of the motion generator performs feature extraction on the audio drive signal to obtain audio features; based on the fusion result of the text features and the audio features, the motion posture representation information of the target speaker is generated.

[0041] Exemplarily, the motion generator can adopt a transformer network structure with two input branches to effectively capture the long-range dependencies between text prompt information and audio driving signals through the self-attention mechanism, and generate the motion posture of the target speaker based on the comprehensive information of text and audio. When fusing text features and audio features, the audio features and text features of each transformer layer of the audio branch and the text branch can be extracted and then fused layer by layer, or the audio features and text features of the last layer output by the audio branch and the text branch can be fused to obtain the fusion result, and then the motion posture representation information is generated through the fusion result. The combination of text and audio information can help generate more realistic and consistent action performance.

[0042] Combine the following Figure 3 The network structure of the motion generator shown in FIG. 1 illustrates the training process of the motion generator. As an implementation method, the motion generator can be trained through the following steps S401-S407:

[0043] like Figure 4As shown, in step S401, masked motion tokens and text prompt samples are input into the text branch of the motion generator for feature extraction to obtain text feature samples.

[0044] In this embodiment, the text branch is trained first, then the weight parameters in the trained text branch are frozen, and the audio branch is trained. The motion gesture representation label is a sequence of multiple motion gestures, and the mask processing can be performed by replacing some motion gestures in the sequence with special masks to randomly cover the motion gestures in the sequence.

[0045] For example, for a sample video, the motion posture of the speaker in each frame of the sample video can be extracted to obtain the motion posture representation label Among them, m t It represents the motion posture of the speaker in the t-th frame, the length of the sequence is T, and the motion posture of some frames is randomly masked, and the semantic information is extracted according to the action of the target speaker in the sample video to obtain the text prompt sample. Assuming that the speaker in the sample video is dancing, the text prompt sample is "the speaker is dancing".

[0046] For the text branch, the training goal is to predict the masked motion gesture based on the given text prompt sample, so that the text branch can learn the mapping from text to motion gesture. Specifically, CLIP (Contrastive Language-Image Pretraining) can be used to extract the initial features of the text sample in the text prompt sample and input it into the text branch to obtain the text feature sample.

[0047] In step S402, a first motion posture prediction result is generated based on the text feature sample.

[0048] Then, the text branch predicts the masked motion posture based on the text feature sample to obtain the first motion posture prediction result

[0049] In step S403, the text branch of the motion generator is trained based on the motion gesture representation label and the first motion gesture prediction result.

[0050] Specifically, the motion posture representation label and the first motion posture prediction result can be used to calculate the loss function, and the text branch can be trained with the goal of minimizing the loss function.

[0051] For example, the training objective can be to maximize the log-likelihood of the first motion posture prediction result To minimize the loss function, It is expressed as follows:

[0052]

[0053] Among them, [mask] t Indicates whether the t-th motion posture is masked. If so, the value is 0, otherwise it is 1, to ensure that the model only needs to predict the masked motion posture instead of the entire sequence.

[0054] The loss function can be expressed as follows:

[0055]

[0056] When the loss function reaches the preset minimum value, or the training reaches the preset number of iterations, the training of the text branch ends and the training of the audio branch begins.

[0057] Next, in step S404, the fixed text prompt is input into the text branch of the motion generator for feature extraction to obtain fixed text features.

[0058] In this stage, the training model predicts motion posture based on audio input and corresponding text. Here, a fixed text input (such as "a person is giving a speech") can be given as the control signal of the text modality.

[0059] In step S405, the audio driving sample is input into the audio branch of the motion generator for feature extraction to obtain an audio feature sample.

[0060] The audio driving sample can be obtained by extracting the audio in the sample video. The audio driving sample can be first processed by wav2vector (waveform to vector) to obtain Audio tokens (initial features of audio samples), and then the Audio tokens are input into the audio branch for feature extraction to obtain audio feature samples.

[0061] In step S406, a second motion posture prediction result is generated based on the fusion result of the fixed text feature and the audio feature sample.

[0062] For example, after the self-attention layer of each transformer in the audio branch and the text branch extracts the audio feature samples and fixed text features of the layer, they are fused layer by layer, and based on the fusion result of the last layer, the second motion posture prediction result is predicted.

[0063] In step S407, the audio branch of the motion generator is trained based on the motion gesture representation label and the second motion gesture prediction result.

[0064] Specifically, the motion posture representation label and the second motion posture prediction result can be used to calculate the loss function, and the audio branch can be trained with the goal of minimizing the loss function. In this process, the weight parameters of the text branch are frozen, that is, the weight parameters of the text branch are not gradient updated during the back propagation process.

[0065] For example, the loss function can use the negative log-likelihood function as shown below:

[0066]

[0067] Among them, m i represents the i-th motion posture, a i Represents the corresponding Audio token, text is a fixed text prompt, p φ (m i ∣a i ,text) is a given i When the text is used, the motion posture m is predicted. i The probability of , represented by the parameter φ. The smaller the negative log likelihood, the closer the model's prediction is to the actual value. The model is trained by minimizing this loss so that the action predicted by the model is as close as possible to the actual action.

[0068] In the above training process, training the text branch first and freezing it means that there is no need to repeatedly train the weight parameters of the text branch, which can greatly reduce the computational complexity. In addition, in multimodal learning tasks, by keeping the parameters of the text modality fixed, it can help the model better focus on the learning of other modalities, avoid overfitting, and accelerate the training process.

[0069] As an implementation method, considering that 3D posture is highly controllable, it can provide more accurate information for controlling the movement of human bones and joints, which is convenient for users to control and edit human movements, and is easy to semantically associate with audio drive signals and text prompt information to better model human movements, while 2D posture contains less detailed information about human movements, such as the lack of key angle information. It is difficult to generate 2D posture through audio drive signals and text prompt information to describe fine human movements. In addition, considering that the correspondence between 2D posture and image pixels is clearer, it can be directly used to guide the generation of details of each pixel during rendering, ensuring that the rendered image looks more natural and accurate. Therefore, this step can generate a 3D posture sequence of the target speaker based on the text prompt information and the audio drive signal, and convert the motion posture representation information of the target speaker based on the 3D posture sequence.

[0070] like Figure 1As shown, the motion generator can use text prompt information and audio drive signals to generate a 3D pose sequence, and then convert it into a 2D pose sequence through a motion translator, so as to use 2D poses to guide pixel generation in the rendering stage. The text prompt information can be first processed by CLIP to obtain the initial text features and then input into the motion generator, and the audio drive signal can also be first processed by wav2vector to obtain the initial audio features and then input into the motion generator. In this example, the motion generator is a neural network that outputs a 3D pose sequence.

[0071] Among them, the 3D posture sequence is a time series that characterizes the 3D positions of key points of the human body. For example, the 3D posture (also called Motion tokens) at each time point in the 3D posture sequence can be represented by the 3D position parameters of the key points of the human body. The position parameters of these key points include the specific coordinates of each joint in the three-dimensional space. For example, the key points of the human body may include joints such as the head, shoulders, elbows, knees, and ankles. The 3D position of each joint can be represented by three coordinates of x, y, and z to form a complete 3D posture. The combination of 3D position parameters of different joints can accurately describe every movement of the target speaker in space. Through these position parameters, the details of the rotation, movement, bending and other movements of the target speaker can be obtained. Exemplarily, SMPL (Skinned Multi-Person Linear Model) can be used to model the shape and posture of the human body. The 3D geometric shape of the human body is represented by linear mixed skinning, so that various postures of the human body can be represented. These 3D postures are controlled by a set of joint parameters, which can be recorded as Among them, J is the number of joints, and each joint uses three parameters to represent its coordinates on the x, y, and z axes respectively.

[0072] Next, after obtaining the 3D posture sequence, it is converted into a time series (i.e., a 2D posture sequence) representing the 2D positions of the key points of the human body for video generation, wherein the 2D motion posture at each time point can be represented by the 2D positions of the key points of the human body.

[0073] This embodiment does not restrict the specific conversion method used for 3D posture to 2D posture. For example, the coordinate information of the key points of the human body in the 3D space at each time point can be projected into the 2D plane to obtain the corresponding 2D motion posture. Specifically, the 3D posture includes the x, y, and z coordinates of each key point of the human body in the 3D space, while the 2D posture converts these 3D coordinates into 2D coordinates (x and y coordinates). In this conversion process, the 3D motion posture at each time point contains a series of three-dimensional positions of the key points of the human body. These positions are converted into corresponding 2D coordinates after projection or mapping. In this way, a 2D posture sequence with time as the dimension can be obtained, in which the 2D coordinates (x, y) of each key point of the human body reflects its position in the image plane. Such a 2D posture sequence can better guide the generation of pixels in the image rendering stage, and because its position is converted from an accurate 3D position, it can ensure the accuracy of human body movements.

[0074] As an implementation method, in order to reduce the computational complexity, motion codes can be used as vectorized discrete representations of the original 3D human motion. In this case, the 3D posture sequence is a motion code sequence used to characterize the 3D positions of key points of the human body, that is, the 3D posture at each time point is represented by a corresponding motion code. Assume that the 3D motion of the human body at the t-th frame (or the t-th time point) is represented by m t Indicates that m t Contains key information such as the position, rotation and speed of each joint of the human body. The 3D motion sequence of the human body in continuous time steps can be expressed as It is understandable that for a 3D human motion sequence, the changes in motion are usually continuous, which will cause the data contained in the 3D human motion sequence to be very complex and the subsequent amount of calculation will be very large. This embodiment maps the continuous 3D human motion to a finite discrete space (codebook). In this way, complex motion data can be represented in a finite, discrete code space, thereby simplifying data processing and reducing computational complexity.

[0075] Considering that directly projecting a 3D posture to obtain a 2D posture will result in stiff and unnatural human body movements, in order to solve this problem, the present embodiment may also construct a relationship library to link the motion code in the codebook with the 2D posture, so that the converted 2D posture is more realistic. In the motion converter, the motion posture representation information matching the 3D posture sequence may be determined based on the pre-established relationship library.

[0076] Among them, the relationship library contains the mapping relationship between the motion codes corresponding to the key points of the human body and the 2D postures. By querying the 2D posture corresponding to each motion code in the motion code sequence in the relationship library, the motion posture representation information composed of multiple 2D postures can be obtained.

[0077] The following first explains the generation of the motion code sequence, and then explains the process of establishing the relationship library.

[0078] First, a VQ-VAE (Vector Quantized Variational AutoEncoder) can be trained as a 3D Human Motion Tokenizer to learn and quantize the discrete representation of human 3D motion, thereby mapping continuous human 3D motion to discrete motion codes.

[0079] Specifically, VQ-VAE uses a learnable codebook of size K The autoencoder is used to reconstruct the original human 3D motion sequence, where each motion code Corresponding to a discretized 3D motion of the human body, d c Represents the dimension of the motion code. Given an original human 3D motion sequence The encoder ε maps it to a latent feature sequence in l is the temporal downsampling rate of ε. For each latent feature z i , by selecting the closest motion code in codebook C to quantize it, thus obtaining a quantized motion code sequence As shown below:

[0080]

[0081] in, Yes i The quantized version of the decoder in VQ-VAE The quantized motion code sequence can be Reconstructing 3D human motion sequences

[0082] For example, a codebook of size 512 can be constructed through the above process. The codebook contains 512 motion codes, and different motion codes can be used to combine into motion code sequences to express the coherent motion of the human body.

[0083] This embodiment does not limit the method for training VQ-VAE. For example, VQ-VAE can be trained by combining motion reconstruction loss with potential embedding loss of quantization layer. The loss function is as follows:

[0084]

[0085] in, represents the motion reconstruction loss, which uses the L1 norm to weigh the input original human 3D motion sequence M and the output reconstructed by the model The difference between represents the latent embedding loss, which uses the L2 norm to measure Z and The difference between them, sg[·] indicates stopping the gradient operation, β is the weighting coefficient, and the codebook can be updated by the exponential moving average method.

[0086] After the 3D human motion segmenter training is completed, the obtained codebook can be used to generate a motion code sequence. Exemplarily, the process of generating a motion code sequence may include the following steps:

[0087] Firstly, based on the audio driving signal and text prompt information, the fusion features are extracted.

[0088] For example, one can use Figure 5 The motion generator of the network structure shown is used to extract fusion features. Figure 5 The network structure of the audio branch for converting audio to motion (Audio to motion) (left figure) and the network structure of the two input branches driven by multimodal control signals (right figure) are shown respectively.

[0089] For ease of understanding, the following first describes the process of extracting audio features from the audio branch and predicting the 3D pose sequence, and then describes the process of extracting fusion features through the network structure of the two input branches.

[0090] In practice, the audio branch transformer can be used Predict the motion code sequence based on the audio driving signal. Specifically, the audio driving signal can be processed to obtain the initial audio features Audio tokens, which are then input into The output of each transformer self-attention layer is collected and represented as in represents the output of the sth layer, S can be set to 8. Then, through a simple linear layer, the audio features output by the last layer are transformed into Convert to audio code probability Where T and K represent the time length of the human 3D motion sequence and the size of the codebook, respectively (for example, when the codebook contains 512 motion codes, K is 512). audioThe code probability distribution of the motion code corresponding to each time t in is as follows:

[0091]

[0092] For each time t, a matching motion code is selected from the codebook according to the code probability distribution. For example, the one with the largest probability in formula (3) is selected. The corresponding c k ), get the motion code sequence

[0093] It should be noted that if the motion code is not used as the vectorized discrete representation of the original human 3D motion, the previously trained VQ-VAE decoder can also be used. Convert motion code sequences into human 3D motion sequences as a sequence of 3D poses.

[0094] The following is an explanation of the network structure of the two input branches. The network structure of the two input branches can be regarded as an audio branch network structure with a text branch transformer added to it. The text branch is used to predict motion code sequences based on text. The network structure used to convert text into motion (text to motion) is the same as that of the audio branch. Specifically, CLIP can be used to extract text prompt information to obtain the initial text features, and then the text prompt information can be input into In the encoder of , the same as the processing of the audio branch, the output of each transformer self-attention layer is collected and represented as in represents the output of the sth layer, such as Figure 5 As shown, it can be The encoder and The text features and audio features of the output of each transformer self-attention layer in the encoder are fused, and the final fused features are obtained in the last layer.

[0095] Then, based on the fusion features, the code probability distribution corresponding to each moment t in the audio driving signal is determined. The code probability distribution can be referred to in formula (6), which is used to represent the probability that the target motion code corresponding to the moment t is each motion code in the codebook, and based on the code probability distribution corresponding to each moment in the audio driving signal, the motion code sequence of the target speaker is determined. Here, the process of determining the matching motion code sequence according to the code probability distribution is the same as that of the aforementioned audio branch, and will not be repeated.

[0096] In this way, the fusion features of text prompt information and audio driving signals can be extracted through the network structure of two input branches, and discrete motion code sequences can be generated. By processing audio and text information at the same time, information from different modalities can be effectively fused. This multimodal learning helps to improve the model's understanding and generation capabilities of complex tasks, especially when it comes to the joint representation of audio and text, which can better capture the relationship between the two and generate matching motion code sequences.

[0097] After obtaining the motion code sequences, the next step is to map them to 2D pose sequences through a relation library. The relation library can be obtained by extracting the 2D pose and motion code from the template video and establishing a mapping relationship between the two.

[0098] Specifically, the SMPL-X (extended version of SMPL) data and 2D posture of the template video can be extracted, and the SMPL-X data (i.e., the original human 3D motion) can be converted into motion codes using the previously trained 3D human motion marker. A mapping relationship between the motion code and the 2D posture corresponding to each video frame is established and saved in a relationship library. In this way, by extracting the 2D postures of the human body in the real world from the template video and aligning them with the motion codes in the codebook, a large number of motion code-2D posture pairs can be created, so that the 2D posture obtained by the motion code conversion can ensure that the generated motion is natural and smooth.

[0099] Next, in step S203, a speaking video of the target speaker is generated based on the reference image, the audio driving signal and the motion posture representation information, wherein the speaking video includes the body movement of the target speaker when expressing the speech content.

[0100] After obtaining the motion posture representation information representing the whole body movement, the motion representation can be converted into an explicit posture sequence, and combined with the audio drive signal and the reference image to generate the final video with accurate lip synchronization, rich gestures and whole body movement. Among them, the reference image is mainly used to indicate the generation of the character image of the target speaker, and the audio drive signal is mainly used to indicate the lip shape of the target speaker. In addition, both the reference image and the audio drive signal can further guide the generation of the facial expression of the target speaker. For example, according to the speech rhythm and emotional changes of the audio, the changes in facial expressions (such as smiling, frowning, mouth movements, etc.) are adjusted.

[0101] Exemplarily, when the reference image is an image of the upper body of the target speaker, the generated speaking video may be a video of the upper body of the target speaker, which includes not only the target speaker's lip shape, gestures, and facial expressions synchronized with the audio drive signal, but also the target speaker's body movements driven by text and audio. Since the positions of the key points of the whole body are taken into account when generating the motion gesture representation information in this embodiment, the generated target speaker will be accompanied by more natural body movements during the speaking process, such as the tilt or shaking of the upper body caused by the movement of the lower body.

[0102] Specifically, this step can use a deep learning model to generate a speaking video, such as a generative adversarial network, VAE (Variational Autoencoder) or a video stream generation model, input a reference image, an audio driving signal and motion posture representation information into the model, and output a speaking video of the target speaker.

[0103] In this embodiment, a multimodal diffusion model can be used to generate a speech video, which generates video data by gradually noise-freeing the input data and recovering from the noise. For example, see Figure 1 , the multimodal diffusion model at least includes: a VAE encoder for generating a noise latent variable (Noise Latent) according to a reference image; a U-Net (a convolutional neural network for image segmentation) encoder for encoding input reference image features, motion posture representation information and audio drive signals to obtain a feature vector, wherein the reference image features can be obtained by extracting features from the reference image using CLIP; a U-Net decoder for decoding the feature vector; and a VAE decoder for decoding the features output by the U-Net decoder to generate a speaking video. It is understandable that in other embodiments, the multimodal diffusion model may also adopt other types of network structures.

[0104] As an implementation method, in order to avoid the influence of the position information of key points related to the face in the motion posture representation information on facial expressions, in this step, facial motion features can be determined based on the audio driving signal; limb motion features can be determined based on the motion posture representation information; and a speaking video of the target speaker can be generated based on the reference image, facial motion features and limb motion features.

[0105] Among them, the facial motion features are the motion features of the facial area, and the limb motion features are the motion features of other body areas outside the face. In this way, the facial motion features can be generated without being affected by the positions of the facial key points in the motion posture representation information, so that the lip movements and facial expressions in the generated video can be driven by the audio driving signal, making the generated facial expressions more natural.

[0106] For example, Figure 1 In the posture condition part of the conditional synergy effect shown, if the motion posture representation information obtained in the previous step is a whole-body posture sequence including the head, the whole-body posture sequence can be redefined as a composite representation. Specifically, the posture sequence of the body below the neck can be retained and the face can be covered with a head mask, and the center of the mask can be located at the midpoint of the face above the neck.

[0107] Then, the composite representation is mapped through PoseNet to obtain the limb motion features, and added element by element to the output of the first convolutional layer of U-Net, so that the influence of the limb motion features is initially superimposed on the noise latent variable to guide the diffusion of each step, so that it can play a role in each step of diffusion to more accurately generate human body movements. Then, in each U-Net block, the original cross-attention layer ( Figure 1 An additional cross-attention layer ( Figure 1 The 4th layer in the image is dedicated to inputting facial motion features to generate face-related dynamics.

[0108] As an implementation method, in order to further ensure the identity consistency between the speaker in the generated video and the target speaker in the reference image, this embodiment can extract the identity features of the target speaker based on the facial image of the target speaker in the reference image; and then generate a speaking video of the target speaker based on the reference image, the audio driving signal, the identity features and the motion posture representation information.

[0109] The identity feature includes the facial feature information of the target speaker and is used to identify the target speaker. For example, it may include information such as the relative position of facial organs, skin texture features, and the three-dimensional structure of the face.

[0110] Specifically, while extracting reference image features of the reference image, a pre-trained facial recognition network is used to detect facial regions in the reference image and extract identity features, which are then input into a newly introduced additional cross-attention layer to interact with facial motion features to accurately generate facial actions while maintaining the identity consistency of the target speaker.

[0111] In other examples, in addition to the audio driving signal, there may be other control signals to drive facial expressions, such as determining facial movement features based on the audio driving signal, facial expression labels (Expression Labels) and / or blink frequency signals (Blink Radio).

[0112] In this example, in addition to the audio driving signal, conditional signals such as facial expression labels related to facial movements and blinking frequency signals are also used as additional control signals to drive facial expressions.

[0113] Among them, facial expression labels usually refer to labels that classify and mark facial expressions, which are used to represent specific facial expressions, such as smile, anger, sadness, surprise, etc., which help to accurately adjust and reproduce specific facial expressions, making the generated video look more natural and real. Facial expression labels can be detected by emotion detectors, such as Figure 1 As shown in the audio condition section, a facial image can be extracted based on a reference image of the target speaker and the facial image can be input into an emotion detector to detect and obtain a facial expression label. For another example, facial images or facial videos of other speakers can be input into the emotion detector to obtain a sequence of facial expression labels.

[0114] Blink frequency refers to the frequency or speed of eye blinking, which is usually used to describe the number of eye blinks in a period of time. It can be extracted from other speakers' speech videos or set manually. Setting the blink frequency can make the target speaker's facial expression more natural and realistic.

[0115] like Figure 1 The audio condition part of the conditional synergy shown in the figure can directly use the above-mentioned conditional signals related to facial movement as additional control signals and input them into the newly added cross-attention layer in U-Net ( Figure 1 In the 4th layer of the encoder, these signals interact with the identity features to accurately generate facial actions such as lip movements, facial expressions, and eye blinks while maintaining the consistency of the target speaker's identity.

[0116] The training method of the multimodal diffusion model is described below. It can be understood that in this embodiment, the motion generator and the multimodal diffusion model can be trained separately. In other embodiments, an end-to-end training method can also be adopted to train the motion generator and the multimodal diffusion model at the same time.

[0117] Considering that the input signals of the multimodal diffusion model come from multiple different modes, in order to ensure the quality and stability of the generated video, this embodiment trains the multimodal diffusion model through a two-stage training method, each stage focusing on the tasks of different modes, so that the model can gradually learn the representation and generation capabilities of different modes, such as Figure 6 As shown, the training process specifically includes steps S601-S604.

[0118] In step S601, a first sample image and a first motion posture representation label are input into a first multimodal diffusion model to generate a first predicted video.

[0119] First, based on the input of a first sample image of a visual modality and a first motion posture representation label, video generation is performed to learn the motion of the whole body.

[0120] Among them, the first sample video is a video of a speaker expressing voice content, the first sample image at least includes an image of the speaker's facial area, the first sample image can be a frame in the first sample video, and the first motion posture representation label is a sequence of multiple motion postures, which can be extracted from the first sample video.

[0121] by Figure 1 For example, the first motion posture representation label can be input into the U-Net of the first multimodal diffusion model after the limb motion features are extracted by PoseNet, the first sample image can be input into the VAE encoder of the first multimodal diffusion model, and the image features of the first sample image extracted by CLIP can be input into the U-Net, and the identity features of the speaker in the first sample image extracted by the facial recognition network can also be input into the U-Net, and the VAE decoder of the first multimodal diffusion model outputs a first predicted video.

[0122] In step S602, based on the difference between the first predicted video and the first sample video, the first multimodal diffusion model is trained to obtain a second multimodal diffusion model.

[0123] For example, the cross entropy loss function can be used to calculate the difference between each frame of the first predicted video and the first sample video, and with the goal of minimizing the loss function, the network parameters of the first multimodal diffusion model are gradually adjusted in back propagation. When the preset training goal is reached, such as when the loss function is less than a certain value, the second multimodal diffusion model is reached. At this time, the model has preliminarily learned how to generate full body movements.

[0124] In step S603, the audio driving sample, the second sample image and the second motion posture representation label are input into a second multimodal diffusion model to generate a second predicted video.

[0125] It should be noted that the description of the first sample image and the second sample image, the first sample video and the second sample video is to distinguish the two-stage training process. In actual applications, the two can be images / videos in the same training set or images / videos in different training sets.

[0126] In addition to the visual modality, this step additionally introduces audio-driven samples of the audio modality as input for video generation, so as to learn audio-driven facial motion generation in addition to learning the motion of the entire body.

[0127] by Figure 1 For example, the audio driving sample can be processed by Wav2vector and then input into the U-Net of the second multimodal diffusion model. It can also be combined with facial expression labels and blinking frequency signals and input into the U-Net. The VAE decoder of the second multimodal diffusion model outputs the second predicted video.

[0128] The second sample video is a video of a speaker expressing speech content, the second sample image at least includes an image of the speaker's facial area, the second sample image may be a frame in the second sample video, and the second motion gesture representation label is a sequence of multiple motion gestures, which may be extracted from the second sample video. The audio-driven sample may be obtained by extracting audio from the second sample video, and the facial expression label and the blinking frequency signal may also be extracted from the speaker's face in the second sample video.

[0129] In step S604, based on the difference between the second predicted video and the second sample video, the second multimodal diffusion model is trained to obtain a multimodal diffusion model.

[0130] Similarly, the cross entropy loss function can be used to calculate the difference between each frame of the second predicted video and the second sample video, and with the goal of minimizing the loss function, the network parameters of the second multimodal diffusion model are gradually adjusted in back propagation. When the preset training goal is reached, for example, the loss function is less than a certain value, the trained multimodal diffusion model is achieved. At this time, the model has learned how to generate full body movements and fine facial movements.

[0131] In other embodiments, in the first stage and / or the second stage of the above-mentioned training process, sample videos of different scales can be further used. These videos cover different body ranges, such as the full body range of a standing posture, the upper body range of a close-up sitting posture, and the head range of a speaking posture, so that the model can have better generation performance under reference images of different scales.

[0132] In the solution provided in the above-mentioned embodiments of the present specification, by combining the audio driving signal and the text prompt information, the target speaker in the reference image is jointly driven to generate a video, and the whole body movement of the target speaker can be flexibly controlled, thereby achieving fine control over the generated video. Moreover, it is not restricted by the movements specified by the template and the body range displayed in the reference image, and the whole body movement can be generated from any reference image, and is no longer limited to the head or upper body, thereby improving the diversity and expressiveness of the generated video.

[0133] Figure 7: is a schematic diagram of the structure of a human body video generation device based on multimodal control in an embodiment of this specification. The device can be applied to any device, platform or device cluster with computing and processing capabilities. The device includes:

[0134] The data acquisition unit 701 is configured to acquire text prompt information, an audio drive signal, and a reference image of a target speaker, wherein the text prompt information includes text information for performing action prompts on the target speaker, and the audio drive signal is audio information including voice content;

[0135] The motion generating unit 702 is configured to generate motion gesture representation information of the target speaker based on the text prompt information and the audio driving signal, wherein the motion gesture representation information is used to represent the motion gesture of the target speaker;

[0136] The video generation unit 703 is configured to generate a speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information, wherein the speaking video includes the body movement of the target speaker when expressing the speech content.

[0137] In one embodiment, the motion generation unit 702 is specifically configured to generate a three-dimensional 3D posture sequence of the target speaker based on the text prompt information and the audio drive signal, where the 3D posture sequence is a time series representing the 3D positions of key points of the human body; based on the 3D posture sequence, the motion posture representation information of the target speaker is converted, where the motion posture representation information is a time series representing the two-dimensional 2D positions of key points of the human body.

[0138] In one embodiment, the 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of the human body. When the motion generation unit 702 converts the motion posture representation information of the target speaker based on the 3D posture sequence, it is specifically configured to determine the motion posture representation information matching the 3D posture sequence based on a pre-established relationship library, and the relationship library contains the mapping relationship between the motion code corresponding to the key points of the human body and the 2D posture.

[0139] In one embodiment, the 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of a human body. The motion generation unit 702 generates a 3D posture sequence of a target speaker based on text prompt information and an audio drive signal, and is specifically configured to extract fusion features based on the audio drive signal and the text prompt information; based on the fusion features, determine the code probability distribution corresponding to each moment in the audio drive signal, the code probability distribution is used to represent the probability that the target motion code corresponding to the moment is each motion code in a code book, and the code book contains multiple motion codes; based on the code probability distribution corresponding to each moment in the audio drive signal, determine the motion code sequence of the target speaker.

[0140] In one embodiment, the motion generation unit 702 is specifically configured to perform feature extraction on the text prompt information by the text branch of the motion generator to obtain text features; perform feature extraction on the audio drive signal by the audio branch of the motion generator to obtain audio features; and generate motion posture representation information of the target speaker based on the fusion result of the text features and the audio features.

[0141] In one embodiment, the motion generator is trained in the following manner: the masked motion posture representation label and text prompt sample are input into the text branch of the motion generator for feature extraction to obtain a text feature sample; based on the text feature sample, a first motion posture prediction result is generated; based on the motion posture representation label and the first motion posture prediction result, the text branch of the motion generator is trained; the fixed text prompt is input into the text branch of the motion generator for feature extraction to obtain a fixed text feature; the audio drive sample is input into the audio branch of the motion generator for feature extraction to obtain an audio feature sample; based on the fusion result of the fixed text feature and the audio feature sample, a second motion posture prediction result is generated; based on the motion posture representation label and the second motion posture prediction result, the audio branch of the motion generator is trained.

[0142] In one embodiment, the video generation unit 703 is specifically configured to determine facial motion features based on an audio drive signal; determine limb motion features based on motion posture representation information; and generate a speaking video of a target speaker based on a reference image, facial motion features, and limb motion features.

[0143] In one embodiment, when the video generation unit 703 determines the facial motion features based on the audio drive signal, it is specifically configured to determine the facial motion features based on the audio drive signal, the facial expression tag and / or the blinking frequency signal.

[0144] In one implementation, the video generation unit 703 is specifically configured to extract the identity features of the target speaker based on the facial image of the target speaker in the reference image; and generate a speaking video of the target speaker based on the reference image, the audio drive signal, the identity features and the motion posture representation information.

[0145] In one embodiment, the video generation unit 703 is specifically configured to generate a speaking video of the target speaker based on a reference image, an audio driving signal and motion posture representation information by a multimodal diffusion model; the multimodal diffusion model is trained in the following manner: a first sample image and a first motion posture representation label are input into a first multimodal diffusion model to generate a first predicted video; based on the difference between the first predicted video and the first sample video, the first multimodal diffusion model is trained to obtain a second multimodal diffusion model; an audio driving sample, a second sample image and a second motion posture representation label are input into a second multimodal diffusion model to generate a second predicted video; based on the difference between the second predicted video and the second sample video, the second multimodal diffusion model is trained to obtain a multimodal diffusion model.

[0146] The present specification also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the following Figure 2 Describe the method.

[0147] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the following is implemented: Figure 2 Describe the method.

[0148] The embodiments of the present specification also provide a computer program product, including a computer program / instruction, which is executed by a processor to implement the following Figure 2 Describe the steps of the method.

[0149] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the multiple embodiments disclosed in this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0150] In some cases, the actions or steps described in the claims may be performed in a different order than in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0151] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above are only specific implementation methods of the multiple embodiments disclosed in this specification, and are not used to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.

Claims

1. A method for generating a human body video based on multimodal control, the method comprising: Acquire text prompt information, an audio drive signal, and a reference image of a target speaker, wherein the text prompt information includes text information for performing action prompts on the target speaker, and the audio drive signal is audio information including voice content; generating, based on the text prompt information and the audio drive signal, movement gesture representation information of the target speaker, wherein the movement gesture representation information is used to represent the movement gesture of the target speaker; A speaking video of the target speaker is generated based on the reference image, the audio driving signal and the motion posture representation information, wherein the speaking video includes body movements of the target speaker when expressing the speech content.

2. The method according to claim 1, wherein: The step of generating the movement gesture representation information of the target speaker based on the text prompt information and the audio drive signal includes: Based on the text prompt information and the audio drive signal, generating a three-dimensional 3D posture sequence of the target speaker, wherein the 3D posture sequence is a time sequence representing the 3D positions of key points of the human body; Based on the 3D posture sequence, the motion posture representation information of the target speaker is converted, and the motion posture representation information is a time series representing the two-dimensional 2D positions of the key points of the human body.

3. The method according to claim 2, wherein: The 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of a human body. The conversion based on the 3D posture sequence to obtain the motion posture representation information of the target speaker includes: Based on a pre-established relationship library, the motion posture representation information matching the 3D posture sequence is determined, wherein the relationship library contains a mapping relationship between the motion code corresponding to the key points of the human body and the 2D posture.

4. The method according to claim 2, wherein: The 3D posture sequence is a motion code sequence for characterizing the 3D positions of key points of a human body. The generating of the 3D posture sequence of the target speaker based on the text prompt information and the audio drive signal includes: Extracting fusion features based on the audio drive signal and the text prompt information; Based on the fusion feature, determine the code probability distribution corresponding to each moment in the audio drive signal, the code probability distribution is used to represent the probability that the target motion code corresponding to the moment is each motion code in the codebook, and the codebook contains multiple motion codes; Based on the code probability distribution corresponding to each moment in the audio drive signal, the motion code sequence of the target speaker is determined.

5. The method according to claim 1, wherein: The step of generating the movement gesture representation information of the target speaker based on the text prompt information and the audio drive signal includes: The text branch of the motion generator performs feature extraction on the text prompt information to obtain text features; The audio branch of the motion generator performs feature extraction on the audio driving signal to obtain audio features; Based on the fusion result of the text feature and the audio feature, the motion gesture representation information of the target speaker is generated.

6. The method according to claim 5, wherein: The motion generator is trained in the following way: Inputting the masked motion gesture representation labels and text prompt samples into the text branch of the motion generator for feature extraction to obtain text feature samples; Based on the text feature sample, generating a first motion posture prediction result; Based on the motion gesture representation label and the first motion gesture prediction result, training a text branch of the motion generator; Inputting the fixed text prompt into the text branch of the motion generator for feature extraction to obtain fixed text features; Input the audio driving sample into the audio branch of the motion generator for feature extraction to obtain an audio feature sample; Generate a second motion posture prediction result based on the fusion result of the fixed text feature and the audio feature sample; Based on the motion posture representation label and the second motion posture prediction result, the audio branch of the motion generator is trained.

7. The method according to claim 1, wherein: The step of generating a speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information comprises: determining facial movement characteristics based on the audio drive signal; Determining limb movement characteristics based on the movement posture representation information; A speaking video of the target speaker is generated based on the reference image, the facial motion features, and the limb motion features.

8. The method according to claim 7, wherein: The step of determining facial movement features based on the audio drive signal comprises: The facial movement feature is determined based on the audio drive signal, the facial expression tag and / or the blink frequency signal.

9. The method according to claim 1, wherein: The step of generating a speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information comprises: Extracting identity features of the target speaker based on the facial image of the target speaker in the reference image; A speaking video of the target speaker is generated based on the reference image, the audio driving signal, the identity feature and the motion posture representation information.

10. The method according to claim 1, wherein: The step of generating a speaking video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information comprises: A multimodal diffusion model is used to generate a speech video of the target speaker based on the reference image, the audio driving signal and the motion posture representation information; the multimodal diffusion model is trained in the following manner: Inputting the first sample image and the first motion posture representation label into a first multimodal diffusion model to generate a first predicted video; Based on the difference between the first predicted video and the first sample video, training the first multimodal diffusion model to obtain a second multimodal diffusion model; Inputting the audio driving sample, the second sample image and the second motion posture representation label into the second multimodal diffusion model to generate a second predicted video; Based on the difference between the second predicted video and the second sample video, the second multimodal diffusion model is trained to obtain the multimodal diffusion model.

11. A human body video generation device based on multimodal control, the device comprising: a data acquisition unit configured to acquire text prompt information, an audio drive signal, and a reference image of a target speaker, wherein the text prompt information includes text information for performing action prompts on the target speaker, and the audio drive signal is audio information including voice content; a motion generating unit, configured to generate motion gesture representation information of the target speaker based on the text prompt information and the audio driving signal, wherein the motion gesture representation information is used to represent the motion gesture of the target speaker; The video generation unit is configured to generate a speaking video of the target speaker based on the reference image, the audio drive signal and the motion posture representation information, wherein the speaking video includes the body movement of the target speaker when expressing the speech content.

12. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the steps of the method described in any one of claims 1 to 10 are implemented.

Citation Information

Cited By

  • Digital human video generation method and device, equipment and storage medium

    CN121217996A