Cooperative voice gesture video generation method based on motion mask guided two-stage network
Through a two-stage network based on motion mask guidance, the problem of synchronous generation of facial expressions and gesture videos in the prior art is solved, and high-quality full-body gesture video generation is realized, ensuring the natural and synchronousness of facial and gesture movements, and improving the authenticity and control accuracy of video generation.
Patent Information
- Application Number
- CN202510388864.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to generate natural facial expressions and gesture videos simultaneously, and the existing data set lacks strong correlation between whole-body movements and speech, which limits the model's expansion ability.
A two-stage network based on motion mask guidance, including a spatial mask guiding audio pose generation network and a motion mask layered audio attention mechanism, is adopted to extract initial posture key points from the audio signal, generate a complete posture sequence containing facial expressions and gesture actions, and generate high-quality full-body gesture video synchronized with speech through dynamic enhancement.
It realizes natural and synchronous control of facial and gesture actions, generates coherent and realistic simultaneous animations, improving the realism generated by video and precise control of output.
Smart Images

Figure CN120375855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and artificial intelligence, and specifically relates to a collaborative speech-gesture video generation method based on a motion mask-guided two-stage network. Background Art
[0002] With the rapid development of virtual reality, digital human technology, and human-computer interaction, speech-driven video generation technology has become a research hotspot. Existing methods are mainly divided into two categories: retrieval-based methods and generation model-based methods. Currently, gesture video generation methods can be roughly divided into two categories: retrieval-based methods and generation-based methods. Although retrieval-based methods have produced satisfactory results with limited computational resources, they face obvious limitations, including the inability to synthesize new gestures and the lack of correlation modeling between the speaker's face and audio. Generation-based methods face common challenges: they cannot generate detailed finger and lip movements simultaneously, exhibit obvious artifacts and deformations, and have limited generality.
[0003] In addition, existing datasets (such as LRS2, VoxCeleb) lack strong correlations between full-body movements and speech, restricting the model's scalability. Therefore, there is an urgent need for a solution that can generate natural facial and gesture videos synchronously without additional priors. Summary of the Invention
[0004] To overcome the defects in the prior art, the present invention proposes a two-stage generation framework (MMGT), which introduces a Spatial Mask Guided Audio Pose Generation Network (SMGA) network. This network is carefully designed to implicitly generate a sequence of key points from audio. By integrating a multi-modal video generation framework with an enhanced Motion Masked Hierarchical Audio Attention (MMHAA) mechanism, this method drives a single static image to simultaneously generate dynamic speech-gesture and talking head videos. In addition to generating speech videos only from audio, our method also supports pose-driven generation and joint driving of pose and audio, providing greater flexibility and precise control over the output.
[0005] The technical solution adopted by the present invention to achieve the above object is as follows:
[0006] A collaborative speech-gesture video generation method based on a motion mask-guided two-stage network, comprising the following steps:
[0007] 1) Through the Spatial Mask Guided Audio Pose Generation Network, generate a complete pose sequence including facial expressions and body movements from the initial pose key points extracted from the audio signal source image;
[0008] 2) Generate a motion mask based on the complete pose sequence, and dynamically enhance the pose sequence, motion mask, and input audio signal through the motion mask hierarchical audio attention mechanism to generate a high-quality full-body gesture video synchronized with the speech.
[0009] Step 1) includes the following steps:
[0010] 1.1) Use the pre-trained key point detector E K to extract the initial pose sequence p (0) from the source image I0, i.e., p (0) = E K (I0);
[0011] 1.2) Add Gaussian noise N(0, σ GT ) to the real pose sequence P 2 to obtain the noisy motion representation:
[0012] x' t = P GT + N(0, σ 2 )
[0013] where x' t represents the noisy pose at time step t;
[0014] 1.3) Copy the initial pose p (0) along the time dimension to make it the same size as P GT , and add it to the noisy motion representation x' t , i.e., x t = x' t + R(p (0) ), where R(·) represents replicating along the time dimension N times, and x t represents the noisy N-frame pose sequence:
[0015]
[0016] where C f represents the key points corresponding to head movement, and C b represents the key points corresponding to the whole body;
[0017] 1.4) Construct the key point mask M f corresponding to head movement and the key point mask M b corresponding to body movement respectively, and apply them to x t ;
[0018] 1.5) Process the audio sequence and image features respectively using different learnable encoders to obtain the complete pose sequence including facial expressions and body movements
[0019]
[0020] f a = E a (a)
[0021]
[0022]
[0023] Among them, D SMGA represents the spatial mask-guided audio pose generation network, f a represents the audio feature embedding, f f and f b respectively represent the facial and body motion features, t is the time embedding, E a is the pre-trained audio encoder, E f and E b are the encoders of the facial and body motion features f f and f b respectively, and a is the audio sequence.
[0024] Step 1.4) includes the following steps:
[0025] 1.4.1) Construct the mask M f ∈R N×C×3 : If C ∈ C f , then M f = 1, otherwise M f = 0;
[0026] 1.4.2) Let the mask M b = 1 - M f ;
[0027] 1.4.3) Apply the mask to x t :
[0028]
[0029] Among them, and correspond to the features of the head and body pose sequences respectively, ⊙ represents element-wise multiplication, and C represents the feature dimension.
[0030] Generate the mask using a method of generating a dynamic mask based on key points, that is, re-partition the feature region of the key points x0 = [p 1 , p 2 , …, p N ∈ R N×C×3 as Let C = C f + C b correspond to the face C f and the body Cb , let the hand key point be C h and the lip key point be C l , where C h ∈C b , C l ∈C f , and then generate a motion mask frame by frame according to the positions of different regions, specifically:
[0031] For each frame in the video sequence, convert the normalized key point coordinates to pixel coordinates, calculate the bounding box of the part where the key point is located, construct a binary mask using the bounding box, set the values inside the bounding box to 255 and the rest to 0. The motion mask generated based on the key points includes those for facial motion for lip motion and for hand motion At the same time, combine the facial and hand masks into one mask, denoted as In addition, set the background mask to
[0032] Optimize the spatial mask-guided audio pose generation network using a multi-component loss function. The total loss is:
[0033] L SMGA = λ f L f + λ b L b
[0034] where λ f and λ b are the weighting factors for the head and gesture motion losses respectively. Each loss term for the head and gesture motion is where L f and L b correspond to the losses of the head and body respectively:
[0035] The reconstruction loss is:
[0036]
[0037] where, x t and represent the ground truth and the predicted motion features respectively, and T is the maximum number of time steps in the diffusion model;
[0038] The velocity loss is:
[0039]
[0040] The acceleration loss is expressed as:
[0041]
[0042] Step 2) includes the following steps:
[0043] 2.1) Extract the semantic features of the source image I0 through the CLIP image encoder E clip Capture the latent representation of the image through the autoencoder E vae Extract the audio feature f of the audio a through E a Extract the pose feature Z of the pose video V a through the pose encoder P ; pose
[0044] 2.2) Obtain the masks obtained through the GT video in the training phase wherein represents the mask of the head and gesture, represents the mask of the mouth, represents the background mask;
[0045] 2.3) Use the motion mask hierarchical audio attention mechanism to process different features and masks, align all features along the time axis, and replace V M and V P with the complete pose sequence and to obtain the final full-body gesture video.
[0046] Optimize the motion mask hierarchical audio attention mechanism using the diffusion-based reconstruction loss :
[0047]
[0048] where T is the maximum number of time steps in the diffusion model, and Z t is the latent representation generated at time step T, is Z t the corresponding target latent representation.
[0049] A collaborative speech-gesture video generation system based on a motion mask-guided two-stage network, comprising:
[0050] A spatial mask-guided audio-pose generation network for generating a complete pose sequence including facial expressions and gesture actions from the initial pose key points extracted from the audio signal source image;
[0051] A motion mask hierarchical audio attention mechanism for generating a motion mask based on the complete pose sequence and dynamically enhancing the pose sequence, the motion mask, and the input audio signal to generate a high-quality full-body gesture video synchronized with the speech.
[0052] The spatial mask-guided audio pose generation network includes:
[0053] At least two motion blocks, respectively used to process features related to facial expressions and gesture movements;
[0054] A self-attention module, used to capture the correlation of motion features themselves;
[0055] A cross-attention mechanism, used to align motion features with audio features;
[0056] A feature-level linear modulation layer, used to dynamically adjust the distribution of motion features according to audio features;
[0057] A multi-layer perceptron, used to fuse facial and gesture motion features to generate a coordinated complete pose sequence.
[0058] The motion mask hierarchical audio attention mechanism includes:
[0059] A cross-attention module, used to align audio features with intermediate hidden states;
[0060] A mask region enhancement module, used to use a motion mask to perform local feature enhancement on a specific region;
[0061] A convolutional adapter module, used to refine the enhanced features through residual connections and convolutional layers to generate spatio-temporally consistent video frames.
[0062] The present invention has the following beneficial effects and advantages:
[0063] 1. The present invention proposes a spatial mask-guided audio pose generation network, which is used to generate a pose video and a dynamic mask from a speech signal and an initial pose. This unified design ensures the natural and synchronous control of facial and gesture movements, thus generating coherent and realistic lip-sync animations.
[0064] 2. The present invention introduces a motion mask hierarchical audio attention module for video generation. By combining audio conditions with a motion mask, the model adaptively focuses on specific semantic regions, such as hands, lips, and faces, thus achieving fine detail enhancement.
[0065] 3. The present invention does not require additional priors. During the inference process, only audio and a single source image are input, and the synchronization of lip shapes, gestures, and speech is precisely controlled through a hierarchical attention mechanism. Description of the Drawings
[0066] Figure 1 It is a schematic diagram of the MMGT framework overview, showing a two-stage process and core modules (SMGA network, MM-HAA mechanism);
[0067] Figure 2It is a detailed architecture diagram of the SMGA network, including self-attention, cross-attention, and FiLM layers. Specific implementation manner
[0068] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0069] A collaborative speech-gesture video generation method based on a motion mask-guided two-stage network is divided into two stages, including:
[0070] The first stage (SMGA network): Generate a complete pose sequence synchronized with the audio through a spatial mask-guided audio pose generation network; this pose sequence not only captures the synchronized face but also naturally incorporates gestures. On this basis, we further calculate the maximum bounding box of each region to obtain a motion mask.
[0071] The second stage (MM-HAA mechanism): Use audio to dynamically enhance the details of specific semantic regions, while the pose sequence drives the motion of static images. Combine the motion mask with a hierarchical audio attention mechanism to generate high-quality videos. By introducing this module, our framework can more accurately control the generation of simultaneous interpretation videos, significantly improving the realism and performance fidelity.
[0072] Both the first stage and the second stage adopt a diffusion model framework, where:
[0073] The first stage generates a pose sequence through a denoising diffusion process;
[0074] The second stage generates a video through a latent space diffusion model and performs multi-modal conditional control by combining the motion mask and audio features.
[0075] In the first stage, the inputs are an audio signal and a source image, and the final output is a complete pose sequence including facial expressions and gestures. The core network structure includes:
[0076] Keypoint detector: Convert the input source image into an initial pose sequence (including the coordinate sequence of human body keypoints in the image). This model is a pre-trained model and does not need to be retrained. It should be noted that the generated initial pose sequence needs to be replicated N times along the time dimension (N is the number of frames of the final generated video), and then added to the noisy pose sequence to obtain a noisy N-dimensional pose sequence;
[0077] Self-attention module: Add a masking mechanism to the above-mentioned N-dimensional pose sequence to make the model only focus on the keypoints of the gestures and the face, and ignore other parts. In particular, in the present invention, the gestures and facial keypoints are separated by a mask, encoded by different encoders respectively, and then passed into their respective self-attention modules to model the temporal dependence of the pose sequence;
[0078] FiLM Layer: The input of this step includes the audio sequence encoded by the WavLM encoder and the output of the self-attention module. By extracting the features of the audio signal and the pose sequence, the feature distribution is dynamically adjusted to enhance the audio-pose alignment.
[0079] Cross-Attention Module: The input is the features modulated by the FiLM layer, which are further aligned with the audio features.
[0080] MLP Fusion Module: The input is the gesture and facial key point sequences that have been separately processed in the above process. It coordinates and merges the above features to generate a unified pose sequence.
[0081] In the second stage, the core innovation is the Motion Mask Hierarchical Audio Attention (MM-HAA). The input includes the pose sequence, audio, and source image generated in the first stage, and the output is a high-fidelity, spatio-temporally consistent full-body digital human video. To achieve the above goal, the present invention provides the following steps:
[0082] Step S1, Motion Mask Generation: Extract key points from the pose sequence output in the first stage above. Calculate the bounding boxes of each region according to the key point coordinates to generate hierarchical masks (face, gesture, lip). It should be noted that since the face mask and the gesture mask do not intersect, the present invention combines the face and gesture masks into one mask for subsequent processing.
[0083] Step S2, Motion Feature Encoding: The pose sequence output in the first stage above needs to be encoded by a learnable encoder and then Gaussian noise is added. The source image is encoded by a clip encoder and a vae encoder respectively to obtain semantic information and latent feature representations, which are used to guide the diffusion model to recover the above noise sequence.
[0084] Step S3, Denoising Unet Network: Input the above audio features, the pose features with added noise, the semantic features and latent representations obtained from the image encoding, and sequentially pass through the spatial attention, cross-attention, and temporal attention mechanisms to obtain intermediate hidden states.
[0085] Step S4, MM-HAA Mechanism: Align the above intermediate hidden states with the audio features through the cross-attention mechanism, and then use the above mask to make it focus on the face and gesture features. Finally, refine the details of specific regions through a convolutional adapter.
[0086] Step S5, vae Decoder: Decode the final output of the above denoising network through a pre-trained decoder into a final high-fidelity, spatio-temporally consistent full-body digital human video.
[0087] The method further includes the following training strategies:
[0088] In the first stage, the pose sequence generation is jointly optimized by reconstruction loss, velocity loss, and acceleration loss;
[0089] In the second stage, the video generation quality is optimized by latent space reconstruction loss and cross-modal alignment loss.
[0090] The source image is a single static image, and the generation process does not rely on additional gesture pose sequences or prior information on body movements.
[0091] The system further includes: a video quality evaluation module for evaluating the quality of the generated co-speech gesture video, including the evaluation of gesture quality, facial quality, lip synchronization quality, and overall video quality; a model optimization module for optimizing the model according to the evaluation results.
[0092] Embodiment: Figure 1 A method for generating a digital human video from audio based on artificial intelligence according to the present invention is given, and an MMGT framework is proposed. This framework ensures the synchronous control of facial and hand movements and also compensates for regional details. The input of this framework requires a given audio sequence a = [a0,..., a n and a source image I0 = R 3×W×H , as inputs, where 3 represents the number of channels, and W and H respectively represent the width and height of the image. The framework of the present invention finally generates a synchronous speech digital human video composed of N frames, denoted as The generation process of the synchronous speech digital human video is divided into two stages, as Figure 1 shown, including motion feature generation (the first stage), motion driving, and detail refinement (the second stage). In the first stage, the Spatial Mask Guided Audio Pose Generation Network (SMGA) generates motion features based on the audio a and the initial pose p (0) , including a pose video and a motion mask The pose video is a human skeleton video composed of key points, containing the motion information of human key points; the motion mask video is a binary image, where the key points of the hand or face correspond to pixels with a value of 255 (white) in the figure, and the pixel values of the remaining parts are 0 (black). Then these features are passed to the second stage to assist in the generation of the synchronous speech digital human video. In the second stage, the input image passes through a pre-trained clip encoder and a vae encoder respectively to obtain the text information Z pose of the image and the latent space representation Z img , and at the same time, the pose video obtained in the first stage (during training, the key point video V P extracted from the label video is directly used, and the above-mentioned The feature Z is obtained through encoding by a learnable encoder pose . Next, the framework of the present invention is divided into two parts: a denoising Unet and a ReferenceNet. The denoising Unet structure uses the above-mentioned encoded feature Z pose and Z text , while the ReferenceNet integrates Z pose , Z text and Z img to generate the final predicted video Therefore, the inference process of the present invention can be formulated as
[0093]
[0094] where E k (·) represents an initial extractor for extracting the motion feature of the initial pose p (0) from the source image I0. The above formula indicates that in the first stage, an SMGA(·) network with the Diffusion Transformer (DIT) architecture as the core uses the audio a as a condition to generate a complete pose sequence (corresponding to the pose video ). In the second stage, it is a video generation model G f (·) with the diffusion model architecture as the core, which uses the motion mask generated in the first stage, the pose video, and the audio as conditions to generate the final digital human video .
[0095] Specifically, in the first stage, the main goal is to generate a pose sequence driven by the audio input. The audio is the main driving signal. The present invention adopts an SMGA network to generate a complete pose sequence. The obtained output sequence is represented as the pose video and the motion mask video for capturing the core motion dynamics. The pose video highlights the key regions with obvious motions, such as the hands and the face, thereby improving the accuracy of the motion representation. Specifically, first, a pre-trained key point detector (E K ) is used to extract the initial pose sequence p (0) from the source image I0, that is, p (0) =E K (I0), where p (0) represents the static pose key point coordinate sequence extracted from the object in the source image. We add Gaussian noise to the real pose sequence P GT to obtain a noisy motion representation
[0096] x t ′ = P GT+N(0, σ 2 )
[0097] where x t ′ represents the noisy pose at time step t. To keep the input size consistent, the present invention repeats the initial pose p (0) along the time dimension to make it the same size as P GT and adds it to the noisy motion representation x t ′, i.e., x t = x t ′ + R(p (0) ), where R(·) means replicating along the time dimension into N copies, thus obtaining a noisy N-frame pose sequence. Next, to improve the generation quality of the hand and face, the motion features are divided into two parts: facial expression motion features and gesture motion features. The noisy N-frame pose sequence is represented as
[0098]
[0099] where C f corresponds to the key points related to head movement, while C b corresponds to the key points of gesture motion. To apply a mask to x t along the feature dimension C, the present invention constructs the mask M f as follows: if C ∈ C f , then M f = 1, otherwise M f = 0, that is, only the key points of head movement are retained after masking, and the hand key points are ignored. Next, let M b = 1 - M f to obtain a complementary mask M b representing only focusing on the key points related to gesture motion and ignoring the head key points. Then these masks are applied to the input x t to separate the feature representations corresponding to the head and gestures, which can be specifically represented as where and correspond to the features of the head and gesture pose sequences respectively, ⊙ represents element-wise multiplication, and then different learnable encoders (E a , E f and E b ) are used to process these input conditions and convert them into separate feature embeddings f a = E a (a), where f a represents the audio feature embedding, E a is a pre-trained audio encoder, and E f and E b are the facial and body motion features f fand f b The encoder of. Finally, the output of the first stage is denoted as
[0100]
[0101] where t is the temporal embedding, and is related to f f , f b , f a These features are passed to the SMGA network D SMGA , D SMGA which decodes them to generate an audio-driven pose sequence
[0102] In terms of mask generation, the present invention introduces a new method for generating dynamic masks based on key points, denoted as where C p = C f + C h + C l corresponding to the face, gesture, and lips respectively. For each frame in the video sequence, the algorithm of the present invention converts the normalized key point coordinates into pixel coordinates and calculates the bounding box of the part where the key points are located. Then, these bounding boxes are used to construct binary masks, setting the values inside the bounding boxes to 255 and the rest to 0. The generated masks include for limb motion and for gesture It should be noted that since the head and hand masks do not overlap, they are combined into one mask, denoted as In addition, we set the background mask to
[0103] The SMGA network aims to capture the temporal and semantic relationships between motion features and audio inputs. As Figure 2 shown, it includes several key components: The Motion block uses self-attention, Feature Linear Modulation (FiLM), and Cross-attention to process motion features, and processes the input features f f or f b representing different body parts respectively. Specifically, the self-attention mechanism simulates temporal dependence to capture the correlation of motion features themselves, and the calculation formula is
[0104]
[0105] where Q f , K f , V f are obtained from the motion features (f f or f b) is derived. Cross-attention can align motion features with audio to ensure synchronization between speech and gestures. Due to the processing of the dynamic mask on x t the cross-attention mechanisms in different Motion blocks will capture different correlations between different body parts and audio embeddings. Through this MASK-guided separate attention mechanism, the present invention realizes coherent and coordinated speech-action generation in different semantic regions (body and face). The specific formula is
[0106]
[0107] where Q f comes from the motion feature (f f or f b ), comes from the audio feature f a . To ensure the coordination of head movement and gestures, the network adopts a fusion strategy. The FiLM layer dynamically adjusts the motion features according to the audio embedding to promote precise cross-modal alignment. To ensure the consistency between the head and gesture movements in the generated video, an MLP layer is used to merge these features, and finally, a FiLM layer dynamically adjusts the combined features. This process is mathematically expressed as
[0108] f merged = FiLM(MLP(f f + f b ))
[0109] where f f and f b represent the head and gesture movement features respectively, and the final output dynamic feature map f merged represents the predicted pose sequence highly aligned with the audio, for predicting the pose sequence This sequence corresponds to the pose sequence generated by audio driving. Through this multi-step integration, the model can generate a pose sequence highly consistent with the input audio and motion features, and is temporally coherent and semantically rich.
[0110] In the training process of the first stage, a multi-component loss function is used to optimize the SMGA network to ensure precise motion generation. The reconstruction loss is defined as where x t and represent the ground truth and the predicted motion feature respectively. The velocity loss is The acceleration loss is expressed as Each loss term for head and gesture movements can be defined as where L f and L bLosses corresponding to the head and gestures respectively. More details of the reconstruction, velocity, and acceleration terms can be extended according to the specific metrics used for evaluation. The total loss is calculated as L SMGA = λ f L f + λ b L b , where λ f = 3 and λ b = 1 are the weighting factors for the head and gesture motion losses. These weights control the relative contribution of each motion type to the overall optimization process.
[0111] In the second stage, the main framework is a diffusion model, which uses motion features, latent space features, and audio generated from the source image to refine the motion details. By integrating spatial attention (SA) for texture details, temporal attention (TA) for smooth transitions, and motion mask-based audio attention (MM-HAA) for aligning audio with the motion mask in the Denoising Unet and ReferenceNet, the ordered combination of these modules ensures temporal and spatial consistency in the generated video.
[0112] The second stage is shown in the Figure 1 lower half. The CLIP image encoder E clip extracts semantic features from the source image I0, while the autoencoder E vae captures the latent representation of the image. For the audio a, the present invention uses E a to extract the corresponding audio feature f a . For the pose video V P , a pose encoder is used to extract the pose feature, denoted as Z pose . In addition, in the training stage, we obtain from the GT video in the formula represents the mask of the head and gestures, represents the mask of the mouth, represents the background mask, which consists of a list of masks of different sizes generated by Gaussian blur and resizing operations. In this way, the model can generate mouth postures synchronized with the speech. Then these masks are input into the model. All features are aligned along the time axis to ensure that the generated video is synchronized with the audio. In the inference stage, we replace V M and V P with the output of the first stage, i.e.: and
[0113] The motion mask hierarchical audio attention module (MM-HAA) module is a key component in the second stage of the framework of the present invention, as shown in Figure 1As shown on the right, MM-HAA optimizes the detailed features of V by aligning the audio feature f a and the cross-attention embedding Z CA to refine the motion features by dynamically aligning the spatial details, motion masks, and audio features. Different from the fixed mask method adopted in previous works, MM-HAA utilizes the motion masks and These masks can adaptively focus attention on specific regions focusing on the face and gestures focusing on lip movements focusing on the background. In addition, MM-HAA uses a cross-attention mechanism to input the intermediate hidden state Z CA and the specific temporal and frequency audio characteristics f a into the cross-attention mechanism to ensure spatio-temporal synchronization. This key operation can be expressed as
[0114]
[0115] where and are the query matrix, key matrix, and value matrix respectively, and d k represents the scaling factor for stability. MM-HAA further integrates these aligned features with the motion masks to enhance the refinement of specific regions. The calculation formulas for the hidden state and the mask are and
[0116] where ⊙ represents element-wise multiplication as defined above. Subsequently, Z′ f+h and Z′ l are processed through a convolutional-based adapter module
[0117] Z MM-HAA = Adapter f+ h(Z f ′ + h)+ Adapter l (Z l ′)+ Adapter b (Z′ b )
[0118] where Adapter f+h , Adapter l , and Adapter b are the convolutional-based refinement modules of and respectively, and the output Z MM-HAAIt is passed to the Temporal Attention (TA) layer in the Denoising Unet, enabling further temporal refinement. The hierarchical design of MM-HAA ensures the precise integration of motion masks, audio features, and spatial details across frames. By combining cross-attention and mask-based refinement, this module aligns audio-driven dynamics with spatially local regions, effectively addressing the limitations of fixed mask methods. The model uses a diffusion-based reconstruction loss for optimization, where T represents the number of time steps, Z t is the latent representation generated at time step T, and is the corresponding target latent representation. Through the decoder D Vae decodes the refined latent representation Z t to reconstruct the final video The decoder D Vae converts the latent space back to the actual video frames to obtain the final digital human video.
[0119] The basic principles of the present invention are described above. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. Additionally, the specific details of the above embodiments are only for illustrative and easy-to-understand purposes, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details for implementation.
[0120] In the embodiments provided by the present invention, it should be understood that the devices and methods used can be implemented in other ways. Additionally, in each embodiment of the present invention, the functional modules can be integrated into one processing unit, or each unit can exist separately, or two or more units can be integrated into one unit.
[0121] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to encompass all changes within the meaning and scope of the equivalent elements of the claims in the present invention.
[0122] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.
Claims
1. A collaborative speech-gesture video generation method based on a motion mask-guided two-stage network, characterized in that Including the following steps: 1) Guided by a spatial mask, an audio pose generation network extracts initial pose key points from an audio signal source image and generates a complete pose sequence including facial expressions and body movements; 2) Based on the complete pose sequence, a motion mask is generated, and a motion mask hierarchical audio attention mechanism is used to dynamically enhance the pose sequence, the motion mask, and the input audio signal to generate a high-quality full-body gesture video synchronized with the speech.
2. The collaborative speech-gesture video generation method based on a motion mask-guided two-stage network according to claim 1, wherein The step 1) includes the following steps: 1.1) Use the pre-trained key point detector E K Extract the initial pose sequence p from the source image I0 (0) , that is, p (0) = E K (I0); 1.2) Add Gaussian noise N(0, σ GT ) to the true pose sequence P 2 to obtain a noisy motion representation: x ′ t = P GT + N(0, σ 2 ) where x ′ t represents the noisy attitude at time step t; 1.3) Copy the initial pose p (0) along the time dimension to make it the same size as P GT and add it to the noisy motion representation x ′ t such that x t = x ′ t + R(p (0) ), where R(·) means copying along the time dimension N times, and x t represents the N-frame pose sequence with noise: Among them, C f represents the key points corresponding to the head movement, and C b represents the key points corresponding to the whole body; 1.4) Construct the key-point mask M corresponding to the head movement f and the key-point mask M corresponding to the body movement b respectively, and apply them to x t ; 1.5) Process the audio sequence and image features using different learnable encoders respectively to obtain a complete pose sequence containing facial expressions and body movements f a = E a (a) Among them, D SMGA represents the spatial mask-guided audio pose generation network, f a represents the audio feature embedding, f f and f b respectively represent the facial and body motion features, t is the time embedding, E a is the pre-trained audio encoder, E f and E b are the encoders of the facial and body motion features f f and f b respectively, and a is the audio sequence.
3. A collaborative speech-gesture video generation method based on a two-stage network guided by a motion mask according to claim 2, wherein The step 1.4) includes the following steps: 1.4.1) Construct the mask M f ∈R N×C×3 : If C ∈ C f , then M f = 1, otherwise M f = 0; 1.4.2) Let the mask M b = 1 - M f ; 1.4.3) Apply the mask to x t : Among them, and correspond to the features of the head and body pose sequences respectively. ⊙ represents element-wise multiplication, and C represents the feature dimension.
4. A collaborative speech gesture video generation method based on a two-stage network guided by a motion mask according to claim 3, wherein Generate a mask using a method of generating a dynamic mask based on key points, that is, re - divide the key points x0 = [p 1 , p 2 , …, p N ∈ R N×C×3 . The feature region is such that C = C f + C b corresponds to the face C f and the body C b respectively. Let the hand key points be C h and the lip key points be C l , where C h ∈ C b , C l ∈ C f . Then, according to the positions of different regions, generate a moving mask frame by frame. Specifically: For each frame in the video sequence, convert the normalized keypoint coordinates to pixel coordinates, calculate the bounding box of the part where the keypoints are located, construct a binary mask using the bounding box, set the value inside the bounding box to 255 and the rest to 0. The motion mask generated based on the keypoints includes for facial motion for lip motion and for hand motion At the same time, combine the facial and hand masks into one mask, denoted as 5. A method for generating collaborative speech gesture videos based on a two-stage network guided by a motion mask according to claim 2, characterized in that Using a multi-component loss function to optimize the spatial mask-guided audio pose generation network, the total loss is: L SMGA = λ f L f + λ b L b Among them, λ f and λ b are the weighting factors for the losses of the head and gesture movements, respectively. Each loss term for the head and gesture movements is where L f and L b correspond to the losses of the head and body, respectively: The reconstruction loss is: where x t and represent the true value and the predicted motion feature respectively, and T is the maximum number of time steps in the diffusion model; The velocity loss is: The acceleration loss is expressed as:
6. A collaborative speech-gesture video generation method based on a motion mask-guided two-stage network according to claim 1, wherein The step 2) includes the following steps: 2.1) Extract the semantic features of the source image I0 through the CLIP image encoder E clip Capture the latent representation of the image through the autoencoder E vae Extract the audio features f of the audio a through E a Extract the audio features f of the audio a through E a Extract the pose features Z of the pose video V through the pose encoder P ; pose ; 2.2) Obtain the mask obtained from the GT video during the training phase Among them, represents the mask of the head and gestures, represents the mask of the mouth, represents the background mask; 2.3) Process different features and masks using a motion mask hierarchical audio attention mechanism, align all features along the time axis, and replace V M and V P with the complete pose sequence and to obtain the final full-body gesture video.
7. A method for generating collaborative speech gesture videos based on a two-stage network guided by a motion mask according to claim 6, wherein Use diffusion-based reconstruction loss Optimize the motion mask hierarchical audio attention mechanism: Among them, T is the maximum number of time steps in the diffusion model, and Z t is the latent representation generated at time step T, and t is the corresponding target latent representation.
8. A collaborative speech-gesture video generation system based on a motion mask-guided two-stage network, characterized in that, Including: A spatial mask-guided audio pose generation network, which is used to extract initial pose key points from an audio signal source image and generate a complete pose sequence including facial expressions and gesture movements; A motion mask hierarchical audio attention mechanism, which is used to generate a motion mask based on the complete pose sequence and dynamically enhance the pose sequence, the motion mask, and the input audio signal to generate a high-quality full-body gesture video synchronized with the speech.
9. A collaborative speech-gesture video generation system based on a motion mask-guided two-stage network according to claim 8, characterized in that, The spatial mask-guided audio pose generation network includes: At least two motion blocks, which are respectively used to process features related to facial expressions and gesture movements; A self-attention module, which is used to capture the correlation of motion features themselves; A cross-attention mechanism, which is used to align motion features with audio features; A feature-level linear modulation layer, which dynamically adjusts the distribution of motion features according to audio features; A multi-layer perceptron, which is used to fuse facial and gesture motion features to generate a coordinated complete pose sequence.
10. A collaborative speech gesture video generation system based on a motion mask-guided two-stage network according to claim 8, characterized in that, The motion mask hierarchical audio attention mechanism includes: A cross-attention module, which is used to align audio features with intermediate hidden states; A mask region enhancement module, which is used to locally enhance features of specific regions using a motion mask; A convolutional adapter module, which is used to refine the enhanced features through residual connections and convolutional layers to generate spatio-temporally consistent video frames.
Citation Information
Cited By
Multi-modal time sequence fusion voice drive gesture generation method
CN121214502A
Speech enhancement method and system
CN121214961A