Digital population type driving method and device and electronic equipment

By extracting audio and facial features and using contrastive learning and denoising networks to generate digital lip-print videos, the problem of unnatural digital lip-print generation in existing technologies is solved, and high-quality and consistent digital human video generation is achieved.

CN121334459APending Publication Date: 2026-01-13CHENGDU ZHIPU HUAZHANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511425699.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies for generating two-dimensional digital lip movements based on pure speech modality lack sufficient semantic guidance, resulting in digital humans with unnatural and unvivid lip movements, making it difficult to meet diverse application needs.

Method used

By extracting audio features from audio signals and facial coefficient features from video frames, a contrastive learning feature fusion network is used for probabilistic fusion processing. Combined with a digital human reference image and noise latent variables, a denoising network is used to generate a sequence of digital human mouth shape video frames. Finally, a digital human video is generated based on the audio signal.

Benefits of technology

It improves the quality and naturalness of digital human videos, ensures high consistency between video and audio in terms of time and content, and generates more realistic and natural digital human videos to meet diverse application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121334459A_ABST
    Figure CN121334459A_ABST
Patent Text Reader

Abstract

The invention provides a digital population type driving method and device and electronic equipment, and the method comprises the steps: extracting an audio feature in an audio signal and a face coefficient feature of a video frame in a video; probabilistic introduction fusion processing is carried out on the audio features and the facial coefficient features through a contrast learning feature fusion network, and feature embedding sequence features are obtained; obtaining a digital human reference figure image and a noise hidden variable; processing the feature embedding sequence features, the digital human reference figure image and the noise hidden variables through a denoising network, and outputting a digital population type video frame sequence; a digital human video is generated based on the sequence of digital population video frames and the audio signal. According to the method, the digital human generation effect can be improved, and diversified application requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a digital population driving method, apparatus and electronic device. Background Technology

[0002] In related technologies, two-dimensional digital lip-reading generation techniques based on pure speech modality typically focus on modeling the correspondence between speech content and lip movement content. This results in the generated digital human's lip movements lacking sufficient semantic guidance, making the lip-reading effect unnatural and vivid, leading to poor digital human performance and difficulty in meeting diverse practical application needs. Summary of the Invention

[0003] This application aims to at least partially address one of the technical problems in the related art.

[0004] The first aspect of this application proposes a digital demographic driving method, including:

[0005] Extract audio features from audio signals and facial coefficient features from video frames;

[0006] The audio features and facial coefficient features are probabilistically fused by a contrastive learning feature fusion network to obtain feature embedding sequence features;

[0007] Obtain the digital human reference image and noise latent variables;

[0008] The feature embedding sequence features, the digital human reference image, and noise latent variables are processed by a denoising network to output a digital human mouth shape video frame sequence.

[0009] A digital human video is generated based on the digital lip-print video frame sequence and the audio signal.

[0010] A second aspect of this application provides a digital phrasing driving device, comprising:

[0011] The feature extraction module is used to extract audio features from audio signals and facial coefficient features from video frames in videos;

[0012] The feature fusion module is used to probabilistically introduce and fuse the audio features and the facial coefficient features through a contrastive learning feature fusion network to obtain feature embedding sequence features;

[0013] The acquisition module is used to acquire the digital human reference image and noise latent variables;

[0014] The sequence output module is used to process the feature embedding sequence features, the digital human reference image and noise latent variables through a denoising network, and output a digital human mouth shape video frame sequence.

[0015] The video generation module is used to generate a digital human video based on the digital lip-print video frame sequence and the audio signal.

[0016] A third aspect of this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0017] The memory stores computer-executed instructions;

[0018] The processor executes computer execution instructions stored in the memory to implement the method as described in any of the first aspects above.

[0019] A fourth aspect of this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method described in any of the first aspects above.

[0020] A fifth aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the method as described in any of the first aspects above.

[0021] In the embodiments of this application, audio features are extracted from the audio signal and facial coefficient features from the video frames; a contrastive learning feature fusion network is used to probabilistically introduce and fuse the audio features and facial coefficient features to obtain a feature embedding sequence feature; a digital human reference image and noise latent variables are obtained; a denoising network is used to process the feature embedding sequence feature, the digital human reference image, and the noise latent variables to output a digital human lip-syncing video frame sequence; and a digital human video is generated based on the digital human lip-syncing video frame sequence and the audio signal. This not only effectively improves the generation quality and naturalness of the digital human video but also ensures a high degree of consistency between the video and audio in terms of time and content. Furthermore, through the fusion of multimodal features and advanced denoising processing, more realistic and natural digital human videos can be generated, improving the digital human generation effect and meeting diverse application needs.

[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0023] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0024] Figure 1 A flowchart illustrating a digital physiognomy-driven method provided in an embodiment of this application;

[0025] Figure 2 A schematic diagram illustrating the execution framework of a digital physiognomy-driven method provided in an embodiment of this application;

[0026] Figure 3 A schematic diagram illustrating the execution framework of a digital physiognomy-driven method provided in an embodiment of this application;

[0027] Figure 4 A schematic diagram illustrating an example of a digital physiognomy driving method provided in an embodiment of this application;

[0028] Figure 5 This is a schematic diagram of the structure of a digital phreatic device provided in an embodiment of this application. Detailed Implementation

[0029] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0030] Among related technologies, voice-driven 2D digital human (2D Digital Human) video generation technology aims to convert a portrait image or video of a given person into a speaking video synchronized with the driving voice. This technology belongs to the field of multimodal generation and is widely used in film production, virtual assistants, online education, video conferencing, and other scenarios. Facial expressions are one of the important ways to present digital humans, and there is a certain correlation between the speaker's facial expressions and the speech signal. However, in related technologies, audio signals are usually only mapped to digital lip movements. Since there is only a weak correlation between the two, end-to-end audio-driven lip-reading generation often needs to introduce additional information related to spatial movements, but it is still difficult to obtain the most suitable result, thus making it difficult to generate natural and coherent lip movements under audio-driven conditions.

[0031] For example, in related technologies, lip-sync generation methods typically employ single-modal audioportrait generation, using the audio signal as the sole embedding to guide a diffusion model in generating lip-sync videos. Alternatively, some methods introduce multimodal information as guiding conditions, combining it with audio information to regulate lip-sync generation. For instance, an Emotion MOE (Emotional Modulation Expert) module can be added to extract emotional features from the audio signal for generation control; textual conditions can be introduced to constrain the generation of digital lip movements and expressions. However, these methods introduce problems such as information redundancy and poor interpretability, reducing the ease of use of the model, and the generation process requires the joint regulation of multiple information sources. Simultaneously, person-related information is strongly interfered with by multimodal conditions, easily leading to problems such as blurred facial features and distorted ID information.

[0032] Based on this, this application provides a digital lip-reading driven method. By introducing audio and facial coefficients, facial expression coefficients are incorporated into the model training process. High-order expression features and audio features are alternately compared and jointly participate in model training, thereby enhancing the digital human's ability to represent facial expressions in speech during generation. This allows for the fusion of representational information from different modalities, enabling the generated digital human's facial expressions to be mapped to highly correlated facial coefficients and effectively aligned with the audio signal. This improves the audio signal's representational ability of facial expressions, thus enhancing the correlation between audio and facial expressions. In this way, efficient alignment between speech features and expression features can be achieved, promoting semantic interoperability between the two and effectively improving the naturalness and expressiveness of digital lip-reading generation.

[0033] The following description, with reference to the accompanying drawings, describes a digital phrasing driving method, apparatus, and electronic device according to embodiments of this application.

[0034] Figure 1 This is a flowchart illustrating a digital phrasing-driven method provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0035] S101, extract audio features from the audio signal and facial coefficient features from the video frames.

[0036] In the embodiments of this application, when executing the digital lip-syncing driven method, feature extraction can be performed first. For example, this includes extracting audio features from audio signals. These audio features can reflect information such as rhythm and intonation, and can reflect the dynamic changes in the audio signal, providing a basis for subsequent lip-syncing generation. Simultaneously, facial coefficient features can be extracted from video frames, such as posture features, expression features, and lip movement features. These features can capture the digital human's facial movements and expression changes, ensuring the visual consistency and naturalness of the generated video frames.

[0037] S102, through a contrastive learning feature fusion network, audio features and facial coefficient features are probabilistically introduced into the fusion process to obtain feature embedding sequence features.

[0038] In the embodiments of this application, after extracting audio features and facial coefficient features, the audio features and facial coefficient features can be input into a contrastive learning feature fusion network to fuse the audio features and facial coefficient features. For example, the contrastive learning feature fusion network can effectively integrate features from different modalities (audio features and facial coefficient features) by probabilistically introducing a fusion mechanism to obtain feature embedding sequence features. This provides rich feature representations for subsequent video generation, enabling the generated features to better reflect the correlation between audio and video.

[0039] S103, obtain the digital human reference image and noise latent variables.

[0040] In embodiments of this application, a digital human reference image can also be used to provide the digital human's appearance features, ensuring that the generated digital human visually matches the target person. The reference image can be an actual photograph of the target person or a pre-set digital human model. Furthermore, a noisy latent variable can be introduced to introduce randomness into the generation process, increasing the diversity and naturalness of the generated results. This noisy latent variable can be sampled from a probability distribution, such as a Gaussian distribution, to provide random initial conditions.

[0041] S104 processes the feature embedding sequence features, digital human reference images, and noise latent variables through a denoising network, and outputs a digital human mouth shape video frame sequence.

[0042] In the embodiments of this application, after obtaining the feature embedding sequence features and introducing the digital human reference image and noise latent variables, denoising processing can be performed. For example, a denoising network can be used to process the feature embedding sequence features, the digital human reference image, and the noise latent variables to generate a digital lip-sync video frame sequence. The denoising network can remove noise from the features, improve the quality of the features, and thus generate clearer and more natural video frames. It is understood that the denoising network typically includes an encoder and a decoder; the encoder extracts high-level features, and the decoder gradually recovers low-level features.

[0043] S105 generates digital human video based on digital lip-shaped video frame sequences and audio signals.

[0044] In the embodiments of this application, after obtaining the digital lip-sync video frame sequence, a digital human video can be generated based on the digital lip-sync video frame sequence and audio signal. For example, the final digital human video can be generated based on the digital lip-sync video frame sequence and audio signal through processing such as time alignment and background compositing.

[0045] In the embodiments of this application, audio features are extracted from the audio signal and facial coefficient features from the video frames; a contrastive learning feature fusion network is used to probabilistically introduce and fuse the audio features and facial coefficient features to obtain a feature embedding sequence feature; a digital human reference image and noise latent variables are obtained; a denoising network is used to process the feature embedding sequence feature, the digital human reference image, and the noise latent variables to output a digital human lip-syncing video frame sequence; and a digital human video is generated based on the digital human lip-syncing video frame sequence and the audio signal. This not only effectively improves the generation quality and naturalness of the digital human video but also ensures a high degree of consistency between the video and audio in terms of time and content. Furthermore, through the fusion of multimodal features and advanced denoising processing, more realistic and natural digital human videos can be generated, improving the digital human generation effect and meeting diverse application needs.

[0046] In one possible implementation, extracting audio features from the audio signal and facial coefficient features from video frames includes:

[0047] Audio features are extracted from the audio signal using an audio extractor, and facial coefficient features are extracted from the video frames using a facial coefficient extractor. The facial coefficient features include pose features, expression features, and lip movement features.

[0048] In the embodiments of this application, a specialized audio extractor can be used to extract key audio features from the audio signal. These features include information such as the frequency, intensity, rhythm, and pitch of the audio, which can reflect the dynamic changes of the audio signal and provide a basis for subsequent lip-sync generation. A facial coefficient extractor can be used to extract facial coefficient features from video frames. These features include pose features, expression features, and lip movement features (e.g., pose feature dimension size 21×3, expression feature dimension size 21×3, lip movement feature dimension size 12×3), which can capture the facial movements and expression changes of the digital human, ensuring the visual consistency and naturalness of the generated video frames.

[0049] In one possible implementation, the contrastive learning feature fusion network includes an attention module, a feature fusion layer, a cross-modal multi-head parameter attention layer, and a feature projection activation layer;

[0050] By using a contrastive learning feature fusion network, audio features and facial coefficient features are probabilistically incorporated into the fusion process to obtain feature embedding sequence features, including:

[0051] The audio features and facial coefficient features are preprocessed and aligned to a preset fusion feature space.

[0052] The aligned audio features and facial coefficient features are input into the attention module for feature enhancement processing;

[0053] The enhanced audio features and facial coefficient features are input into the feature fusion layer for flattening and activation processing, as well as layer normalization and linear projection processing.

[0054] The audio features and facial coefficient features processed by the feature fusion layer are input into the cross-modal multi-head parameter attention layer to calculate the similarity score and multi-head attention weight;

[0055] The audio features and facial coefficient features processed by the cross-modal multi-head parameter attention layer are input into the feature projection activation layer for feature recombination, activation and reshaping to obtain the feature embedding sequence features.

[0056] In the embodiments of this application, when probabilistically introducing and fusing audio features and facial coefficient features into the feature embedding sequence features through a contrastive learning feature fusion network, the audio features and facial coefficient features can first be preprocessed, such as normalization and noise reduction, and then aligned to a preset fusion feature space. For example, audio features may need to be aligned with the timestamps of video frames using a time alignment algorithm, and facial coefficient features may need to be aligned with the preset feature space using a spatial alignment algorithm. Then, the aligned audio features and facial coefficient features can be input into the attention module of the contrastive learning feature fusion network. The attention module can dynamically allocate weights by calculating the similarity scores between features, enhancing the representational power of important features while suppressing unimportant features. For example, the attention module can use dot product similarity to calculate the similarity score between query and key, and then normalize these scores into weights using a softmax function. These weights are then multiplied by the value to generate a weighted feature representation. In this way, key information can be highlighted, and the expressive power and robustness of features can be improved.

[0057] Next, the enhanced audio features and facial coefficient features can be input into the feature fusion layer. In the feature fusion layer, a flattening operation can be performed to flatten the feature tensor into a one-dimensional vector; then, a non-linear transformation using an activation function (such as ReLU) is applied to increase the expressive power of the features; next, layer normalization can be performed to stabilize the training process; finally, linear projection can be used to map the features to a new feature space. This makes multimodal features more suitable for subsequent fusion operations. For example, the flattening operation converts the multi-dimensional feature tensor into a one-dimensional vector, the activation function introduces non-linearity, layer normalization ensures the stability of the feature distribution, and linear projection maps the features to a new feature space to better capture the relationships between features. The audio features and facial coefficient features processed by the feature fusion layer can also be input into a cross-modal multi-head parameter attention layer. In the cross-modal multi-head parameter attention layer, the similarity score between the query end (i.e., audio features) and the key end (i.e., facial coefficient features) can be calculated, and multi-head attention weights can be generated. The multi-head attention mechanism can learn the relationships between features from multiple perspectives simultaneously, improving the effect of feature fusion. Specifically, multi-head attention mechanisms can segment features into multiple "heads," each independently calculating a similarity score and attention weights. These results are then combined to generate richer feature representations. This allows for a better capture of the complex relationships between audio and video features, resulting in more comprehensive feature representations.

[0058] Finally, the audio features and facial coefficient features processed by the cross-modal multi-head parameter attention layer can be input into the feature projection activation layer. The feature projection activation layer can perform recombination, activation, and reshaping of the features. For example, recombination can reorganize the features into a format suitable for subsequent generation tasks; activation functions (such as ReLU or Tanh) can further enhance the non-linear expressive power of the features; finally, reshaping can adjust the features to the desired tensor shape, and the resulting feature embedding sequence will serve as input for subsequent generation tasks. As an example, recombination can involve feature concatenation, stacking, etc., activation functions can introduce non-linearity, and reshaping can adjust the features into a format suitable for the generation model input. This not only improves the feature fusion effect but also enhances the quality and naturalness of the generated video, further improving the digital human generation effect.

[0059] In one possible implementation, a denoising network processes the feature embedding sequence features, the digital human reference image, and noise latent variables to output a digital human mouth-printing video frame sequence, including:

[0060] The embedded sequence features, the digital human reference image, and the noise latent variables are input into the denoising network;

[0061] A denoising network is used to perform multimodal feature fusion processing on the feature embedding sequence features, digital human reference images, and noise latent variables to generate enhanced features. The multimodal feature fusion processing includes spatial cross-attention processing, semantic cross-attention processing, and self-attention processing in the temporal dimension.

[0062] The enhanced features are decoded using a denoising network to obtain a sequence of digital phreatic video frames. The decoding process includes recovering low-level features using a skip connection mechanism and upsampling using global semantic information.

[0063] In the embodiments of this application, when processing the feature embedding sequence features, the digital human reference image, and noise latent variables through a denoising network to output a digital lip-sync video frame sequence, the feature embedding sequence features, the digital human reference image, and noise latent variables can be input into the denoising network. The denoising network can generate enhanced features through multimodal feature fusion processing such as spatial cross-attention processing, semantic cross-attention processing, and temporal self-attention processing. For example, the spatial cross-attention mechanism can enhance the representational ability of features in the spatial dimension, ensuring the spatial consistency and accuracy of the generated video frames; the semantic cross-attention mechanism can fuse semantic information from different modalities, enhancing the representational ability of features in the semantic dimension, ensuring the semantic consistency and accuracy of the generated video frames; and the temporal self-attention mechanism can enhance the representational ability of features in the temporal dimension, ensuring the temporal consistency and continuity of the generated video frame sequence. Then, the enhanced features can be decoded using a denoising network. For example, skip connection mechanisms can be used to recover low-level features, improving the quality of the generated video frames. Furthermore, global semantic information processing provides overall semantic support for the upsampling path, ensuring the consistency of the generated video frames in global semantics. This not only improves the effect of feature fusion but also enhances the quality and naturalness of the generated video, resulting in more realistic and natural digital lip-syncing videos. This further improves the digital human generation effect and better meets diverse practical application needs.

[0064] In a further possible implementation, a denoising network is used to perform multimodal feature fusion processing on the feature embedding sequence features, the digital human reference image, and the noise latent variables to generate enhanced features, including:

[0065] By using a denoising network and a cross-attention mechanism, the features of the digital human reference image are combined with the features embedded in the sequence features for spatial cross-attention processing to generate spatially enhanced features.

[0066] By using a denoising network, the spatial information of the reference frame and the noise latent variables are semantically cross-attention processed to generate semantically enhanced features.

[0067] By using a denoising network and employing a self-attention mechanism in the temporal dimension, semantic enhancement features are processed to generate temporally enhanced features.

[0068] In the embodiments of this application, a denoising network is used to perform multimodal feature fusion processing on the feature embedding sequence features, the digital human reference image, and the noise latent variables to generate enhanced features. First, a cross-attention mechanism is employed in the denoising network to perform spatial cross-attention processing on the features of the digital human reference image and the feature embedding sequence features, generating spatially enhanced features. For example, weights can be dynamically allocated by calculating the similarity score between the reference image features and the feature embedding sequence features, enhancing the spatial representation capability of the features. As an example, the cross-attention mechanism can highlight key regions in the reference image, ensuring that the generated video frames are spatially consistent with the reference image. Then, a denoising network can perform semantic cross-attention processing on the spatial information of the reference frame and the noise latent variables to generate semantically enhanced features. For example, weights can be dynamically allocated by calculating the similarity score between the spatial information of the reference frame and the noise latent variables, enhancing the semantic representation capability of the features. As an example, the semantic cross-attention mechanism can fuse the spatial information of the reference frame and the semantic information of the noise latent variables to generate richer semantically enhanced features. Subsequently, a denoising network can be used with a temporal self-attention mechanism to process the semantic enhancement features and generate temporal enhancement features. For example, the similarity score of the semantic enhancement features in the temporal dimension can be calculated, and weights can be dynamically assigned to enhance the representational power of the features in the temporal dimension. As an example, the temporal self-attention mechanism can capture the temporal continuity and consistency of the semantic enhancement features, generating temporal enhancement features and ensuring that the generated video frame sequence has natural transitions and coherence in time. This allows for the generation of more realistic and natural digital human mouth-printing video frame sequences, providing better data support for digital human video generation. This further improves the effect of feature fusion, enhances the quality and naturalness of the generated digital human videos, and ultimately improves the digital human generation effect, better meeting the diverse practical application needs.

[0069] In a further possible implementation, the enhanced features are decoded using a denoising network to obtain a sequence of digital lip-printing video frames, including:

[0070] The intermediate feature map output by the encoder is input into the decoder through the denoising network, and the intermediate feature map is combined with the temporal enhancement feature using a skip connection mechanism to obtain the recovered feature; wherein, the intermediate feature map is obtained based on feature embedding sequence features, spatial enhancement features and semantic enhancement features;

[0071] By processing the bottleneck layer of the denoising network, the recovered features are processed to extract global semantic information. The global semantic information is then combined with the semantic enhancement features to obtain enhanced global semantic features.

[0072] Based on the restored features and enhanced global semantic features, a digital phreatic video frame sequence is obtained.

[0073] In the embodiments of this application, when decoding the enhanced features through a denoising network to obtain a digital lip-sync video frame sequence, the intermediate feature map output by the encoder can be input into the decoder of the denoising network. This intermediate feature map is generated based on feature embedding sequence features, spatial enhancement features, and semantic enhancement features, containing rich, multi-layered information. For example, the decoder can employ a skip connection mechanism to combine the intermediate feature map with temporal enhancement features. This skip connection mechanism helps recover low-level detail information by passing intermediate layer features from the encoder to the corresponding layer of the decoder, thereby generating a higher-quality feature representation. This generates recovered features, which are more complete and accurate in both spatial and temporal dimensions. Then, the recovered features can be processed through the bottleneck layer of the denoising network to extract global semantic information. The bottleneck layer is typically located in the middle part of the network and is responsible for compressing and abstracting the features, extracting the most essential semantic information. For example, the extracted global semantic information can be combined with semantic enhancement features to generate enhanced global semantic features. This can further enrich the semantic representation of the features and ensure the consistency and accuracy of the generated video frames in terms of global semantics.

[0074] Subsequently, a sequence of digital human mouth-printing video frames can be generated based on the restored features and enhanced global semantic features. For example, the upsampling operation of the decoder can be used to progressively restore the feature maps to their original resolution, generating the final video frame sequence. In this way, by combining restored features and enhanced global semantic features, the generated video frames are not only rich in detail but also consistent in global semantics, ensuring the naturalness and coherence of the digital human video. This further enhances the realism and naturalness of the generated digital human video, thus improving its overall effect.

[0075] In some possible implementations, digital human videos are generated based on digital lip-printing video frame sequences and audio signals, including:

[0076] The digital lip-sync video frame sequence is time-aligned with the audio signal to obtain a time-aligned video frame sequence.

[0077] Background synthesis is performed on time-aligned video frame sequences, and digital human images are synthesized into preset background video frames to generate a synthesized video frame sequence.

[0078] The synthesized video frame sequence is combined with the audio signal to generate a digital human video.

[0079] In the embodiments of this application, when generating a digital human video based on a digital lip-syncing video frame sequence and an audio signal, the digital lip-syncing video frame sequence and the audio signal can first be time-aligned to ensure that each video frame and its corresponding audio segment are precisely synchronized in time. For example, a time alignment algorithm can be used to adjust the timestamps of the video frames to make them consistent with the timestamps of the audio signals, ensuring that the lip movements and sounds match when the digital human video and audio are played. Then, background compositing can be performed on the time-aligned video frame sequence, compositing the digital human image into a preset background video frame. For example, image compositing technology can be used to place the digital human image in a suitable background (preset background video frame) to enhance the realism and visual effects of the video, improve the visual quality of the video, make the digital human appear to be moving in a specific scene, and increase the immersion of the video. Afterwards, the composited video frame sequence can be combined with the audio signal to generate the final digital human video. For example, video editing technology can be used to combine the video frame sequence with the audio signal to generate a complete video file, ensuring that the generated digital human video is consistent visually and aurally, guaranteeing the naturalness and coherence of the video.

[0080] To make the digital phrasing-driven method provided in this disclosure clearer, it will be described below with reference to the following examples.

[0081] The execution framework of the human-driven digital intelligence approach is as follows: Figure 2 and Figure 3 As shown, Figure 2 In this context, the following components are used: Reference Frames, Motion Frames, ReferenceNet, Face Region Mask, A2EEmbedding, A2P Embedding, Denoising Unet, Generated Frames, WAV2VEC, Coefficient extractor, ExpPoseLip module, Concentrate CL Module, Self-Attention, Motion Attention, AE-Attention, AP-Attention, AL-Attention, Temporal-Attention, Concat, and NOISE. Figure 3In this context, the components are: Coefficients extractor (facial coefficient extractor), Probabilistic Sampling, ExpPoseLip coeffs, EmbeddingFusion, Learnable Parameter, Linear (linear layer), ParamAttention, Feature Activate Project, A2E / A2P / A2LEmbedding, and Concatenate CL Module.

[0082] In this application, embodiments can replace the audio encoding module before embedding and denoising the U-net with an audio-motion coefficient contrastive learning module to complete the training of digital lip shape generation based on contrastive learning. During data preprocessing, the facial expression coefficient encoder uses Liveportrait as the face coefficient encoder to obtain facial coefficients for each frame (e.g., pose feature dimension 21×3, exp expression feature dimension 21×3, lip lip movement feature dimension 12×3). Audio embedding uses an audio encoding model such as wav2vec to obtain the audio signal embedding (e.g., size 12×768). The contrastive learning encoding block aims to probabilistically introduce expression coefficient features representing facial information and audio signals during training, allowing expression coefficients to fully participate in feature guidance and fusion during training. The introduction probability is then reduced over time, making the entire training feature approximate the audio signal. The introduction probability is fixed in the later stages of training, with the audio introduction probability reaching its maximum value until training is complete. The specific audio-expression coefficient contrastive learning framework is shown in the figure. The extracted audio features and three types of face parameters are input into a Concatenate CL Module, which contains three parallel feature alignment sub-modules (corresponding to pose, expression, and lip movement, respectively). Each sub-module has a consistent structure and is implemented as a ContrastiveLearningFeatureModule. Each input, after random sampling, first passes through a feature fusion and projection layer (Embedding Fusion) to unify the high-dimensional representations of audio and target features. Then, it enters a multi-head cross-modal attention layer (Param Attention), where learnable KV layer parameters are trained to align the input fused features Q. After residual connections and layer normalization, deep fusion of target features guided by audio is achieved. Finally, a linear transformation and sequence recombination activation module (Feature Activate Project) outputs the final feature embedding sequence features for subsequent digital lip shape generation. The module needs to optimize the contrastive similarity loss and joint loss for individual features as follows:

[0083]

[0084] in, For contrastive similarity loss of a single feature (contrast loss of single face coefficient features), These are positive / negative samples, representing the speaker's pose, exp, and lip coefficients, respectively. The positive and negative samples are represented by the audio coefficients, respectively. sim(·) is the similarity calculation function, exp(·) is the exponential calculation, and log(·) is the logarithmic calculation. The core of this calculation method is to maximize the similarity between the expression coefficient and the positive audio sample pair, while minimizing the similarity between the anchor point and the negative sample. The contrastive loss (joint loss) representing the overall loss, λ n The constants set for training (which can be finely adjusted manually according to the actual facial expression task), for example, λ1 = 0.2, λ2 = 0.3, λ3 = 0.5, and τ is a temperature parameter commonly used in contrastive learning.

[0085] After training is complete, the contrastive learning module will be added to replace the original audio projection activation layer. For the two types of digital human-driven tasks, the digital human video lip-sync generation process is as follows: Figure 4 As shown, the following processes are included:

[0086] 1. Preprocessing of audio features and facial expression coefficient features

[0087] This embodiment utilizes a pre-trained contrastive learning feature fusion module to divide the digital lip-reading task into a dual-task flow model. The features extracted by the audio feature extractor and the facial expression coefficient extractor are aligned to a 16×5×12×768 fusion feature space. Calculations are performed on the three types of features respectively, and the results are input into the three newly added attention modules of Unet.

[0088] Two heterogeneous features are input into the feature fusion layer (Embedding Fusion). After flattening and activation, the layer is normalized and linearly projected.

[0089] x audio-merge =reshape(x) audio ,[-1,C×D])

[0090] x feat-merge =reshape(x) feat [-1, features_dim])

[0091] h a =ReLU(x W) a +b a )

[0092] h f =ReLU(x W) f +b f )

[0093]

[0094] Where, x audio With x featThese represent two heterogeneous features, respectively. `reshape` represents the operation of flattening the feature dimensions, and `ReLU` represents the activation function. W a b a Let represent the learnable weight matrix and bias matrix of the heterogeneous features during the activation process, respectively.

[0095] Finally, the query matrix of heterogeneous features is obtained by using the layer normalization operation LayerNorm and the linear layer projection Linear in the feature fusion layer.

[0096] In the cross-modal multi-head parametric attention layer, a learnable parameter matrix P is defined. l Each passed through its own channel containing W k / v b k / v A linear layer with weight and bias matrices is used to compute the key (k) and value (v) sides of the attention mechanism. The resulting attention calculation A is then compared with the audio / expression latent features h. a / f The intermediate hidden layer output feature h is obtained after residual connection and normalization operations. out :

[0097] k = P l ·W k +b k

[0098] v = P l ·W v +b v

[0099]

[0100] h out =LayerNorm(h a / f +A)

[0101] Finally, the Feature Activate Projection layer takes the result h after processing by the Param Attention module and applies it. out Flattening along the feature dimensions, i.e., batch number B, frame number F, window number W, channel number C, and dimension D are each flattened into two dimensions to obtain z. flat The intermediate results z1 and z2 are obtained by applying two ReLU layers and a linear activation layer, respectively. Then, z2 is linearly projected and reshaped to [B×F,M,d]. o Where M is the category dimension of the output feature, d o The embedding dimension of the U-net structure, in this embodiment, can be, for example, 32 or 768, to obtain z. ctxt After normalization, it is reshaped into [B,F,M,d]. oThe shape of the tensor z is obtained, which conforms to the input to the U-net structure. final :

[0102] z flat =reshape(h out ,[B×F,W×C×D])

[0103] z1 = ReLU(z flat W1+b1)

[0104] z2 = ReLU(z1 W2 + b2)

[0105] z ctxt =reshape(z2 W3+b3,[B×F,M,d) o ])

[0106] z final =reshape(LayerNorm(z ctxt ),[B,F,M,d o 2. Reference Image Coding Module

[0107] The reference frame encoding module used in this embodiment includes two parts: a reference frame spatial encoding compression module and a reference frame semantic encoding compression module. The reference frame spatial encoding compression module is an image encoder VAE pre-trained by Stability AI, which compresses the reference frame of size 512 into a latent variable of size 512 as the input of the reference network. The reference frame semantic compression module uses an ID Encoder pre-trained by Arcface[5] to generate an embedding that can represent the person ID information, and maps the representation of the reference frame in the latent space into a vector embedding of length 512 with high-level person ID information.

[0108] 3. Reference network and denoising network

[0109] Both the reference network and the denoising network employ the same Unet architecture to ensure that the denoising network can selectively introduce reference frame features from the reference network that are within the same level and dimension of the feature space in different layers. The overall architecture of Unet is symmetrical, consisting of a downsampling path (encoder), a bottleneck layer, and an upsampling path (decoder). Each layer of the decoder uses the same design in terms of spatial dimension and number of channels as its corresponding encoder module. Through a skip connection mechanism, each layer receives the output from the forward pass of the encoder module, serving as part of the input features to the decoder module. This architecture facilitates efficient extraction of high-level features while effectively recovering low-level features. The bottleneck layer further processes the encoded features, enabling the network to understand global semantic information on the smallest resolution feature map, providing overall semantic support for decoding and restoration to the initial resolution in the upsampling path.

[0110] Meanwhile, in different layers of Unet modules, the reference network utilizes a cross-attention mechanism to ensure the consistency of the obtained features at the high-level semantic level during the feature extraction process. The denoising network first performs spatial cross-attention between the reference frame feature map extracted by the reference network and the fusion latent variable output by the audio module to obtain the spatial information of the reference frame; then, it fuses semantic information through cross-attention with the reference network; finally, it employs a self-attention mechanism in the temporal dimension to ensure the consistency and continuity of content between frames in the generated sequence.

[0111] 4. Video frame decoding module

[0112] Corresponding to the reference frame space compression module, this module uses the VAE decoding module to reconstruct the latent variables into the pixel space, generating video frames, and finally obtaining a video frame sequence corresponding to the audio length, synthesizing a digital human motion video.

[0113] The specific implementation and technical effects of each part in this embodiment are similar to those in the above embodiments, and will not be repeated here.

[0114] To implement the above embodiments, this application also proposes a digital population type driving device. Figure 5 This is a schematic diagram of a digital mouth-type driving device provided in an embodiment of this application. Figure 5 As shown, the digital phreography actuator 500 includes:

[0115] The feature extraction module 510 is used to extract audio features from audio signals and facial coefficient features from video frames in videos;

[0116] The feature fusion module 520 is used to probabilistically introduce and fuse the audio features and the facial coefficient features through a contrastive learning feature fusion network to obtain feature embedding sequence features;

[0117] Module 530 is used to acquire the digital human reference image and noise latent variables;

[0118] The sequence output module 540 is used to process the feature embedding sequence features, the digital human reference image and noise latent variables through a denoising network, and output a digital human mouth shape video frame sequence.

[0119] The video generation module 550 is used to generate a digital human video based on the digital lip-print video frame sequence and the audio signal.

[0120] In some possible implementations, the feature extraction module 510 is used for:

[0121] Audio features are extracted from the audio signal using an audio extractor, and facial coefficient features are extracted from the video frames using a facial coefficient extractor; wherein, the facial coefficient features include posture features, expression features, and lip movement features.

[0122] In some possible implementations, the contrastive learning feature fusion network includes an attention module, a feature fusion layer, a cross-modal multi-head parameter attention layer, and a feature projection activation layer;

[0123] The feature fusion module 520 is used for:

[0124] The audio features and facial coefficient features are preprocessed and aligned to a preset fusion feature space.

[0125] The aligned audio features and facial coefficient features are input into the attention module for feature enhancement processing;

[0126] The enhanced audio features and facial coefficient features are input into the feature fusion layer for flattening and activation processing, as well as layer normalization and linear projection processing.

[0127] The audio features and facial coefficient features processed by the feature fusion layer are input into the cross-modal multi-head parameter attention layer to calculate the similarity score and multi-head attention weight.

[0128] The audio features and facial coefficient features processed by the cross-modal multi-head parameter attention layer are input into the feature projection activation layer for feature recombination, activation and reshaping to obtain feature embedding sequence features.

[0129] In some possible implementations, the sequence output module 540 is configured to:

[0130] The embedded sequence features, the digital human reference image, and the noise latent variables are input into the denoising network.

[0131] The denoising network performs multimodal feature fusion processing on the feature embedding sequence features, the digital human reference image, and the noise latent variables to generate enhanced features; wherein, the multimodal feature fusion processing includes spatial cross-attention processing, semantic cross-attention processing, and self-attention processing in the temporal dimension;

[0132] The enhanced features are decoded using the denoising network to obtain a sequence of digital phreatic video frames; wherein the decoding process includes recovering low-level features using a skip connection mechanism and upsampling using global semantic information.

[0133] In some possible implementations, the sequence output module 540 is specifically used for:

[0134] Through the denoising network, a cross-attention mechanism is used to perform spatial cross-attention processing on the features of the digital human reference image and the feature embedding sequence features to generate spatially enhanced features;

[0135] The spatial information of the reference frame and the noise latent variables are semantically cross-attention processed by the denoising network to generate semantically enhanced features.

[0136] The semantic enhancement features are processed by the denoising network using a self-attention mechanism in the temporal dimension to generate temporal enhancement features.

[0137] In some possible implementations, the sequence output module 540 is specifically used for:

[0138] The intermediate feature map output by the encoder is input into the decoder through the decoder of the denoising network, and the intermediate feature map is combined with the temporal enhancement feature using a skip connection mechanism to obtain the recovered feature; wherein, the intermediate feature map is obtained based on the feature embedding sequence feature, the spatial enhancement feature and the semantic enhancement feature;

[0139] The recovered features are processed through the bottleneck layer of the denoising network to extract global semantic information. The global semantic information is then combined with the semantic enhancement features to obtain enhanced global semantic features.

[0140] Based on the restored features and the enhanced global semantic features, a digital phreatic video frame sequence is obtained.

[0141] In some possible implementations, the video generation module 550 is used for:

[0142] The digital lip-sync video frame sequence is time-aligned with the audio signal to obtain a time-aligned video frame sequence.

[0143] Background synthesis is performed on the time-aligned video frame sequence to synthesize the digital human image into a preset background video frame, thereby generating a synthesized video frame sequence.

[0144] The synthesized video frame sequence is combined with the audio signal to generate a digital human video.

[0145] The specific implementation and technical effects of each module in this embodiment are similar to those in the above method embodiments, and will not be repeated here.

[0146] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0147] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0148] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0149] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0150] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0151] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this application is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.

[0152] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0153] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0154] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0155] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0156] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0157] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0158] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0159] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A digital population profile driving method, characterized by, The method comprises the following steps: extracting audio features in an audio signal and face coefficient features of video frames in a video; performing probabilistic introduction fusion processing on the audio features and the face coefficient features by a contrastive learning feature fusion network to obtain feature embedding sequence features; obtaining a digital human reference figure image and a noise hidden variable; processing the feature embedding sequence features, the digital human reference figure image and the noise hidden variable by a denoising network to output a digital human figure video frame sequence; generating a digital human video based on the digital human figure video frame sequence and the audio signal.

2. The digital demographic-driven method of claim 1, wherein, The method comprises the following steps: extracting audio features in an audio signal and face coefficient features of video frames in a video by an audio extractor and a face coefficient extractor; wherein the face coefficient features comprise posture features, expression features and lip movement features.

3. The digital population-driven method of claim 2, wherein, The contrastive learning feature fusion network comprises an attention module, a feature fusion layer, a cross-modal multi-head parameter attention layer and a feature projection activation layer. The method comprises the following steps: preprocessing the audio features and the face coefficient features, and aligning the audio features and the face coefficient features to a preset fusion feature space; inputting the aligned audio features and face coefficient features into the attention module for feature enhancement processing; inputting the feature-enhanced audio features and face coefficient features into the feature fusion layer for flattening and activation processing, layer normalization and linear projection processing; inputting the audio features and face coefficient features processed by the feature fusion layer into the cross-modal multi-head parameter attention layer to calculate similarity scores and multi-head attention weights; inputting the audio features and face coefficient features processed by the cross-modal multi-head parameter attention layer into the feature projection activation layer for feature reorganization, activation and reshaping processing to obtain feature embedding sequence features.

4. The digital demographic drive method of claim 1, wherein, The method comprises the following steps: inputting the feature embedding sequence features, the digital human reference figure image and the noise hidden variable into the denoising network; performing multi-modal feature fusion processing on the feature embedding sequence features, the digital human reference figure image and the noise hidden variable by the denoising network to generate enhanced features; wherein the multi-modal feature fusion processing comprises spatial cross-attention processing, semantic cross-attention processing and self-attention processing in the time dimension; performing decoding processing on the enhanced features by the denoising network to obtain a digital human figure video frame sequence; wherein the decoding processing comprises restoring low-level features by using a skip connection mechanism and performing up-sampling processing by using global semantic information.

5. The digital population profile driving method of claim 4, wherein, The multi-modal feature fusion processing of the feature embedding sequence feature, the digital human reference character image and the noise hidden variable is performed through the denoising network to generate an enhanced feature, including: The feature of the digital human reference character image and the feature embedding sequence feature are subjected to spatial cross-attention processing through the denoising network to generate spatial enhanced features; The spatial information of the reference frame and the noise hidden variable are subjected to semantic cross-attention processing through the denoising network to generate semantic enhanced features; The semantic enhanced features are processed through the self-attention mechanism in the time dimension of the denoising network to generate time enhanced features.

6. The digital population profile driving method of claim 5, wherein, The decoding processing of the enhanced features is performed through the denoising network to obtain a digital human portrait video frame sequence, including: The intermediate feature map output by the encoder is input into the decoder of the denoising network, and the intermediate feature map and the time enhanced features are combined and processed through the skip connection mechanism to obtain restored features; wherein the intermediate feature map is obtained based on the feature embedding sequence feature, the spatial enhanced feature and the semantic enhanced feature; The restored features are processed through the bottleneck layer of the denoising network to extract global semantic information, and the global semantic information and the semantic enhanced features are combined to obtain enhanced global semantic features; The digital human portrait video frame sequence is obtained based on the restored features and the enhanced global semantic features.

7. The digital demographic-driven method of claim 1, wherein, The digital human video is generated based on the digital human portrait video frame sequence and the audio signal, including: The digital human portrait video frame sequence and the audio signal are time-aligned to obtain a time-aligned video frame sequence; The time-aligned video frame sequence is subjected to background synthesis to synthesize the digital human image into a preset background video frame to generate a synthesized video frame sequence; The synthesized video frame sequence and the audio signal are synthesized to generate a digital human video.

8. A digital population profile driving device, characterized by, It includes: A feature extraction module is configured to extract audio features in an audio signal and facial coefficient features of video frames in a video; A feature fusion module is configured to perform probabilistic introduction fusion processing on the audio features and the facial coefficient features through a contrast learning feature fusion network to obtain feature embedding sequence features; An acquisition module is configured to acquire a digital human reference character image and a noise hidden variable; A sequence output module is configured to process the feature embedding sequence features, the digital human reference character image and the noise hidden variable through a denoising network to output a digital human portrait video frame sequence; A video generation module is configured to generate a digital human video based on the digital human portrait video frame sequence and the audio signal.

9. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

10. A computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1-7. The computer instructions are for causing the computer to perform the method according to any one of claims 1-7.