Digital human mouth shape synchronization method and device, and electronic device

By using audio feature extraction and multi-scale feature fusion rendering technology, the problems of unnatural lip shape and identity drift in traditional digital human lip-syncing methods have been solved, achieving high-resolution, high-fidelity digital human lip-syncing and improving the realism and visual coherence of digital human interaction.

CN122415803APending Publication Date: 2026-07-17AIJUN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AIJUN TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional 2D digital human mouth-syncing methods produce stiff lip changes, have a low degree of matching with real pronunciation mouth shapes, lack natural fluency, and the digital human identity features are prone to drift, affecting visual coherence. It is difficult to generate high-resolution, high-fidelity digital human images, especially in terms of detail representation and anti-artifact performance in the lip area.

Method used

Audio signals are processed using Fourier transform and Mel filter banks to extract Mel spectrum and Mel cepstral coefficients. Pre-trained speech models and recurrent neural networks are used to capture audio pronunciation details and generate lip shape parameters. Feature fusion is performed by combining encoder and decoder, and multi-scale feature fusion is performed using cross-modal attention mechanism and adaptive instance normalization. Fully convolutional neural networks and adaptive affine transformations are applied to render digital human images.

Benefits of technology

It significantly improves the synchronization accuracy of digital human mouth shape and voice, enhances the realism of digital human images and user experience, and ensures the consistency of digital human identity and high-quality lip image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415803A_ABST
    Figure CN122415803A_ABST
Patent Text Reader

Abstract

This invention provides a digital human lip-syncing method, apparatus, and electronic device. The method includes: extracting audio features from the original audio signal and mapping the audio features to lip-sync parameters; fusing the lip-sync parameters with at least one preset reference image frame to obtain fused features; and rendering the fused features into a digital human image. This significantly improves the realism of digital human interaction and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital media technology, specifically to a digital lip-syncing method, apparatus, and electronic device. Background Technology

[0002] With the rapid development of artificial intelligence and computer graphics technology, digital human technology has shown enormous application potential in fields such as virtual reality, augmented reality, film and television production, and online education. Among these, achieving precise synchronization between lip movements and speech in digital humans is a crucial step in enhancing the realism and interactive experience of digital humans. Traditional 2D digital human lip-syncing methods often produce stiff lip movements, with low matching accuracy to real-life pronunciation, lacking natural fluency. During animation generation, the digital human's identity features (such as facial structure and texture details) are prone to drift, leading to inconsistencies in the digital human image between different frames, affecting visual continuity, especially in the lip area. Furthermore, it is difficult to generate high-resolution, high-fidelity 2D digital human images, particularly in terms of detail representation and artifact resistance in the lip area. Summary of the Invention

[0003] In view of this, embodiments of the present invention aim to provide a digital human mouth synchronization method, apparatus and electronic device, which solves the above-mentioned technical problems.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a digital human lip-syncing method, comprising: extracting audio features based on an original audio signal and mapping the audio features to lip-sync parameters; performing feature fusion on the lip-sync parameters and at least one preset reference image frame to obtain fused features; and rendering the fused features into a digital human image.

[0005] In one embodiment, the step of extracting audio features from the original audio signal includes: converting the original audio signal to the frequency domain through Fourier transform and processing it through a Mel filter bank to obtain the Mel spectrum; and / or, performing a discrete cosine transform on the Mel spectrum to extract Mel cepstral coefficients; and / or, using context embedding features extracted by a pre-trained speech model.

[0006] In one embodiment, mapping the audio features to lip shape parameters includes: applying a preset recurrent neural network to process the audio features, capturing pronunciation details in the audio, and outputting lip shape parameters.

[0007] In one embodiment, the mouth shape parameters include 2D or 3D lip key points and / or mouth shape parameters.

[0008] In one embodiment, the step of fusing the lip shape parameters with at least one preset reference image frame to obtain fused features includes: extracting identity features from the reference image frame using an encoder; and performing deep fusion of the identity features and the lip shape parameters in the feature space to obtain fused features.

[0009] In one embodiment, the step of deeply fusing the identity features and the lip shape parameters in the feature space to obtain fused features includes: fusing the identity features and the lip shape parameters at multiple scales through splicing, cross-modal attention mechanisms, or adaptive instance normalization to obtain fused features.

[0010] In one embodiment, rendering the fused features into a digital human image includes: applying a preset renderer to generate a digital human image frame by frame based on the fused features to form a mouth-sync video, wherein a decoder is applied to capture subtle texture and shape changes in the mouth area based on the fused features to enhance local details.

[0011] According to another aspect of the present invention, an embodiment of the present invention provides a digital human lip-syncing device, the digital human lip-syncing device comprising: a feature extraction module, configured to extract audio features based on an original audio signal and map the audio features to lip-syncing parameters; a feature fusion module, configured to fuse the lip-syncing parameters with at least one preset reference image frame to obtain fused features; and an image rendering module, configured to render the fused features into a digital human image.

[0012] According to another aspect of the present invention, an embodiment of the present invention provides a digital mouth-tracking synchronization device, comprising: a memory for storing executable program code; and a processor for calling and running the executable program code from the memory, such that the processor executes the aforementioned digital mouth-tracking synchronization device.

[0013] The present invention provides a digital human lip-syncing method, apparatus, and electronic device, which extracts audio features from the original audio signal and maps the audio features to lip-sync parameters; fuses the lip-sync parameters with at least one preset reference image frame to obtain fused features; and renders the fused features into a digital human image, significantly improving the realism of digital human interaction and user experience. Attached Figure Description

[0014] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0015] Figure 1 The diagram shown is a flowchart of a digital human mouth synchronization method provided in an embodiment of this application.

[0016] Figure 2 The diagram shown is a structural schematic of a digital mouth-shaped synchronization device provided in an embodiment of this application.

[0017] Figure 3 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Furthermore, in exemplary embodiments, since the same reference numerals denote the same components having the same structure or the same steps of the same method, if one embodiment has been described by way of example, then in other exemplary embodiments only structures or methods different from those described in the embodiment will be described.

[0020] Throughout the specification and claims, when a component is described as being “connected” to another component, that component may be “directly connected” to the other component or “electrically connected” to the other component via a third component. Furthermore, unless explicitly stated otherwise, the term “comprising” and its corresponding terms should be understood only to include the stated component and not to exclude any other component.

[0021] With the rapid development of artificial intelligence and computer graphics technology, digital human technology has shown enormous application potential in fields such as virtual reality, augmented reality, film and television production, and online education. Among these, achieving precise synchronization between lip movements and speech in digital humans is a crucial step in enhancing the realism and interactive experience of digital humans. Traditional two-dimensional (2D) digital human lip-syncing methods often produce stiff lip movements, poor matching with real-life pronunciation, and a lack of natural fluency. During animation generation, the digital human's identity features (such as facial structure and texture details) are prone to drift, leading to inconsistencies in the digital human image between different frames, affecting visual coherence, especially in the lip area. Furthermore, it is difficult to generate high-resolution, high-fidelity 2D digital human images, particularly in terms of detail representation and artifact resistance in the lip area.

[0022] To address these issues, researchers have proposed various technical solutions. For example, some methods utilize deep learning models to directly generate image sequences from audio, but often struggle to balance the naturalness of lip shapes with identity consistency. Other methods drive 2D image generation through 3D intermediate representations, which improves lip shape control but still leaves room for improvement in rendering details and real-time performance, such as applying a 3D Morphable Face Model (3DMM) to drive 2D image generation. Therefore, there is an urgent need for a 2D digital human lip-syncing method that can effectively combine audio information, maintain digital human identity consistency, and generate high-quality lip images.

[0023] To address the technical problems of stiff lip movements in digital human lip shapes, poor matching with real pronunciation, and lack of natural fluency in related technologies, this application provides a digital human lip shape synchronization method. Figure 1 The diagram shown is a flowchart illustrating a digital mouth-tracking synchronization method provided in an embodiment of this application. Figure 1 As shown, the digital mouth synchronization method includes: Step S11: Extract audio features from the original audio signal and map the audio features to lip shape parameters.

[0024] The original audio signal can be data input by the user when the digital human is triggered. A digital human refers to a virtual character created using computer technologies such as computer graphics, artificial intelligence, machine learning, and natural language processing, typically possessing human-like appearance, communication abilities, or motor skills. In this embodiment, the digital human, by simulating real human facial expressions or body movements, can provide users with a more realistic and immersive interactive experience.

[0025] After receiving the raw audio signal input by the user, audio features can be extracted from the raw audio signal. In this embodiment, the raw audio signal can be converted to the frequency domain using Fourier transform and processed by a Mel-filter bank to obtain the Mel-spectrogram. And / or, a discrete cosine transform can be performed on the Mel-spectrogram to extract Mel-frequency cepstral coefficients (MFCCs) to obtain a more compact feature representation. And / or, contextual embedding features extracted by a pre-trained speech model can also be utilized. The speech model can be any existing model capable of extracting features from audio signals, such as Wav2Vec 2.0, HuBERT (Hidden-unit BERT), etc. Contextual embedding features contain richer semantic and emotional information. In this embodiment, one or more features can be selected for extraction according to actual needs and normalized.

[0026] The extracted audio features are used as input, and a pre-defined recurrent neural network is applied to process these features, capturing pronunciation details and outputting lip shape parameters. Specifically, a pre-trained neural network model predicts lip shape parameters corresponding to the audio content. An architecture based on a temporal correlation extraction network can be adopted. For example, a sequence-to-sequence (Seq2Seq) model can be used, such as a Transformer-based encoder-decoder structure or a recurrent neural network containing multiple gated recurrent units (GRUs). The input is a sequence of audio features, and the output is a sequence of lip shape parameters corresponding to the time step. To capture pronunciation details in the audio, an attention mechanism or a multi-head self-attention mechanism can be introduced into the network. Lip shape parameters are feature representations related to facial movement, containing information about the movement of muscles such as the lips and cheeks during pronunciation, and are the driving force behind the "speaking" of a static image. Lip shape parameters include 2D or 3D lip keypoints and / or mouth shape parameters, etc. For example, 2D lip keypoints can directly represent the lip contour, and 3D mouth shape parameters can be parameterized based on the mouth region of a 3DMM. The embodiments of this application primarily predict mouth shape changes related to pronunciation. Mouth shape parameters are typically low-dimensional vectors, such as 20-dimensional or 30-dimensional.

[0027] Before applying recurrent neural networks, a large amount of audio-video data is required for training. The video data is fitted with lip keypoints or 3DMM mouth regions to extract accurate mouth shape parameters. The training objective is to minimize the absolute difference loss (L1) or squared error loss (L2) between the predicted mouth shape parameters and the actual mouth shape parameters. Adversarial loss can be introduced to improve the realism of the generated mouth shape.

[0028] Step S12: Perform feature fusion on the mouth shape parameters and at least one preset reference image frame to obtain fused features.

[0029] Reference image frames are typically high-quality still images of the digital human or keyframes from video sequences. They provide the digital human's identity information and texture details of other facial areas. These frames can be pre-selected and stored, such as by choosing frames based on different facial poses and mouth opening degrees. The reference image frames contain the digital human's identity information (such as facial texture, hairstyle, skin tone, etc.), forming the basic "persona" for the generated video. To enhance robustness, at least one reference image frame should be selected.

[0030] After receiving the lip shape parameters and one or more reference image frames, an encoder extracts identity features from the reference image frames. The identity features and the lip shape parameters are then deeply fused in the feature space to obtain fused features, which contain intermediate feature representations of lip shape and identity information. The encoder can use a ResNet-based or Visual Geometry Group (VGG) feature extraction network. High-level identity features are extracted from the reference image frames by the encoder. Simultaneously, a small multilayer perceptron (MLP) converts the previously obtained lip shape parameters into a vector representation that can be fused with image features.

[0031] The identity features and lip-shape parameters are fused at multiple scales using concatenation, cross-modal attention mechanisms, or adaptive instance normalization (AdaIN) to obtain fused features. Concatenation directly joins the identity feature vector and lip-shape parameter vector along the channel dimension, forming a longer feature vector containing both types of information. Cross-modal attention mechanisms offer a more intelligent feature fusion approach, allowing the model to dynamically learn the correlation between identity features and lip-shape parameters. For example, the model can learn which parts of the lips need to move more when pronouncing a certain sound, and how these movements combine with the lip shape of a specific person. Adaptive instance normalization focuses more on style transfer and fusion, treating identity features as "content" and lip-shape parameters (or their derived style features) as "style." By adjusting the mean and variance of the content features to match the statistical properties of the style features, the dynamic style of speech is "rendered" onto the static identity content.

[0032] The fused features contain the lip shape information and digital human identity information required for the current frame, laying a solid foundation for subsequent video frame generation. To further enhance the fusion effect, a multi-scale feature fusion strategy can be introduced to fuse features at different levels. By injecting reference image frames, the renderer can utilize the rich texture and detail information of the reference image frames, avoiding image generation from scratch, thereby significantly improving the quality and identity consistency of the generated images, especially in the lip region. The digital human image generated based on the fused features can also serve as a reference image frame for subsequent feature fusion. The use of so many reference image frames also enhances robustness to fluctuations in the quality of reference image frames.

[0033] Step S13: Render the fused features into a digital human image.

[0034] In this embodiment of the invention, a preset renderer can be applied to convert the fused feature representation into a final 2D digital human image. Optionally, the preset renderer is applied to generate a digital human image frame by frame based on the fused features, forming a lip-sync video. A decoder is applied to capture subtle textures and shape changes in the mouth region based on the fused features, performing local detail enhancement. A fully convolutional neural network (Unet) can be used, and an adaptive affine transformation (AdaAT) can be used to align the reference image with the original image. The input is the fused features, and the output is a 2D digital human image. Adaptive affine transformation is a neural network module for image generation designed to solve the spatial misalignment problem between feature maps. It not only integrates information from different modalities but also spatially deforms and aligns features, thereby generating a more accurate and natural output. By explicitly handling spatial misalignment through adaptive affine transformation (AdaAT), it ensures that the movement trajectory and shape of the lips are highly synchronized with the audio signal. Adaptive Affine Transformation (AdaAT) predicts a set of affine transformation parameters (θ) based on fused features, including rotation, scaling, translation, etc. Using these affine transformation parameters, spatial warping is applied to the feature maps of the source image. This adaptive affine transformation (AdaAT) forcibly transforms the facial features of the source person to the position of the target pose or mouth shape. For example, if the audio requires the mouth to be open wide, AdaAT stretches the feature regions of the closed mouth to roughly align them spatially with the opening mouth action. Because AdaAT only performs simple geometric stretching, the stretched image will have artifacts or blurring in edges or occluded areas. UNet's task is to repair these details and generate clear, coherent 2D digital human images. Specifically, the feature maps aligned by AdaAT are fed into UNet for final generation. The encoder in UNet further extracts geometrically corrected features to capture the global structure, and the decoder in UNet progressively upsamples to restore the image resolution, generating the 2D digital human image.

[0035] To generate high-resolution images, a decoder (SPADEDecoder) based on Spatially-Adaptive Normalization (SPADE) is applied for decoding, and a multi-scale discriminator is used for correction. The main function of SPADEDecoder is to transform coarse semantic layouts (such as semantic segmentation maps) into high-resolution, realistic, and detailed RGB images, i.e., digital human images. SPADEDecoder receives the fused features and uses the spatially adaptive normalization method to recover high-frequency details and textures, generating the final high-definition video frames.

[0036] In this embodiment, to generate high-fidelity images, the decoder employs a residual design with more parameters to capture subtle textures and shape variations in the lip region. Additionally, a local detail enhancement method can be introduced, where a local discriminator separately identifies the mouth region and enhances loss calculations for this region, enabling refined rendering of the mouth area while preserving the original features of other facial areas.

[0037] In addition to traditional pixel-level losses (such as absolute difference loss / squared error loss), perceptual loss can be introduced during training to improve the sharpness of the generated image, and adversarial loss can be used to eliminate artifacts, so as to ensure the visual quality and realism of the generated image.

[0038] The digital human lip-syncing method of this application can be applied to a virtual anchor system. For example, a user inputs a voice clip of an anchor, and the system generates a 2D digital human anchor video synchronized with the voice content in real time. Specifically, audio feature extraction uses Wav2Vec 2.0 to extract audio embedding features; audio-to-lip-sync parameter generation uses a Transformer-based encoder-decoder network to predict 20-dimensional lip-sync parameters; multi-reference frame injection and feature fusion use a ResNet-based encoder to extract reference frame features and perform feature fusion using AdaIN; high-fidelity adaptive rendering uses a StyleGAN2-based generator to generate a 1024x1024 resolution digital human image, where only the lip region is rendered according to the lip-sync parameters, while other facial regions retain the original features of the reference frames.

[0039] This invention provides a digital human lip-syncing method that extracts audio features from the original audio signal and maps the audio features to lip-sync parameters; fuses the lip-sync parameters with at least one preset reference image frame to obtain fused features; and renders the fused features into a digital human image, which can significantly improve the realism of digital human interaction and user experience.

[0040] Figure 2 The diagram shown is a structural schematic of a digital mouth-tracking synchronization device according to an embodiment of this application. Figure 2 As shown, the digital mouth-shaped synchronization device 200 includes: Feature extraction module 201 is used to extract audio features based on the original audio signal and map the audio features to lip shape parameters; The feature fusion module 202 is used to perform feature fusion on the mouth shape parameters and at least one preset reference image frame to obtain fused features; Image rendering module 203 is used to render the fused features into a digital human image.

[0041] In one embodiment, the feature extraction module 201 is used to: convert the original audio signal to the frequency domain by Fourier transform and process it by Mel filter bank to obtain Mel spectrum; and / or, perform discrete cosine transform on the Mel spectrum to extract Mel cepstral coefficients; and / or, use context embedding features extracted by a pre-trained speech model.

[0042] In one embodiment, the feature extraction module 201 is further configured to: apply a preset recurrent neural network to process the audio features, capture pronunciation details in the audio, and output mouth shape parameters.

[0043] In one embodiment, the mouth shape parameters include 2D or 3D lip key points and / or mouth shape parameters.

[0044] In one embodiment, the feature fusion module 202 is used to: extract identity features from the reference image frame using an encoder; and perform deep fusion of the identity features and the lip shape parameters in the feature space to obtain fused features.

[0045] In one embodiment, the feature fusion module 202 is further configured to: perform multi-scale feature fusion of the identity features and the lip shape parameters by splicing, cross-modal attention mechanism or adaptive instance normalization to obtain fused features.

[0046] In one embodiment, the image rendering module 203 is used to: apply a preset renderer to generate a digital human image frame by frame based on the fusion features to form a mouth-sync video, wherein a decoder is applied to capture subtle texture and shape changes in the mouth area based on the fusion features to perform local detail enhancement.

[0047] According to another aspect of the present invention, one embodiment of the present invention provides an electronic device, Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0048] For example, such as Figure 3As shown, the electronic device includes a memory 301 and a processor 302. The memory 301 stores executable program code 3011, and the processor 302 is used to call and execute the executable program code 3011 from the memory, so that the processor 302 executes the digital mouth-sync method.

[0049] This embodiment can divide the electronic device into functional modules according to the above method embodiment. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and is only a logical functional division. In actual implementation, there may be other division methods.

[0050] When each functional module is divided according to its corresponding function, the electronic device may include: a feature extraction module, a feature fusion module, and an image rendering module, etc.

[0051] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0052] The electronic device provided in this embodiment is used to execute the above-described method for co-synchronizing digital human mouth shapes, and thus can achieve the same effect as the above-described implementation method.

[0053] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of program code and data by the electronic device.

[0054] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits as disclosed in this application. The processor may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.

[0055] This embodiment also provides a computer-readable storage medium (including but not limited to disk storage, CD-ROM, optical storage, etc.) storing computer program code. When the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement the digital human mouth synchronization method provided in the above embodiment. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, Digital Video Discs (DVDs), Compact Disc Read-Only Memory (CD-ROM), microdrives, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), dynamic random access memory (DRAM), video random access memory (VRAM), flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0056] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to realize the digital human mouth synchronization method provided in the above embodiment.

[0057] The beneficial effects of the above embodiments can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0058] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0059] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0060] It should also be noted that in the apparatus or equipment of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0061] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0062] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A digital lip-syncing method, characterized in that, The method includes: Audio features are extracted from the original audio signal and then mapped to lip shape parameters. The mouth shape parameters are fused with at least one preset reference image frame to obtain fused features; The fused features are rendered into a digital human image.

2. The method according to claim 1, characterized in that, The extraction of audio features based on the original audio signal includes: The original audio signal is converted to the frequency domain using Fourier transform, and then processed using a Mel filter bank to obtain the Mel spectrum; and / or, Based on the Mel spectrum, perform a discrete cosine transform to extract the Mel cepstral coefficients; and / or, Contextual embedding features extracted using a pre-trained speech model.

3. The method according to claim 1, characterized in that, The step of mapping the audio features to lip shape parameters includes: The audio features are processed using a pre-defined recurrent neural network to capture pronunciation details in the audio and output mouth shape parameters.

4. The method according to claim 1, characterized in that, The mouth shape parameters include 2D or 3D lip key points and / or mouth shape parameters.

5. The method according to claim 1, characterized in that, The step of fusing the mouth shape parameters with at least one preset reference image frame to obtain fused features includes: The encoder is used to extract identity features from the reference image frame; The identity features and the lip shape parameters are deeply fused in the feature space to obtain fused features.

6. The method according to claim 5, characterized in that, The step of deeply fusing the identity features and the lip shape parameters in the feature space to obtain fused features includes: The identity features and the lip shape parameters are fused using multi-scale feature fusion through splicing, cross-modal attention mechanisms, or adaptive instance normalization to obtain fused features.

7. The method according to claim 1, characterized in that, The step of rendering the fused features into a digital human image includes: The application uses a preset renderer to generate digital human images frame by frame based on the fusion features, forming a mouth-synced video. The application uses a decoder to capture subtle textures and shape changes in the mouth area based on the fusion features, and performs local detail enhancement.

8. A digital human mouth-shaped synchronization device, characterized in that, The device includes: The feature extraction module is used to extract audio features from the original audio signal and map the audio features to lip shape parameters; The feature fusion module is used to fuse the mouth shape parameters with at least one preset reference image frame to obtain fused features; An image rendering module is used to render the fused features into a digital human image.

9. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor for calling and running the executable program code from the memory, such that the processor performs the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 7.