Digital human head generation method based on real-time audio drive

By introducing trainable embedded tags and dynamic Gaussian functions into the digital avatar generation method, combined with 3DGS technology, the problems of real-time rendering and dynamic scene rendering in the existing technology are solved, and efficient real-time digital avatar generation is achieved, which significantly improves the rendering efficiency.

CN119006663BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411061800.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2025-05-27
Estimated Expiration
2044-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively use audio-driven digital avatar generation method in real-time rendering and dynamic scenes, especially in scenarios such as live streaming, which cannot meet the needs of real-time rendering.

Method used

The digital avatar generation method based on real-time audio driver is adopted. By introducing trainable embedded labels and dynamic Gaussian functions, combined with 3D GuassionSplatting (3DGS) technology, dynamic scene rendering of the talk head is realized, and the audio feature extraction model is aligned with the driver signal to realize real-time rendering.

Benefits of technology

It significantly improves the efficiency of vocal head generation, achieves efficient real-time rendering, achieves the best inference efficiency at present, and achieves nearly 50 times the rendering efficiency improvement under the conditions of ensuring the same rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006663B_ABST
    Figure CN119006663B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a digital human head based on real-time audio driving, including: introducing a learnable embedding code to indirectly represent 3DGS to complete the training of the head rendering model, calculating the loss function according to the key facial features and audio coding features to complete the training of the audio feature extraction model, achieving the control of the audio on the modeling dynamic scene through the alignment of the real-time audio coding features and the key facial features, and finally completing the rendering of the talking head through Splatting, thereby realizing the generation of the speech-driven talking head. The present invention introduces a trainable embedding label as a position condition, uses a dynamic Gaussian function and audio input to drive the talking head for modeling, realizes the dynamic scene rendering of the digital human head, and has high rendering efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual digital technology, and particularly relates to a method for generating a digital human head based on real-time audio driving. Background Art

[0002] Speaker head generation is a key technology in the process of virtual digital human synthesis. Generally, audio, video, or text is used to drive the movement of parts such as the mouth, head, and body of a person. With the proposal of concepts such as digital twins and the metaverse, a large number of research teams have devoted their efforts to the research of talking heads, and many high-quality and high-efficiency research results on talking head generation have been produced.

[0003] In recent years, the technology for generating high-naturalness talking head videos driven by audio has received increasing attention and has broad application prospects in virtual simulation images, film dubbing, virtual reality, etc. The task of audio-driven talking head synthesis aims to synthesize an audio-related portrait action video based on an input reference image and an arbitrary piece of audio. Currently, mainstream research can be divided into methods based on technologies such as GAN and Pix2Pix, and methods based on technologies such as neural rendering represented by Neural Radiance Field (NeRF).

[0004] StyleGAN is one of the methods with the best effects among the current GAN-based methods. It is trained based on a large-scale audio-visual dataset and uses a pre-trained lip discriminator to generate individual-independent lip movements, thus obtaining a good lip-sync effect. However, such methods not only perform poorly when rendering large-scale movements but also incur high training costs because generating images of different resolutions requires repeated training.

[0005] The NeRF-based methods make full use of the capabilities of NeRF such as generating real new perspective images to greatly improve the fidelity of the generated talking heads. AD-NeRF is one of the representative works. It first introduces the neural radiance field (NeRF) as an intermediate representation to implicitly learn human features in this task, and uses the features extracted by DeepSpeech to regress human-related actions, achieving impressive results in terms of human restoration. ER-NeRF expands such methods and greatly improves the computational efficiency. Although the NeRF-based methods achieve high-fidelity video rendering, the training cost for a specific individual is also extremely high, which limits its application in scenarios such as real-time rendering. At the same time, the disadvantage of NeRF in rendering efficiency leads to the above methods being unable to well meet the requirements of real-time talking head rendering in scenarios such as live streaming push.

[0006] A method for modeling a scene by using a set of closed functions as representation primitives. Among them, 3DGS (Three-Dimensional Gaussian Sputtering) is a representative method in this direction. It initializes Gaussian functions through a set of random or structured SFM point clouds, and adaptively learns image features through the density control of adaptively optimized Gaussian functions. Thanks to the fast rasterization mechanism based on sorting, this method can achieve much higher training and inference efficiency than NeRF. Although the training process of this method is adaptive, it is a static scene reconstruction scheme. If you need to model a dynamic scene, you need to model the static scene under a monocular camera frame by frame. The static Gaussian function cannot be driven by other data, which cannot be applied to the task of the present invention.

[0007] Therefore, how to provide a digital human head generation method that can use audio to drive Gaussian functions for dynamic scene rendering is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0008] In view of the above research status, the present invention provides a digital human head generation method based on real-time audio driving, introducing a trainable embedding label f s As a position condition, a dynamic Gaussian function and audio input are used to drive the talking head for modeling, realizing the dynamic scene rendering of the digital human head with high rendering efficiency.

[0009] A digital human head generation method based on real-time audio driving provided by the present invention includes the following steps:

[0010] Head rendering model training stage:

[0011] S100: Extract N head images of consecutive frames from the head dynamic action video frames, N>1;

[0012] S101: Input the N head images into the 3DDFA model to extract the 3DMM point cloud data of the head and N groups of face key features corresponding to the N head images as driving signals;

[0013] S102: The 3DMM point cloud data is initialized by 3DGS to obtain a static Gaussian distribution describing the 3DMM point cloud data, calculate the spatial position semantic information of the Gaussian distribution, and use it as the embedding label of the static Gaussian distribution to generate a head dynamic Gaussian distribution;

[0014] S103: Input the driving signal and the spatial position semantic information corresponding to the current head image into the motion controller to predict and output the head motion offset;

[0015] S104: After the head pose offset is superimposed and updated to the head dynamic Gaussian function, project the Gaussian distribution corresponding to the updated head dynamic Gaussian function in the 3D space onto the two-dimensional image plane to obtain the rendered head image;

[0016] S105: Calculate the loss function based on the rendered head image and the current head image. After reversely optimizing and updating the spatial position semantic information according to the loss function calculation result, repeat S103 - S105 until the training of the N head images is completed to obtain a trained head rendering model;

[0017] Audio feature extraction model training stage:

[0018] S200: Input the audio of a given duration into the audio feature extraction model to extract audio encoding features;

[0019] S201: Calculate the loss function based on the facial key features and the audio encoding features. After reversely optimizing and updating the audio feature extraction model according to the loss function calculation result, obtain a trained audio feature extraction model;

[0020] Real-time audio-driven digital human head generation stage:

[0021] Input the N head images into the trained head rendering model to obtain N sets of facial key features, that is, N sets of driving signals;

[0022] Input the real-time audio into the trained audio feature extraction model to extract real-time audio encoding features, align and replace the real-time audio encoding features with the driving signals to achieve the driving of the audio to the facial semantic prior; use the trained head rendering model to output the rendering of the real-time digital human head.

[0023] Preferably, in the head rendering model training stage, the steps of extracting the 3DMM point cloud data of the head include:

[0024] The 3DDFA model screens the N head images to obtain a standard head image, and extracts the 3DMM point cloud data from the standard head image.

[0025] Preferably, in the head rendering model training stage, the steps of extracting the facial key features as driving signals include:

[0026] The 3DDFA model extracts the facial expression feature f for each of the head images exp as the driving signal.

[0027] Preferably, in the training stage of the avatar rendering model, the 3DDFA model is further used to extract the head pose feature f for each of the avatar images pos .

[0028] Preferably, in the training stage of the avatar rendering model, the steps of calculating the spatial position semantic information of the Gaussian distribution include:

[0029] Calculating the point coordinates of the spatial center points of all Gaussian ellipsoidal domains in the static Gaussian distribution;

[0030] Performing Fourier encoding on the point coordinates to obtain the semantic information f s .

[0031] Preferably, in the training stage of the avatar rendering model, the steps of S103 include:

[0032] The 3DDFA model composites the facial expression feature f exp and the head pose feature f pos to generate f w :

[0033] ;

[0034] The motion controller adopts a two-layer MLP to form an attention mechanism, inputs the f w into an MLP to obtain an encoded vector f a , combines the semantic information f s , and uses an attention mechanism along the dimension to obtain a control vector f v :

[0035] ;

[0036] ;

[0037] Obtaining the attention weight f a through the motion controller, and obtaining the spatial position offset p, the rotation offset r, and the change in the size of the ellipsoidal domain of the Gaussian distribution s:

[0038] .

[0039] Preferably, in S105, the loss function is further used to reversely optimize and update the 3DGS parameters and the motion controller parameters.

[0040] Preferably, in S105, the steps of calculating the loss function based on the rendered avatar image and the current avatar image include:

[0041] The loss function is expressed as follows:

[0042] ;

[0043] where λ 1 and λ 2 respectively represent and weights;

[0044] ;

[0045] where is the rendered avatar image, is the corresponding avatar image;

[0046] ;

[0047] where the features of the rendered avatar image and the corresponding avatar image are extracted using the VGG19 network, represents the output of the i-th layer of the VGG19 network.

[0048] Preferably, in the S200, the steps of extracting the audio coding features include:

[0049] The audio feature extraction model segments the audio according to the frame rate of the head dynamic action video frame to obtain an audio sequence containing N audio segments;

[0050] DeepSpeech is used to extract the features of the audio sequence and combined with an end-to-end audio codec for feature encoding and decoding; the calculation result of the loss function is used to reversely optimize and update the audio codec.

[0051] Preferably, in the S201, the steps of calculating the loss function according to the face key features and the audio coding features include:

[0052] The loss function calculates the loss value between the audio coding features of each audio segment and the face key features of the avatar image of its corresponding frame.

[0053] The digital human avatar generation method based on real-time audio driving proposed by the present invention has the following beneficial effects compared with the prior art:

[0054] The present invention proposes a FastTalker method, which for the first time introduces the high-efficiency 3D scene rendering method 3D GuassionSplatting (3DGS) into the talking head generation task, greatly improving the efficiency of talking head generation.

[0055] The present invention proposes a phased processing strategy. In the first phase, a learnable spatial position semantic information is introduced as an embedding label to indirectly represent the 3D Guassion. In the second phase, audio and driving signals are aligned to achieve the control of audio over the modeled dynamic scene. Finally, the talking head is rendered through Splatting, thereby realizing the generation of a voice-driven talking head.

[0056] The present invention conducts a comparative experiment on an open-source talking head generation dataset. The results show that FastTalker achieves the current optimal effect in terms of inference efficiency. Compared with the state-of-the-art NeRF method, it realizes a rendering efficiency nearly 50 times higher under the condition of ensuring the same rendering quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the provided drawings without creative efforts.

[0058] Figure 1 It is a schematic diagram of the method for generating a digital human head driven by real-time audio provided by an embodiment of the present invention;

[0059] Figure 2 It is a flowchart of the method for generating a digital human head driven by real-time audio provided by an embodiment of the present invention;

[0060] Figure 3 It is a detailed block diagram of the method for generating a digital human head driven by real-time audio provided by an embodiment of the present invention;

[0061] Figure 4 is a schematic diagram of the trainable semantic label embedding control vector based on the semantic attention mechanism provided by an embodiment of the present invention;

[0062] Figure 5 is a schematic diagram of the alignment of audio and driving signals provided by an embodiment of the present invention;

[0063] Figure 6 is a comparison diagram of the rendering effects of multiple methods provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0065] The present invention provides a method for generating a digital human head portrait driven by real-time audio, as Figures 1-3 shown, which includes the following steps:

[0066] Head portrait rendering model training stage:

[0067] S100: Extract N head portrait pictures of consecutive frames from the video frames of the dynamic actions of the human head, where N>1;

[0068] S101: Input the N head portrait pictures into the 3DDFA model to extract the 3DMM point cloud data of the head portrait and N groups of key facial features corresponding to the N head portrait pictures as driving signals;

[0069] S102: The 3DMM point cloud data is initialized by 3DGS to obtain a static Gaussian distribution describing the 3DMM point cloud data, calculate the spatial position semantic information of the Gaussian distribution, and use it as the embedding label of the static Gaussian distribution to generate a dynamic Gaussian distribution of the head portrait;

[0070] S103: Input the driving signal and the spatial position semantic information corresponding to the current head portrait picture into the motion controller to predict and output the head portrait motion offset;

[0071] S104: After the head portrait motion offset is superimposed and updated to the dynamic Gaussian function of the head portrait, project the Gaussian distribution in the 3D space corresponding to the updated dynamic Gaussian function of the head portrait onto the two-dimensional image plane to obtain the rendered head portrait picture;

[0072] S105: Calculate the loss function according to the rendered head portrait picture and the current head portrait picture, and reversely optimize and update the spatial position semantic information according to the calculation result of the loss function, then repeat S103-S105 until the training of the N head portrait pictures is completed to obtain a trained head portrait rendering model;

[0073] Audio feature extraction model training stage:

[0074] S200: Input the audio of a given duration into the audio feature extraction model to extract the audio coding features;

[0075] S201: Calculate the loss function according to the key facial features and the audio coding features, and reversely optimize and update the audio feature extraction model according to the calculation result of the loss function to obtain a trained audio feature extraction model;

[0076] Real-time audio-driven digital human head portrait generation stage:

[0077] Input the N head portrait pictures into the trained head portrait rendering model to obtain N groups of key facial features, that is, N groups of driving signals;

[0078] Input the real-time audio into the trained audio feature extraction model to extract the real-time audio coding features, align and replace the real-time audio coding features with the driving signals to achieve the driving of the audio to the face semantic prior; use the trained avatar rendering model to output the rendering diagram of the real-time digital human avatar.

[0079] It should be noted that the modeling of the solution proposed in the embodiment of the present invention is short and includes two parts. The first part first initializes the static Gaussian function using the input image head point cloud generated by 3DMM, and then uses a set of driving signals to train the head modeling on the time series, so as to establish the mapping relationship between the embedded label and the head dynamic Gaussian function. The second stage focuses on how to align the input audio and the embedded label, and by training the accuracy of the change of the embedded code driven by the change of the audio in the audio feature extraction model, finally, it is possible to realize the deformation of the head dynamic Gaussian function driven by the mapping relationship established in the first part, and then generate the talking head video corresponding to the input audio.

[0080] In one embodiment, in the avatar rendering model training stage, the steps of extracting the 3DMM point cloud data of the avatar include:

[0081] The 3DDFA model screens N avatar pictures to obtain a standard head image, and extracts the 3DMM point cloud data from the standard head image.

[0082] In one embodiment, in the avatar rendering model training stage, the steps of extracting the key face features as the driving signals include:

[0083] The 3DDFA model extracts the facial expression feature f exp from each avatar picture as the driving signal.

[0084] In one embodiment, in the avatar rendering model training stage, the 3DDFA model is also used to extract the head pose feature f pos from each avatar picture.

[0085] In one embodiment, in the avatar rendering model training stage, the static Gaussian distribution representation process of the avatar picture is as follows:

[0086] For a given set of avatar pictures, the 3DGS uses a set of sparse spatial points to initialize the Gaussian function, and adaptively optimizes the density and position of the Gaussian function during the training process. A standard 3DGS function can be expressed as follows:

[0087] ;

[0088] where p represents the position (center point) of the Gaussian function, represents the covariance of the Gaussian function. Since It is only meaningful in the positive semi - definite case. Therefore, 3DGS uses the following equivalent representation for easier generation and optimization:

[0089] ;

[0090] For the fast rasterization of Gaussian functions, 3DGS uses an affine transformation to map the 3DGS in the camera coordinate system to a special ray space and then projects it onto a two - dimensional plane:

[0091] ;

[0092] where W represents the camera view matrix and J is the Jacobian matrix. By ignoring the third row and column of the covariance matrix, the final result for rasterization can be obtained. Finally, 3DGS can be represented by the following formula:

[0093] ;

[0094] where c represents the color corresponding to the Gaussian function, which is represented by a set of 3 - order spherical harmonic functions in this embodiment, represents the transparency.

[0095] In one embodiment, in the avatar rendering model training stage, the steps of calculating the spatial position semantic information of the Gaussian distribution include:

[0096] Calculating the point coordinates of the spatial center points of all Gaussian ellipsoidal domains in the static Gaussian distribution;

[0097] Performing Fourier encoding on the point coordinates to obtain the semantic information f s .

[0098] In one embodiment, in the avatar rendering model training stage, based on 3DGS, a learnable embedding label f s is additionally introduced into the Gaussian function to construct the representation basis of the face. It includes:

[0099] The 3DDFA model composites the facial expression feature f exp and the head pose feature f pos to generate f w :

[0100] ;

[0101] Different parts of the human face often have different degrees of association with the same input action control latent code. For example, lip movements change more than those in other facial regions. If only an MLP is used as the action driver, it may cause the action control latent code to be parsed into ambiguous and unclear results, making the algorithm unable to converge stably and producing blurred results. This embodiment proposes a lightweight local semantic attention mechanism based on semantic tags.

[0102] As Figure 4 shown, this embodiment uses a two-layer MLP to form an attention mechanism. The action controller uses a two-layer MLP to form an attention mechanism, inputting f w into an MLP to obtain the encoded vector f a , combining the semantic information f s , and using an attention mechanism along the dimension to obtain the control vector f v :

[0103] ;

[0104] ;

[0105] This process can be understood as a process of attention allocation. The learnable f s can be regarded as the semantic information of the Gaussian function. Regions with a higher degree of association with facial expression features are assigned greater weights, while regions with a lower degree of association are assigned lower weights, thus achieving more precise facial action control.

[0106] Obtain the attention weights through the action controller and get the spatial position offset p, the rotation offset r, and the change in the size of the ellipsoidal domain of the Gaussian distribution s:

[0107] .

[0108] In dynamic human face modeling, DSG uses a lightweight action controller based on the local semantic attention mechanism to dynamically change the appearance of the Gaussian function ( ), thus achieving the facial action control of the generated character. Taking 's composite features as conditions to predict the spatial position offset p and the rotation offset r, and the size of the semi-major axis of the ellipsoidal domain of the Gaussian distribution s, thus realizing the dynamic change of the human face.

[0109] ;

[0110] Finally, the obtained dynamic DSG can be summarized as the following five - tuple:

[0111] .

[0112] In this embodiment, by embedding the label f s more abundant information is expressed. At the same time, the proposed local semantic attention mechanism is combined to obtain a clearer rendering result. The DSG is initialized by using a standard 3DMM model to extract spatial points, as Figure 1 shown. The DSG is initially initialized on the face surface and then optimized using the adaptive density control in 3DGS to achieve automatic parameter adjustment of the avatar rendering model.

[0113] In one embodiment, in S105, the loss function is also used to reversely optimize and update the 3DGS parameters and the action controller parameters.

[0114] The 3DGS parameters include: covariance, color, and transparency; the action controller parameters include MLP parameters.

[0115] In one embodiment, in S105, the steps of calculating the loss function according to the rendered avatar picture and the current avatar picture include:

[0116] The loss function is expressed as follows:

[0117] ;

[0118] where λ 1 and λ 2 represent and weights respectively;

[0119] The image texture loss from coarse to fine is introduced. Coarsely speaking, each frame of the generated video should be consistent with the original picture. Therefore, this loss function is:

[0120] ;

[0121] where is the rendered avatar picture, is the corresponding avatar picture;

[0122] To further improve the authenticity of the generated picture, a finer - grained loss function based on intuitive perception is introduced:

[0123] ;

[0124] where the features of the rendered avatar picture and the corresponding avatar picture are extracted using the VGG19 network, Denotes the output of the i-th layer of the VGG19 network.

[0125] In one embodiment, in S200, the steps of extracting audio coding features include:

[0126] The audio feature extraction model segments the audio according to the frame rate of the head dynamic action video frames to obtain an audio sequence containing N audio segments;

[0127] Use DeepSpeech to extract the features of the audio sequence and combine with the audio encoder for feature coding; the calculation result of the loss function is used to reverse-optimize and update the audio encoder.

[0128] The audio encoder uses a temporal attention network for action coding to obtain rich context features, so as to achieve a smoother action control effect.

[0129] In this embodiment, in order to solve the huge cost brought by training on large-scale data in existing methods (such as StyleGAN) and the underfitting problem caused by training on a small amount of data (such as AD-NeRF), this embodiment is based on the data augmentation strategy proposed in DEP-Talker, and randomly cuts video segments of different lengths and positions for data augmentation. In addition, a new augmentation strategy is introduced to adjust the video rate at different rates of 0.8, 1.0, and 1.25 to enhance the adaptability of the model to different speech rates.

[0130] In one embodiment, in S201, the steps of calculating the loss function according to the facial key features and the audio coding features include:

[0131] The loss function calculates the loss value between the audio coding feature of each audio segment and the facial key feature of the avatar picture of its corresponding frame respectively.

[0132] In this embodiment, as Figure 5 shown, use the L2 norm as the loss

[0133] ;

[0134] where f is the facial key feature, including the facial expression feature f exp , is the audio coding feature.

[0135] To further illustrate the technical effects of the present invention, the following specific comparative experimental examples of the present invention are given:

[0136] This experiment was tested on the dataset provided by AD-NeRF. All videos in the dataset were resampled to 25fps, and the sampling rate of the audio was also adjusted to 16000HZ.

[0137] Set the depth D of the action encoder M_action to 5 and the hidden layer dimension W to 256. M_att is set to 2 layers with the hidden layer dimension set to 128. The Adam optimizer is used, and the learning rate experiences exponential decay from 1.6e^(-4) to 1.6e^(-6).

[0138] In this experiment, two mainstream research - commonly used metrics, SSIM and PSNR, are used to evaluate the quality of talking - head generation. For the accuracy of mouth shape and lip synchronization, Landmark Distance (LMD) and ConfidenceScore (Conf) are used for evaluation.

[0139] Comparisons are made with two existing methods, including IP_LAP, one of the state - of - the - art methods based on GAN, and representative methods based on NeRF, AD - NeRF and ER - NeRF.

[0140] Table 1 Performance comparison of different methods

[0141]

[0142] According to Table 1, a detailed analysis and comparison of the performance of each method are as follows:

[0143] Image rendering quality: Regarding the image rendering quality index PSNR (Peak Signal - to - Noise Ratio) and the image structure recovery index SSIM (Structural Similarity Index Measure), the method of the present invention achieves results similar to the current state - of - the - art methods based on NeRF (Neural Radiance Fields). This means that the method of the present invention is comparable to the state - of - the - art methods in terms of the clarity and structure preservation of the generated images.

[0144] Lip - synchronization metrics: In terms of lip - synchronization metrics, the IP - LAP method achieves the best results. However, IP - LAP requires training using a large - scale talking - head dataset.

[0145] Generation efficiency: The method of the present invention achieves a rendering efficiency of over 100 FPS at a resolution of 480×480 on an RTX3090 24GB GPU. In contrast, the state - of - the - art method based on NeRF, ER - NeRF, can only reach 25 FPS, and the rendering efficiency of IP - LAP is even lower than 10 FPS. This means that the method of the present invention has significant advantages in terms of real - time performance and processing speed, and is more suitable for application scenarios that require fast generation and rendering.

[0146] Model complexity: The method of the present invention has relatively low model complexity, thus achieving higher generation efficiency. This means that the method of the present invention is more easily deployable on devices with limited resources, further expanding its application scope.

[0147] Dataset size: Although the method of the present invention is slightly inferior to IP-LAP in lip synchronization, the dataset used in the present invention is smaller, and under the same dataset size, it can meet the same lip synchronization index requirements as IP-LAP.

[0148] Trade-off between real-time performance and accuracy: In real-time generation and rendering applications, the method of the present invention provides the possibility for real-time applications by improving the generation efficiency.

[0149] In summary, the method of the present invention has achieved results comparable to the current state-of-the-art methods in terms of image rendering quality, action accuracy, and structure recovery, while having a significant advantage in terms of generation efficiency.

[0150] To further verify the effectiveness of the method of the present invention, an ablation experiment was conducted on the embodiments. In the ablation experiment, the processing method of the dataset and the network structure were kept unchanged. First, a single position encoding method was used to replace the learnable embedding f proposed by the present invention. s Then, the action control code and the position code were directly concatenated. Finally, the concatenated result was sent to a multi-layer perceptron (MLP) for processing. The experimental results obtained under the above experimental conditions were named "Ours-P".

[0151] To better verify the influence of semantic attention on the experimental results, an additional experiment was conducted under the experimental conditions of "Ours-P", that is, the semantic attention module was removed. Therefore, only position encoding was used in "Ours-P" and it was directly sent to the MLP for action prediction. This ablation experiment was labeled as "Ours-Att".

[0152] Table 2 Performance comparison of ablation experiments

[0153]

[0154] As shown in Table 2, the clarity of the mouth area is significantly reduced. This observation further verifies the importance of semantic attention in improving image rendering quality, especially in capturing and presenting facial details. By maintaining the semantic attention mechanism, the method of the present invention can effectively capture and render clearer and more realistic facial features, including the mouth area. After removing the learnable embedding f s not only does the quality of the rendered image (measured by the PSNR metric) decrease, but also the accuracy of action synchronization is affected. This is mainly because, after removing f sAfter that, only using positional encoding cannot well represent the semantics of DSG (possibly referring to a certain data structure or feature).

[0155] From the experimental results in Table 2 and Figure 6 it can be observed that although removing the learnable embedding f s and the semantic attention mechanism can reduce the computational complexity (the output frame rate has increased), there are many blurred and distorted parts in the generated human head images. Combining the above experimental results, it can be seen that each component of the semantic attention plays a positive role in the model performance.

[0156] The above embodiments of the present invention can produce high-fidelity lip-sync results in single-shot and few-shot settings. It has unique technical advantages:

[0157] (1) The rendering method based on 3DGS achieves the best current training and rendering efficiency;

[0158] (2) Utilize the advantages of the stable 3DMM for optimization, thereby obtaining more accurate modeling results;

[0159] (3) The proposed local semantic attention mechanism can perform more accurate facial motion control.

[0160] The above has introduced in detail a method for generating a digital human head driven by real-time audio provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0161] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

Claims

1. A method for generating a digital human portrait based on real-time audio drive, characterized in that: The steps include: Avatar rendering model training phase: S100: extracting N head portrait images of consecutive frames according to the dynamic action video frames of the head, where N>1; S101: inputting the N head portrait images into the 3DDFA model, extracting the 3DMM point cloud data of the head portrait, and N groups of key facial features corresponding to the N head portrait images as driving signals; S102: The 3DMM point cloud data is initialized by 3DGS to obtain a static Gaussian distribution describing the 3DMM point cloud data, calculate the spatial position semantic information of the Gaussian distribution and use it as an embedded label of the static Gaussian distribution to generate a dynamic Gaussian function of the avatar; S103: the driving signal and the spatial position semantic information corresponding to the current avatar image are input into the action controller, and the avatar action offset is predicted and output; S104: after the avatar motion offset is superimposed and updated to the avatar dynamic Gaussian function, the updated avatar dynamic Gaussian function is projected onto a two-dimensional image plane corresponding to the Gaussian distribution in the 3D space to obtain a rendered avatar image; S105: Calculate a loss function based on the rendered avatar image and the current avatar image, and after reversely optimizing and updating the spatial position semantic information based on the loss function calculation result, repeat S103-S105 until the training of the N avatar images is completed to obtain a trained avatar rendering model; The step of calculating the loss function includes: The loss function is expressed as follows: Among them, λ1 and λ2 represent and Weight; in, is the rendered avatar picture, and I is the corresponding avatar picture; Among them, the VGG19 network is used to extract the features of the rendered avatar image and the corresponding avatar image. represents the output of the i-th layer of the VGG19 network; Audio feature extraction model training phase: S200: Inputting audio of a given duration into an audio feature extraction model to extract audio coding features; S201: Calculate a loss function according to the key features of the face and the audio coding features, and reversely optimize and update the audio feature extraction model according to the loss function calculation result to obtain a trained audio feature extraction model; Real-time audio-driven digital portrait generation stage: The trained avatar rendering model retrieves the 3DMM point cloud data and the driving signal; The real-time audio is input into the trained audio feature extraction model to extract the real-time audio coding features, and the real-time audio coding features are aligned and replaced with the driving signal to realize the audio driving of the facial semantic prior; the trained avatar rendering model is used to output the rendering image of the real-time digital human avatar.

2. The method for generating a digital human portrait based on real-time audio drive according to claim 1, characterized in that: In the avatar rendering model training stage, the step of extracting 3DMM point cloud data of the avatar includes: The 3DDFA model screens the N head portrait images to obtain a standard human head image, and extracts 3DMM point cloud data from the standard human head image.

3. The method for generating a digital human portrait based on real-time audio drive according to claim 1, characterized in that: In the avatar rendering model training stage, the step of extracting key features of the face as a driving signal includes: The 3DDFA model extracts facial expression features f for each avatar image. exp as a driving signal.

4. The method for generating a digital human portrait based on real-time audio drive according to claim 3, characterized in that: In the avatar rendering model training phase, the 3DDFA model is also used to extract head posture features f for each avatar image. pos .

5. The method for generating a digital human portrait based on real-time audio drive according to claim 1, characterized in that: In the avatar rendering model training phase, the step of calculating the spatial position semantic information of the Gaussian distribution includes: Calculate the point coordinates of the spatial center points of all Gaussian ellipsoid domains in the static Gaussian distribution; Perform Fourier encoding on the point coordinates to obtain semantic information f s .

6. The method for generating a digital human portrait based on real-time audio drive according to claim 4, characterized in that: In the avatar rendering model training phase, step S103 includes: The 3DDFA model is used to model the facial expression feature f exp and the head posture feature f pos Compound to generate f w : f w =Cat(f exp ,f pos ); The action controller uses a two-layer MLP to form an attention mechanism. w Input to an MLP att To obtain the encoding vector f a , combined with the semantic information f s , using an attention mechanism along the dimension to obtain the control vector f v : f a =MLP att (f w ); f v =f a ⊙f s ; The attention weight f is obtained by the action controller a , and obtain the spatial position offset △p, rotation offset △r, and Gaussian distribution ellipsoid domain size change △s: △p,△r,△s=MLP action (f v )。 7. The method for generating a digital human portrait based on real-time audio drive according to claim 1, characterized in that: In S105, the loss function is also used to reversely optimize and update the 3DGS parameters and the motion controller parameters.

8. The method for generating a digital human portrait based on real-time audio drive according to claim 1, characterized in that: In S200, the step of extracting audio coding features includes: The audio feature extraction model segments the audio according to the frame rate of the human head dynamic action video frame to obtain an audio sequence including N audio segments; DeepSpeech is used to extract features of the audio sequence and combined with an audio-motion encoder for feature encoding; the loss function calculation result is used to reversely optimize and update the audio-motion encoder.

9. The method for generating a digital human portrait based on real-time audio drive according to claim 4, characterized in that: In S201, the step of calculating the loss function according to the key features of the face and the audio coding features includes: The loss function respectively calculates the loss value between the audio coding features of each audio segment and the key facial features of the avatar image of its corresponding frame.

Citation Information

Patent Citations

  • Monocular face avatar generation method based on Gaussian point rendering

    CN117974867A

  • Three-dimensional face generating and driving method based on Gaussian splashing method

    CN118071898A