A High-Fidelity Voice-Driven Digital Human Synthesis Method Based on 3DGS

By introducing optimized global prompt modules, progressive conditional attribute prediction networks and dual discriminator architectures, combined with 3D Gaussian Splatting technology, the problems of structural drift and unnatural expressions in the existing methods are solved, and efficient and high-fidelity voice-driven digital human head synthesis is achieved, suitable for fields such as virtual reality and augmented reality.

CN119991888BActive Publication Date: 2025-07-29JIANGSU XUNGAO INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457933.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing 3DGS-based speech-driven digital human head synthesis method has problems such as structural drift, improper attribute processing, difficulty in taking into account expression and efficiency, and lack of fine-grained expression control, resulting in low synthesis efficiency, low quality and unnatural animation.

Method used

Adopting the optimized global prompt module, the progressive conditional attribute prediction network module and the dual discriminator architecture, the digital human animation generation is driven through voice signals, and combined with 3D Gaussian Splatting technology, static and dynamic digital human models are built to achieve high-fidelity and efficient three-dimensional digital human reconstruction.

Benefits of technology

It significantly improves the rendering efficiency and image quality of voice-driven digital human animations, ensures the stability of animation structure and natural expressions, and realizes real-time and high-fidelity synthesis, which is suitable for fields such as virtual reality and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991888B_ABST
    Figure CN119991888B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-fidelity voice-driven digital human synthesis method based on 3DGS. First, a static digital human model is trained, constructed based on 3D Gaussian Splatting, and the image quality is improved using a spatial discriminator to capture the basic shape and appearance of the digital human. Subsequently, a dynamic driving network is trained, which includes an optimizable global prompt, a progressive conditional attribute prediction network module, and a dual discriminator architecture. Among them, the optimizable global prompt module is used to stabilize the facial geometry of the digital human and prevent drift during the animation process; the progressive conditional attribute prediction network module is used to efficiently and temporally coherently predict the dynamic Gaussian parameters of the digital human model; the dual discriminator architecture module is used to improve the realism and temporal consistency of the synthesized digital human animation. The present invention is applicable to voice-driven digital human animation synthesis, can effectively improve the realism, efficiency, and structural coherence of the synthesized digital human animation, and achieve real-time rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human reconstruction, and particularly to a high-fidelity voice-driven digital human synthesis method based on 3DGS. Background Art

[0002] The voice-driven digital human head synthesis technology, as an important innovative direction in the fields of computer graphics and artificial intelligence, has attracted much attention in recent years. It shows great potential in fields such as virtual reality, augmented reality, online education, intelligent customer service, and digital media content creation. Its aim is to generate natural and realistic digital human head animations in real time according to voice signals, realizing intelligent human-computer interaction and vivid content expression. Therefore, this technology has become a key component of modern digital content generation and human-computer interaction systems, and is also a research hotspot in the fields of computer graphics and artificial intelligence.

[0003] Early research mainly relied on two-dimensional generative adversarial networks (2D GANs). Although progress was made in audio-visual synchronization and generating images with acceptable quality, 2D GANs were difficult to ensure the consistency of three-dimensional heads under perspective changes and were also difficult to construct real three-dimensional structures, restricting their application in immersive interaction scenarios. Subsequent research tried to introduce three-dimensional information such as three-dimensional model parameters (3DMM) or facial key points (Landmark) to improve the control accuracy, but the errors in preprocessing and estimation affected the authenticity and stability of the synthesis results.

[0004] In recent years, the neural radiance field (NeRF) technology has achieved view-consistent three-dimensional head modeling and rendering. The audio-driven neural radiance field (Audio-driven NeRF) further integrates voice signals to achieve higher-quality voice-driven digital human head synthesis. Existing research has improved the rendering efficiency and image quality through strategies such as audio-conditioned input and mesh optimization. However, NeRF still has problems such as low rendering efficiency, difficulty in real-time application, and room for improvement in the control of fine expressions and lip details. At the same time, NeRF is prone to blurred lip shapes and dull eyes when dealing with complex lip shape changes and subtle expressions, restricting the vividness and naturalness of digital human heads.

[0005] As an emerging explicit three-dimensional scene representation, 3D Gaussian Splatting (3DGS) has attracted much attention for its fast rendering and excellent quality. Applying 3DGS to the synthesis of speech-driven digital human heads is expected to improve efficiency while maintaining or even surpassing the image quality of NeRF. Existing research has preliminarily verified the potential of 3DGS in the real-time high-quality synthesis of speech-driven digital human heads. However, existing 3DGS-based methods for speech-driven digital human head synthesis still face challenges: the lack of global structure guidance is likely to cause facial drift and unstable animations; the processing method of Gaussian basis element attributes needs to be improved. When predicting and controlling attributes such as position, scale, rotation, and color, there is a lack of effective modeling of the internal dependencies and generation order, and geometric structure instability and unnatural expressions are likely to occur when dealing with complex mouth shapes and fine expressions; there is still a need for a better balance among synthesis efficiency, image quality, and fine expression control. For example, existing methods usually process Gaussian basis element attributes independently, ignoring the integrity and coherence of the facial structure; when predicting Gaussian basis element attributes, there is a lack of effective modeling of the dependency relationships and generation order among attributes. Summary of the Invention

[0006] The object of the present invention is to provide a high-fidelity speech-driven digital human synthesis method based on 3DGS, aiming to solve the problems existing in the prior art in the synthesis of speech-driven three-dimensional digital human heads, such as structural drift, improper attribute processing, difficulty in balancing expressiveness and efficiency, and lack of fine-grained expression control, and finally realizing real-time three-dimensional digital human reconstruction with high fidelity, high efficiency, and structural coherence.

[0007] To achieve the above functions, the present invention designs a high-fidelity speech-driven digital human synthesis method based on 3DGS, which executes the following steps S1 - S3 to generate a digital human animation driven by a speech signal:

[0008] Step S1: Perform feature encoding and static Gaussian parameter prediction for the digital human to construct a static digital human model. Use 3D Gaussian Splatting software to capture and render the basic shape and appearance of the digital human for the static digital human model, and train the static digital human model using the backpropagation method to obtain a trained static digital human model;

[0009] Step S2: Construct and train a speech-driven digital human synthesis system, including an optimizable global hint module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; among them, the optimizable global hint module generates a global hint input, and the progressive conditional attribute prediction network module takes an audio signal, expression parameters, view parameters, and the global hint input as inputs, predicts dynamic deformations in stages, combines the dynamic deformations with the static digital human model to obtain a dynamic digital human model, and the dual discriminator architecture module discriminates the dynamic digital human model from real dynamic face images;

[0010] Step S3: Input the voice signal into the trained voice-driven digital human synthesis system to output the voice-driven digital human animation, completing the synthesis of the digital human animation.

[0011] Beneficial effects: The present invention proposes a high-fidelity voice-driven digital human synthesis method based on 3DGS, significantly improving the rendering efficiency of voice-driven digital human animations, reducing the computational cost, and achieving real-time and high-fidelity synthesis. The method innovatively introduces an optimizable global hint module to stabilize the facial geometry structure, reduce structural drift, and ensure the stability of the animation structure and the consistency of the facial contour. A progressive conditional attribute prediction network is constructed to efficiently and precisely predict dynamic Gaussian parameters, enhancing the richness of expressions and the control accuracy, and realizing natural and vivid expression animations. A dual discriminator architecture is designed, with the spatial and temporal discriminators collaborating to enhance the pixel-level realism and temporal consistency of the images, significantly improving the visual quality and realism of the animations. The present invention has achieved significant improvements in aspects such as rendering efficiency, image quality, structural stability, and expression control fineness, providing an efficient and high-quality solution for real-time high-fidelity voice-driven three-dimensional digital human synthesis, and can be widely applied in fields such as virtual reality and augmented reality. Description of the Drawings

[0012] Figure 1 is a flowchart of a high-fidelity voice-driven digital human synthesis method based on 3DGS provided by an embodiment of the present invention;

[0013] Figure 2 is a schematic diagram of the optimizable global hint finally learned according to an embodiment of the present invention. Detailed Embodiments

[0014] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0015] A high-fidelity voice-driven digital human synthesis method based on 3DGS provided by an embodiment of the present invention, referring to Figure 1 , performs the following steps S1 - S3 to generate a digital human animation driven by a voice signal:

[0016] Step S1: Perform feature encoding and static Gaussian parameter prediction on the digital human, construct a static digital human model, capture and render the basic shape and appearance of the digital human using 3D Gaussian Splatting software for the static digital human model, and train the static digital human model using the backpropagation method to obtain a trained static digital human model;

[0017] The specific steps of Step S1 are as follows:

[0018] Step S1.1: The digital human spatial position Input multi-resolution three-planes Perform encoding to obtain a feature vector , where the multi-resolution three-planes are composed of three orthogonal 2D feature grids . The shape of each 2D feature grid is , where H represents the feature hidden dimension and R represents the dimension resolution;

[0019] Step S1.2: Input the feature vector into the static network . The static network is constructed based on a multi-layer perceptron, and maps the feature vector to static Gaussian parameters , including the mean position , the mean scale , the mean rotation , the mean spherical harmonic coefficients and the mean opacity value . The static Gaussian parameters constitute the static digital human model ;

[0020] Step S1.3: Use 3D Gaussian Splatting rendering software to render the static digital human model based on the static Gaussian parameters to generate a synthesized static face image ;

[0021] Step S1.4: Input the synthesized static face image and the real static face image into the spatial discriminator to obtain the discrimination result output by the spatial discriminator;

[0022] Step S1.5: Based on the discrimination result and the static loss function between the synthesized static face image and the real static face image , perform backpropagation to optimize and train the static digital human model to obtain a trained static digital human model.

[0023] The static loss function in Step S1.5 includes a segment importance balance loss function , a structural similarity D-SSIM loss function , a perceptual similarity LPIPS loss function and an adversarial loss function , where:

[0024] Utilize the segment importance balance loss function Constrained synthetic static face image For the pixel-level photorealism, the formula for the segment importance balanced loss function is as follows:

[0025] ;

[0026] Among them, represents the number of pixels, represents the value of the th pixel in the synthetic static face image i , represents the value of the th pixel in the real face image i ;

[0027] Using the structural similarity D-SSIM loss function to constrain the synthetic static face image , the formula for the structural similarity D-SSIM loss function is as follows:

[0028] ;

[0029] Among them, represents the differentiable structural similarity D-SSIM function;

[0030] Using the perceptual similarity LPIPS loss function to constrain the perceptual realism of the synthetic static face image , the formula for the perceptual similarity LPIPS loss function is as follows:

[0031] ;

[0032] Among them, represents the number of network layers, represents the feature extraction operation of the j th layer of the pre-trained Alex-Net network, and respectively represent the height and width of the feature map of the j th layer, represents the pixel position in the feature map;

[0033] Using the adversarial loss function to perform adversarial training through the spatial discriminator to enhance the detailed realism of the synthetic static face image , the formula for the adversarial loss function is as follows:

[0034] ;

[0035] ;

[0036] ;

[0037] Among them, is the discrimination result of the spatial discriminator for the synthesized static face image , is the discrimination result of the spatial discriminator for the real static face image , is the label of the real static face image , is the label of the synthesized static face image . BCE represents the binary cross-entropy loss function, and MSE represents the mean square error loss function;

[0038] Static loss function The calculation formula is as follows:

[0039] ;

[0040] Among them, , , and respectively represent the weight coefficients of each loss function.

[0041] Step S2: Construct and train a speech-driven digital human synthesis system, including an optimizable global prompt module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; among them, the optimizable global prompt module generates global prompts, and the progressive conditional attribute prediction network module takes the audio signal, expression parameters, perspective parameters, and global prompts as inputs, predicts dynamic deformations in stages, combines the dynamic deformations with the static digital human model to obtain a dynamic digital human model, and the dual discriminator architecture module discriminates the dynamic digital human model and the real dynamic face image; the schematic diagram of the finally learned optimizable global prompt refers to Figure 2 ;

[0042] The specific steps of Step S2 are as follows:

[0043] Step S2.1: Input the audio signal a , expression parameter e, perspective parameter v, and the global prompt generated by the optimizable global prompt module into the progressive conditional attribute prediction network module to predict dynamic deformations in stages; among them is the global position offset, is the scale change amount, is the rotation adjustment amount, is the change amount of the opacity value, and is the change amount of the spherical harmonic coefficients;

[0044] The specific steps of step S2.1 are as follows:

[0045] Step S2.1.1: Predict the global position offset of the dynamic digital human model , and establish a spatial anchor point;

[0046] Step S2.1.2: Based on the global position offset , predict the scale change amount and the rotation adjustment amount of the dynamic digital human model, and refine the facial geometry;

[0047] Step S2.1.3: Based on the facial geometry, predict the opacity change amount of the dynamic digital human model, and optimize the detailed appearance;

[0048] Step S2.1.4: Based on the opacity change amount , predict the spherical harmonic coefficient change amount of the dynamic digital human model, and capture the lighting and material details.

[0049] Step S2.2: Combine the dynamic deformation with the trained static digital human model to obtain the dynamic Gaussian parameters . The dynamic Gaussian parameters constitute the dynamic digital human model;

[0050] Step S2.3: Use the 3D Gaussian Splatting software to render and generate a synthetic dynamic face image sequence based on the dynamic Gaussian parameters ;

[0051] Step S2.4: Input the synthetic dynamic face image sequence and the real dynamic face image sequence into the dual discriminator architecture module respectively. The dual discriminator architecture module includes a spatial discriminator and a temporal discriminator , and obtain the discrimination results output by the spatial discriminator and the temporal discriminator;

[0052] The discrimination of the spatial discriminator includes the following steps:

[0053] Step S2.4.1.1: Extract features of different scales of the input image through a convolutional neural network to obtain multi-scale features ;

[0054] Step S2.4.1.2: Input the multi-scale features into a multi-layer perceptron to obtain the discriminant results output by the spatial discriminator; where is the feature map obtained after the original resolution image passes through a convolutional neural network, is the feature map obtained after the original image is downsampled by 1 / 2 and then passes through a convolutional neural network, is the feature map obtained after the original image is downsampled by 1 / 4 and then passes through a convolutional neural network.

[0055] The discrimination of the temporal discriminator includes the following steps:

[0056] Step S2.4.2.1: Perform a 2D Fourier transform on the input synthetic dynamic face image sequence to extract frequency domain features , and then concatenate the original frame with its frequency domain features to obtain the concatenated features ;

[0057] Step S2.4.2.2: Apply 3D convolution to the concatenated features to capture temporal correlation features, and use a local attention module to refine the obtained temporal correlation features; enabling it to focus on key facial regions crucial for perceptual coherence to judge the temporal coherence of the animation sequence;

[0058] Step S2.4.2.3: Generate a temporal quality score by combining 2D convolution and a multi-layer perceptron.

[0059] Step S2.5: Based on the discriminant results and the dynamic loss function between the synthetic dynamic face image sequence and the real dynamic face image sequence , perform backpropagation to optimize and train the progressive conditional attribute prediction network module, the optimizable global hint module, and the dual discriminator architecture module to obtain a trained speech-driven digital human synthesis system.

[0060] The dynamic loss function described in Step S2.5 includes a static loss function , a temporal adversarial loss function and a temporal consistency loss function , where:

[0061] Use the static loss function to constrain the realism of the single-frame image of the synthesized static face image , and the static loss function ​The calculation formula is the same as the static loss function in step S1.5 The calculation formula is the same;

[0062] Use the temporal adversarial loss function Through the time discriminator Adversarial training to improve the temporal coherence of the synthesized dynamic face image sequence The calculation formula of the temporal adversarial loss function is the same as the adversarial loss function in step S1.5 The calculation formula is the same;

[0063] Use the temporal consistency loss function Explicitly constrain the inter-frame temporal consistency of the synthesized dynamic face image sequence to further improve the smoothness and coherence of the animation. The calculation formula of the temporal consistency loss function is as follows:

[0064] ;

[0065] where T represents the number of frames of the image sequence, and respectively represent the th t frame and the t +1th frame of the synthesized dynamic face image sequence and respectively represent the th t frame and the t +1th frame of the real dynamic face image sequence;

[0066] The calculation formula of the dynamic loss function is as follows:

[0067] ;

[0068] where and respectively represent the weight coefficients of the temporal adversarial loss function and the temporal consistency loss function.

[0069] Step S3: Input the speech signal into the trained speech-driven digital human synthesis system to output the speech-driven digital human animation, and complete the synthesis of the digital human animation.

[0070] The specific steps of step S3 are as follows:

[0071] Step S3.1: Input the speech signal, expression parameters, perspective parameters, and the global hint generated by the optimizable global hint module Input into the progressive conditional attribute prediction network module to predict dynamic deformation ;

[0072] Step S3.2: Combine the dynamic deformation with the static digital human model to obtain dynamic Gaussian parameters ;

[0073] Step S3.3: Use the 3D Gaussian Splatting software to generate voice-driven digital human animations in real time based on the dynamic Gaussian parameters .

[0074] To verify the effectiveness of the method of the present invention, the following comparative experiments and ablation experiments were carried out:

[0075] First, the used datasets and training details are introduced, then the comparative experiment results of different algorithms on the datasets are shown, and the effectiveness of the optimizable global prompt, progressive conditional attribute prediction network module, and dual discriminator architecture is evaluated through a series of ablation experiments. The training process of the model is divided into two stages: static model initialization and dynamic model training. In the static model initialization stage, training was carried out for 8000 iterations, and the batch size was set to 1. The weight coefficients of the loss function in this stage were set as follows:

[0076] The coefficient of the segment importance balance loss function was set to 0.8, the coefficient of the structural similarity D-SSIM loss function was set to 0.2, the coefficient of the perceptual similarity LPIPS loss function was set to 0.01, and the coefficient of the adversarial loss function was set to 0.01.

[0077] In the dynamic model training stage, the number of iterations was 20000, and the batch size was 16. The weight coefficients of the loss function in this stage were set as follows:

[0078] The coefficient of the temporal consistency loss function was set to 0.04, and the coefficient of the temporal adversarial loss function was set to 0.01.

[0079] The method proposed in the present invention was compared with several current mainstream 3D reconstruction technologies. The experimental dataset was sourced from RAD-NeRF. In the comparative experiment, the better results in the experimental data and pre-trained models were selected, covering the current state-of-the-art models. The experimental results are shown in Tables 1 and 2:

[0080] Table 1. Comparative Experiment (Self-Driving)

[0081]

[0082] Table 2. Comparative Experiments (Cross-Drive)

[0083]

[0084] To further verify the effectiveness of each module in the model, ablation experiments were conducted. In the experiments, the optimizable global prompt, the progressive conditional attribute prediction network module, and the dual discriminator architecture were removed respectively to evaluate the impact of these modules on the overall effect, and the performance of the model after removing different modules was compared with that of the complete model. The results of the ablation experiments are shown in Tables 3, 4, 5, and 6. In the tables indicates the retention of the corresponding module:

[0085] Table 3. Ablation Experiment for Learnable Global Prompt (Self-Drive)

[0086]

[0087] Table 4. Ablation Experiment for Learnable Global Prompt (Cross-Drive)

[0088]

[0089] Table 5. Ablation Experiment (Self-Drive)

[0090]

[0091] Table 6. Ablation Experiment (Cross-Drive)

[0092]

[0093] Among the evaluation metrics, PSNR and SSIM represent peak signal-to-noise ratio and structural similarity respectively, while LPIPS measures perceptual similarity. PSNR mainly reflects the clarity of the synthesized image, and the higher the value, the better the quality of the synthesized image; SSIM takes into account the structural similarity of the synthesized image, and the higher the value, the higher the structural retention of the synthesized image; LPIPS is used to measure the perceptual difference of the synthesized image, and the lower the value, the better the perceptual quality of the synthesized image. LMD is used to measure the deviation of the facial key point movement amplitude of the synthesized animation from the real animation, and the lower the value, the more accurate the movement amplitude of the synthesized animation; AUE is used to quantitatively evaluate the lip-sync accuracy of the synthesized animation, and the lower the value, the higher the lip-sync accuracy; Sync-C is used to evaluate the audio-visual synchronization confidence of the synthesized animation, and the higher the value, the better the audio-visual synchronization effect; FPS is used to measure the model inference speed and rendering efficiency, and the higher the value, the higher the model efficiency.

[0094] As can be seen from Table 1, under the self-driven scenario setting, the performance of this method in terms of the three metrics of PSNR, SSIM, and LPIPS is 33.51, 0.944, and 0.038 respectively, and it is particularly outstanding in terms of the frame rate (FPS) metric, reaching 167 FPS. These results indicate that this method can achieve higher-quality image rendering in the self-driven scenario while maintaining a significantly higher real-time rendering efficiency than other comparison methods. Compared with other comparison methods, this method achieves a better balance between image quality and efficiency, can greatly improve the rendering speed while ensuring that the image quality is close to the optimal level, and is more suitable for real-time interactive application scenarios. In addition, in terms of motion quality and lip-sync accuracy, the LMD metric of this method is 2.681, the AUE metric is 1.062, and the Sync-C metric is 6.205. All metrics are at the same level as advanced 3DGS-based methods such as TalkingGaussian and GaussianTalker, indicating that this method does not sacrifice motion quality and lip-sync accuracy while maintaining high rendering efficiency.

[0095] As can be seen from Table 2, under the cross-driven scenario setting, that is, when the model is driven by the audio of unseen subjects, this method still demonstrates excellent performance and robustness. On the two test sets of Testset A and Testset B, the three lip-sync related metrics of Sync, LMD, and AUE of this method are significantly better than 2D-driven methods such as Wav2Lip and PC-AVS, and are at a competitive level with advanced 3D-driven methods such as RAD-NERF, ER-NeRF, TalkingGaussian, and GaussianTalker. These results indicate that this method can still maintain stable and reliable lip-sync performance in the cross-driven scenario, reflecting the good generalization ability and robustness of the method of the present invention when dealing with unseen audio inputs. It is particularly worth mentioning that compared with other 3D-driven methods, this method still has a significant rendering efficiency advantage while maintaining the same or better lip-sync performance (as shown in Table 1), which makes this method more advantageous and potential in practical applications.

[0096] As can be seen from the data in Table 3 under the self-driven scenario, after introducing the optimizable global hint module, the proposed method has improvements in the PSNR, SSIM, and Sync-C metrics. The PSNR is increased to 33.48, the SSIM is increased to 0.944, and the Sync-C is significantly increased to 6.141, indicating that the optimizable global hint module can effectively improve the image quality and lip-sync accuracy in the self-driven scenario.

[0097] From the data in Table 4 under the cross-drive scenario, it can be seen that after introducing the optimizable global hint module, the proposed method has achieved more significant improvements in two lip-sync related metrics, namely LMD and Sync-C. The LMD metric remains the best, and Sync-C has significantly increased to 4.478, indicating that the optimizable global hint module plays a more prominent role in improving lip-sync accuracy and enhancing model robustness under the cross-drive scenario. From the data in Table 3 and Table 4 under the self-drive and cross-drive scenarios, it can be seen that on the basis of introducing the optimizable global hint module, by further introducing the progressive conditional attribute prediction network, the image quality metrics and lip-sync accuracy metrics of the proposed method in both scenarios have been further significantly improved, and the frame rate still remains at a highly competitive level.

[0098] From the data in Table 5 under the self-drive scenario, it can be seen that after gradually introducing the progressive conditional attribute prediction network, the spatial discriminator, and the temporal discriminator module, the PSNR, SSIM, and LPIPS metrics of the proposed method all show a continuous upward trend, indicating that each module has made a positive contribution to improving image quality. From the data in Table 6 under the cross-drive scenario, it can be seen that after gradually introducing the progressive conditional attribute prediction network, the spatial discriminator, and the temporal discriminator module, the LMD, AUE, and Sync-C metrics of the proposed method still show a continuous improvement trend, indicating that each module has also made an important contribution to improving the lip-sync performance under the cross-drive scenario. From the full-module ablation experiment data in Table 5 and Table 6 under the self-drive and cross-drive scenarios, it can be seen that each module proposed in the present invention, such as the optimizable global hint module, the progressive conditional attribute prediction network, and the dual discriminator architecture, has brought continuous and significant performance improvements to the proposed method in both scenarios, strongly verifying the effectiveness and necessity of each module.

[0099] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Without departing from the spirit of the present invention, various changes can be made within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A high-fidelity voice-driven digital human synthesis method based on 3DGS, characterized in that, Perform the following steps S1 - S3 to generate a digital human animation driven by a voice signal: Step S1: Perform feature encoding and static Gaussian parameter prediction for the digital human, construct a static digital human model, capture and render the basic shape and appearance of the digital human using the 3D Gaussian Splatting software for the static digital human model, and train the static digital human model using the backpropagation method to obtain a trained static digital human model; Step S2: Construct and train a voice-driven digital human synthesis system, including an optimizable global hint module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; among them, the optimizable global hint module generates global hints, the progressive conditional attribute prediction network module takes the audio signal, expression parameters, perspective parameters, and global hints as inputs, predicts dynamic deformations in stages, combines the dynamic deformations with the static digital human model to obtain a dynamic digital human model, and the dual discriminator architecture module discriminates the dynamic digital human model from real dynamic face images; The specific steps are as follows: Step S2.1: Input the audio signal a , the processed expression features e , the perspective feature v and the optimizable global hint into the progressive conditional attribute prediction network module to predict dynamic deformation in stages ; where is the global position offset,[[]] is the scale change amount,[[]] is the rotation adjustment amount,[[]] is the opacity value change amount,[[]] is the spherical harmonic coefficient change amount; Step S2.2: Combine the dynamic deformation with the trained static digital human model to obtain dynamic Gaussian parameters , is the average position, is the average scale, is the average rotation, is the average spherical harmonic coefficient, is the average opacity value; the dynamic Gaussian parameters constitute the dynamic digital human model; Step S2.3: Using 3D Gaussian Splatting software, based on the dynamic Gaussian parameters render and generate a synthetic dynamic face image sequence ; Step S2.4: Input the synthesized dynamic face image sequence and the real dynamic face image sequence into the dual discriminator architecture module respectively. The dual discriminator architecture module includes a spatial discriminator and a temporal discriminator , and obtain the discrimination results output by the spatial discriminator and the temporal discriminator; Step S2.5: Based on the discrimination result and the synthesized dynamic face image sequence and the real dynamic face image sequence the dynamic loss function between them , backpropagate to optimize and train the progressive conditional attribute prediction network module, the optimizable global prompt module, and the dual discriminator architecture module to obtain a trained voice-driven digital human synthesis system; Step S3: Input the voice signal into the trained voice-driven digital human synthesis system, output the voice-driven digital human animation, and complete the synthesis of the digital human animation.

2. The high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, wherein, The specific steps of Step S1 are as follows: Step S1.1: Input the digital human's spatial position into the multi-resolution tri-planar for encoding to obtain a feature vector , where the multi-resolution tri-planar is composed of three orthogonal 2D feature grids . The shape of each 2D feature grid is , where H represents the feature hidden dimension and R represents the dimension resolution; Step S1.2: Input the feature vector into the static network . The static network is constructed based on a multi-layer perceptron, and maps the feature vector to static Gaussian parameters , including the average position , average scale , average rotation , average spherical harmonic coefficients and average opacity value . The static Gaussian parameters constitute the static digital human model ; Step S1.3: Using 3D Gaussian Splatting rendering software, based on the static Gaussian parameters , render the static digital human model to generate a synthetic static face image ; Step S1.4: Input the synthesized static face image and the real static face image into the spatial discriminator to obtain the discrimination result output by the spatial discriminator; Step S1.5: Based on the discrimination result and the synthesized static face image and the real static face image the static loss function between them , backpropagate to optimize and train the static digital human model to obtain a trained static digital human model.

3. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 2, characterized in that The static loss function described in step S1.5 includes a segment importance balancing loss function , a structural similarity D-SSIM loss function , a perceptual similarity LPIPS loss function and an adversarial loss function , where: Fragment importance balanced loss function The calculation formula is as follows: ; Among them, represents the number of pixels, represents the synthesized static face image in the i value of the th pixel; represents the real face image i value of the th pixel. Structural Similarity D-SSIM Loss Function The calculation formula is as follows: ; Among them, represents the differentiable structural similarity D-SSIM function; Perceptual Similarity LPIPS Loss Function The calculation formula is as follows: ; Among them, represents the number of network layers, represents the feature extraction operation of the j th layer of the pre-trained Alex-Net network, and respectively represent the height and width of the feature map of the j th layer, represents the pixel position in the feature map; Adversarial loss function The calculation formula is as follows: ; ; ; Among them, is the spatial discriminator for the discriminant result of the synthesized static face image ; is the spatial discriminator for the discriminant result of the real static face image ; is the label of the real static face image ; is the label of the synthesized static face image ; BCE represents the binary cross-entropy loss function, and MSE represents the mean square error loss function; Static loss function The calculation formula is as follows: ; Among them, , , and respectively represent the weight coefficients of each loss function.

4. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that The specific steps of Step S2.1 are as follows: Step S2.1.1: Predict the global position offset of the dynamic digital human model ; Step S2.1.2: Based on the global position offset predict the scale change amount and the rotation adjustment amount of the dynamic digital human model to refine the facial geometry; Step S2.1.3: Predict the change amount of the opacity of the dynamic digital human model based on the facial geometry ; Step S2.1.4: Based on the opacity change amount , predict the change amount of the spherical harmonic coefficients of the dynamic digital human model .

5. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that, The discriminator in step S2.4 The discrimination includes the following steps: Step S2.4.1.1: Extract features of different scales of the input image through a convolutional neural network to obtain multi-scale features ; where is the feature map obtained after the original resolution image passes through the convolutional neural network, is the feature map obtained after the original image is downsampled by 1 / 2 and then passes through the convolutional neural network, is the feature map obtained after the original image is downsampled by 1 / 4 and then passes through the convolutional neural network; Step S2.4.1.2: Input the multi-scale features into a multi-layer perceptron to obtain the discrimination result output by the spatial discriminator.

6. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that The time discriminator in step S2.4 The discrimination includes the following steps: Step S2.4.2.1: Perform 2D Fourier transform on the input synthetic dynamic face image sequence to extract frequency domain features , and then splice the original frame with its frequency domain features to obtain the spliced features ; Step S2.4.2.2: Apply 3D convolution to the splicing feature to capture temporal correlation features, and use a local attention module to refine the obtained temporal correlation features; Step S2.4.2.3: Generate corresponding scores by combining two-dimensional convolution and a multi-layer perceptron.

7. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that The dynamic loss function described in step S2.5 includes a static loss function , a temporal adversarial loss function and a temporal consistency loss function , where: Static loss function The calculation formula of which is the same as that of the static loss function in step S1.5 The calculation formula is consistent; Temporal adversarial loss function has the same calculation formula as the adversarial loss function in step S1.5; Temporal Consistency Loss Function The calculation formula is as follows: ; Among them, T represents the number of frames of the image sequence, and respectively represent the th t frame and the t +1th frame of the synthesized dynamic face image sequence and respectively represent the th t frame and the t +1th frame of the real dynamic face image sequence; Dynamic loss function The calculation formula is as follows: ; Among them, and respectively represent the weight coefficients of the temporal adversarial loss function and the temporal consistency loss function.

8. A high-fidelity voice-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that, The specific steps of Step S3 are as follows: Step S3.1: Input the voice signal, expression parameters, perspective parameters, and the global hint generated by the optimizable global hint module into the progressive conditional attribute prediction network module to predict dynamic deformation ; Step S3.2: Combine the dynamic deformation with the static digital human model to obtain the dynamic Gaussian parameters ; Step S3.3: Using 3D Gaussian Splatting software, based on the dynamic Gaussian parameters Generate voice-driven digital human animations in real time.

Citation Information

Patent Citations

  • Digital human head portrait generation method based on real-time audio driving

    CN119006663A

  • High-definition digital human video generation method and system for customizing real human image, storage medium and equipment

    CN119729130A