High-fidelity voice-driven digital human synthesis method based on 3DGS

By adopting a 3DGS-based method in voice-driven digital human synthesis, combined with the optimized global prompt module and the progressive conditional attribute prediction network module, the problem of structural drift and improper attribute processing is solved, and efficient and real digital human animation synthesis is achieved.

CN119991888AActive Publication Date: 2025-05-13JIANGSU XUNGAO INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510457933.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art has problems such as structural drift, improper attribute processing, difficulty in taking into account both expression and efficiency, and lack of fine-grained expression control in speech-driven three-dimensional digital human head synthesis.

Method used

A high-fidelity voice-driven digital human synthesis method based on 3DGS is adopted to build a static digital human model through feature coding and static Gaussian parameter prediction. Combining the optimized global prompt module, an incremental conditional attribute prediction network module and a dual discriminator architecture module, dynamic deformation prediction and synthesis are realized.

Benefits of technology

It significantly improves the rendering efficiency of voice-driven digital human animations, reduces calculation costs, and realizes real-time and high-fidelity synthesis, ensures the stability of the animation structure and the consistency of facial contours, improves the richness of expressions and control accuracy, and significantly improves the visual quality and authenticity of animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991888A_ABST
    Figure CN119991888A_ABST
Patent Text Reader

Abstract

The invention discloses a high-fidelity voice-driven digital human synthesis method based on 3DGS, and the method comprises the steps: firstly, training a static digital human model, carrying out the construction based on 3D Gaussian Splitting, improving the image quality through a space discriminator, and capturing the basic shape and appearance of a digital human; then, a dynamic driving network is trained, the dynamic driving network comprises an optimizable global prompt module, a progressive condition attribute prediction network module and a double discriminator framework, and the optimizable global prompt module is used for stabilizing the geometric structure of the digital human face and preventing drifting in the animation process; the progressive condition attribute prediction network module is used for efficiently predicting dynamic Gaussian parameters of the digital human model in a time sequence coherent manner; and the double-discriminator architecture module is used for improving the sense of reality and time consistency of the synthesized digital human animation. The method is suitable for voice-driven digital human animation synthesis, the sense of reality, the efficiency and the structural continuity of the synthesized digital human animation can be effectively improved, and real-time rendering is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human reconstruction, and in particular to a high-fidelity speech-driven digital human synthesis method based on 3DGS. Background Art

[0002] As an important innovation direction in the field of computer graphics and artificial intelligence, speech-driven digital head synthesis technology has attracted much attention in recent years. It has shown great potential in the fields of virtual reality, augmented reality, online education, intelligent customer service and digital media content creation. It aims to generate natural and realistic digital head animations in real time based on speech signals, and realize intelligent human-computer interaction and vivid content expression. Therefore, this technology has become a key component of modern digital content generation and human-computer interaction systems, and is also a research hotspot in the fields of computer graphics and artificial intelligence.

[0003] Early research mainly relied on two-dimensional generative adversarial networks (2D GANs). Although progress has been made in audio and video synchronization and generating images of acceptable quality, 2D GANs have difficulty ensuring the consistency of the three-dimensional head under changes in perspective, and it is also difficult to construct a true three-dimensional structure, which limits its application in immersive interactive scenarios. Subsequent studies have attempted to introduce three-dimensional information such as three-dimensional model parameters (3DMM) or facial landmarks to improve control accuracy, but preprocessing and estimation errors affect the authenticity and stability of the synthesis results.

[0004] In recent years, NeRF technology has achieved 3D head modeling and rendering with consistent perspective. Audio-driven NeRF further integrates speech signals to achieve higher-quality speech-driven digital head synthesis. Existing research has improved rendering efficiency and image quality through strategies such as audio conditional input and mesh optimization. However, NeRF still has problems such as low rendering efficiency, difficulty in real-time application, and room for improvement in fine expression and lip shape detail control. At the same time, NeRF is prone to fuzzy mouth shape and dull eyes when dealing with complex lip shape changes and subtle expressions, which limits the vividness and naturalness of digital human heads.

[0005] As an emerging explicit three-dimensional scene representation, 3D Gaussian Splatting (3DGS) has attracted much attention for its fast rendering and excellent quality. 3DGS is applied to speech-driven digital head synthesis, which is expected to improve efficiency while maintaining or even exceeding NeRF image quality. Existing research has preliminarily verified the potential of 3DGS in real-time, high-quality speech-driven digital head synthesis. However, existing speech-driven digital head synthesis methods based on 3DGS still face challenges: the lack of global structural guidance can easily lead to facial drift and animation instability; the processing method of Gaussian primitive attributes needs to be improved, and when predicting and controlling attributes such as position, scale, rotation, and color, there is a lack of effective modeling of intrinsic dependencies and generation order. When processing complex lip shapes and fine expressions, it is easy to have unstable geometric structures and unnatural expressions; synthesis efficiency, image quality, and expression control precision still need to be better balanced. For example, existing methods usually process Gaussian primitive attributes independently, ignoring the integrity and coherence of facial structure; when predicting Gaussian primitive attributes, there is a lack of effective modeling of dependencies between attributes and generation order. Summary of the invention

[0006] The purpose of the present invention is to provide a high-fidelity voice-driven digital human synthesis method based on 3DGS, aiming to solve the problems existing in the prior art in voice-driven three-dimensional digital human head synthesis, such as structural drift, improper attribute processing, difficulty in balancing expressiveness and efficiency, and lack of fine-grained expression control, and ultimately achieve real-time three-dimensional digital human reconstruction with high fidelity, high efficiency and structural coherence.

[0007] To achieve the above functions, the present invention designs a high-fidelity voice-driven digital human synthesis method based on 3DGS, and executes the following steps S1 to S3 to generate a digital human animation driven by a voice signal: Step S1: feature encoding and static Gaussian parameter prediction are performed on the digital human to construct a static digital human model. The basic shape and appearance of the digital human are captured and rendered using 3D Gaussian Splatting software. The static digital human model is trained using a back propagation method to obtain a trained static digital human model. Step S2: construct and train a voice-driven digital human synthesis system, including an optimizable global prompt module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; wherein the optimizable global prompt module generates a global prompt input, the progressive conditional attribute prediction network module takes an audio signal, an expression parameter, a viewing angle parameter, and a global prompt input as input, predicts dynamic deformation in stages, combines the dynamic deformation with a static digital human model, and obtains a dynamic digital human model, and the dual discriminator architecture module discriminates between the dynamic digital human model and a real dynamic face image; Step S3: Input the speech signal into the trained speech-driven digital human synthesis system, output the speech-driven digital human animation, and complete the synthesis of the digital human animation.

[0008] Beneficial effects: The present invention proposes a high-fidelity voice-driven digital human synthesis method based on 3DGS, which significantly improves the rendering efficiency of voice-driven digital human animation, reduces computing costs, and realizes real-time, high-fidelity synthesis. The method innovatively introduces an optimizable global prompt module to stabilize the facial geometry, reduce structural drift, and ensure the stability of the animation structure and the consistency of facial contours. A progressive conditional attribute prediction network is constructed to efficiently and accurately predict dynamic Gaussian parameters, improve expression richness and control accuracy, and achieve natural and vivid expression animation. A dual discriminator architecture is designed, and spatial and temporal discriminators work together to improve the pixel-level realism and temporal consistency of the image, significantly improving the visual quality and realism of the animation. The present invention has achieved significant improvements in rendering efficiency, image quality, structural stability, and expression control precision, providing an efficient and high-quality solution for real-time, high-fidelity voice-driven three-dimensional digital human synthesis, which can be widely used in virtual reality, augmented reality and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a flow chart of a high-fidelity voice-driven digital human synthesis method based on 3DGS according to an embodiment of the present invention; Figure 2 It is a schematic diagram of an optimizable global prompt finally learned according to an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0011] A high-fidelity speech-driven digital human synthesis method based on 3DGS is provided in an embodiment of the present invention. Figure 1 , perform the following steps S1 to S3 to generate a digital human animation driven by a voice signal: Step S1: feature encoding and static Gaussian parameter prediction are performed on the digital human to construct a static digital human model. The basic shape and appearance of the digital human are captured and rendered using 3D Gaussian Splatting software. The static digital human model is trained using a back propagation method to obtain a trained static digital human model. The specific steps of step S1 are as follows: Step S1.1: Position the digital human in space Input multi-resolution triplanar Encode and obtain the feature vector , where the multi-resolution triplane consists of three orthogonal 2D feature grids The shape of each 2D feature grid is , H represents the feature hidden dimension, and R represents the dimension resolution; Step S1.2: Transform the feature vector Enter a static network , the static network Based on the multi-layer perceptron, the feature vector Mapping to static Gaussian parameters , including the average position , average scale , average rotation , average spherical harmonic coefficients and the average opacity value , static Gaussian parameters Constructing a static digital human model ; Step S1.3: Using 3D Gaussian Splatting rendering software, based on static Gaussian parameters , for static digital human models Rendering to generate synthetic static face images ; Step S1.4: Synthesize a static face image Compared with real static face images Input the spatial discriminator to obtain the discrimination result output by the spatial discriminator; Step S1.5: Based on the discrimination result and synthesized static face image Compared with real static face images The static loss function between , back-propagation optimization training of the static digital human model to obtain a trained static digital human model.

[0012] The static loss function described in step S1.5 Including fragment importance balance loss function , structural similarity D-SSIM loss function , Perceptual Similarity LPIPS Loss Function And the adversarial loss function ,in: Using fragment importance to balance the loss function Constrained Synthesis of Static Face Images Pixel-level realism, fragment importance balance loss function The calculation formula is as follows: ; in, Indicates the number of pixels, Represents a synthetic static face image Middle i The value of pixels, Represents a real face image Middle i The value of pixels; Using structural similarity D-SSIM loss function Constrained Synthesis of Static Face Images , structural similarity D-SSIM loss function The calculation formula is as follows: ; in, Represents the differentiable structural similarity D-SSIM function; Using perceptual similarity LPIPS loss function Constrained Synthesis of Static Face Images Perceptual realism, perceptual similarity LPIPS loss function The calculation formula is as follows: ; in, Indicates the number of network layers, Represents the first j The feature extraction operation of the layer, and Respectively represent j The height and width of the layer feature map, Represents the pixel position in the feature map; Using adversarial loss function Through the spatial discriminator Adversarial training to improve synthetic static face images Detailed realism, adversarial loss function The calculation formula is as follows: ; ; ; in, is a spatial discriminator For synthesizing static face images The judgment result of is a spatial discriminator For real static face images The judgment result of Is a real static face image Tags, It is a synthetic static face image The label of , BCE represents the binary cross entropy loss function, and MSE represents the mean square error loss function; Static loss function The calculation formula is as follows: ; in, , , and Represent the weight coefficients of each loss function respectively.

[0013] Step S2: Construct and train a speech-driven digital human synthesis system, which includes an optimizable global prompt module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; wherein the optimizable global prompt module generates global prompts, the progressive conditional attribute prediction network module takes audio signals, expression parameters, viewing angle parameters, and global prompts as input, predicts dynamic deformation in stages, combines dynamic deformation with a static digital human model, and obtains a dynamic digital human model; the dual discriminator architecture module discriminates between the dynamic digital human model and the real dynamic face image; the schematic diagram of the optimizable global prompt finally learned is shown in FIG. Figure 2 ; The specific steps of step S2 are as follows: Step S2.1: Convert the audio signal a , expression parameter e, viewing angle parameter v, and global prompts generated by the global prompt module Input progressive conditional attribute prediction network module to predict dynamic deformation in stages ;in is the global position offset, is the scale change, is the rotation adjustment amount, is the opacity value change, is the variation of spherical harmonic coefficients; The specific steps of step S2.1 are as follows: Step S2.1.1: Predict the global position offset of the dynamic digital human model , establish a spatial anchor point; Step S2.1.2: Offset at global position Based on the prediction of the scale change of the dynamic digital human model and rotation adjustment , refine the facial geometry; Step S2.1.3: Predict the opacity change of the dynamic digital human model based on the facial geometry , optimize the appearance of details; Step S2.1.4: Change the opacity Based on the above, the change of spherical harmonic coefficients of the dynamic digital human model is predicted. , capturing lighting and material details.

[0014] Step S2.2: Dynamic deformation Compared with the trained static digital human model Combined to obtain dynamic Gaussian parameters , dynamic Gaussian parameters Constructing a dynamic digital human model; Step S2.3: Using 3D Gaussian Splatting software, based on dynamic Gaussian parameters Rendering to generate synthetic dynamic face image sequences ; Step S2.4: Synthesize dynamic face image sequence and real dynamic face image sequences Input the dual discriminator architecture module respectively, the dual discriminator architecture module includes a spatial discriminator and time discriminator , obtain the discrimination results output by the spatial discriminator and the temporal discriminator; Spatial Discriminator The identification includes the following steps: Step S2.4.1.1: Extract features of different scales of the input image through a convolutional neural network to obtain multi-scale features ; Step S2.4.1.2: Multi-scale features Input to the multi-layer perceptron to obtain the spatial discriminator The output judgment result; It is the feature map obtained after the original resolution image passes through the convolutional neural network. It is the feature map obtained after the original image is downsampled by 1 / 2 and then passed through the convolutional neural network. It is the feature map obtained after the original image is downsampled by 1 / 4 and then passed through the convolutional neural network.

[0015] Temporal Discriminator The identification includes the following steps: Step S2.4.2.1: Perform 2D Fourier transform on the input synthetic dynamic face image sequence to extract frequency domain features , and then the original frame and its frequency domain features Perform splicing to obtain splicing features ; Step S2.4.2.2: Splicing features Three-dimensional convolution is applied to capture temporal correlation features, and the local attention module is used to purify the obtained temporal correlation features; so that it can focus on the key facial areas that are crucial to perceptual coherence to judge the temporal coherence of the animation sequence; Step S2.4.2.3: Generate a temporal quality score by combining two-dimensional convolution with a multi-layer perceptron.

[0016] Step S2.5: Based on the discrimination results and synthesized dynamic face image sequence Real dynamic face image sequence The dynamic loss function between , back-propagation optimization training of progressive conditional attribute prediction network module, optimizable global prompt module and dual discriminator architecture module to obtain a trained voice-driven digital human synthesis system.

[0017] The dynamic loss function described in step S2.5 Including static loss function , temporal adversarial loss function And the temporal consistency loss function ,in: Using static loss function Constrained Synthesis of Static Face Images The realism of a single frame image, static loss function The calculation formula is the same as the static loss function in step S1.5 The calculation formula is consistent; Using time series to fight against loss function Through the time discriminator Adversarial training to improve the synthesis of dynamic face image sequences Temporal coherence, temporal adversarial loss function The calculation formula is the same as the adversarial loss function in step S1.5 The calculation formula is consistent; Using the temporal consistency loss function Synthesis of dynamic face image sequences with explicit constraints The temporal consistency between frames further improves the smoothness and coherence of the animation. The temporal consistency loss function The calculation formula is as follows: ; in, T Indicates the number of frames in the image sequence, and Respectively represent the synthetic dynamic face image sequence Middle t Frame and t +1 frame image, and Represents real dynamic face image sequences Middle t Frame and t +1 frame image; Dynamic loss function The calculation formula is as follows: ; in, and They represent the weight coefficients of the temporal adversarial loss function and the temporal consistency loss function respectively.

[0018] Step S3: Input the speech signal into the trained speech-driven digital human synthesis system, output the speech-driven digital human animation, and complete the synthesis of the digital human animation.

[0019] The specific steps of step S3 are as follows: Step S3.1: Speech signal, expression parameter, viewing angle parameter and global prompt generated by the global prompt module can be optimized Input into the progressive conditional attribute prediction network module to predict dynamic deformation ; Step S3.2: Dynamic deformation Static Digital Human Model Combined to obtain dynamic Gaussian parameters ; Step S3.3: Using 3D Gaussian Splatting software, based on dynamic Gaussian parameters Generate speech-driven digital human animation in real time.

[0020] In order to verify the effectiveness of the method of the present invention, comparative experiments and ablation experiments are carried out as follows: First, the dataset and training details used are introduced. Then, the comparative experimental results of different algorithms on the dataset are presented. A series of ablation experiments are performed to evaluate the effectiveness of the optimizable global hint, the progressive conditional attribute prediction network module, and the dual discriminator architecture. The model training process is divided into two stages: static model initialization and dynamic model training. In the static model initialization stage, 8000 iterations of training were performed, and the batch size was set to 1. The weight coefficient of the loss function in this stage was set as follows: Coefficients of the fragment importance balance loss function Set to 0.8, the coefficient of the structural similarity D-SSIM loss function Set to 0.2, the coefficient of the perceptual similarity LPIPS loss function Set to 0.01, the coefficient of the adversarial loss function Set to 0.01.

[0021] The number of iterations in the dynamic model training phase is 20,000, the batch size is 16, and the weight coefficient of the loss function in this phase is set as follows: Coefficients of the temporal consistency loss function Set to 0.04, the coefficient of the time series adversarial loss function Set to 0.01.

[0022] The method proposed in this paper is compared with several current mainstream 3D reconstruction technologies. The experimental data set is from RAD-NeRF. In the comparative experiment, the better results in the experimental data and pre-trained models are selected, covering the most advanced models currently. The experimental results are shown in Table 1 and Table 2: Table 1. Comparative experiment (self-driving)

[0023] Table 2. Comparative experiment (cross-driver)

[0024] In order to further verify the effectiveness of each module in the model, an ablation experiment was conducted. In the experiment, the optimizable global hint, the progressive conditional attribute prediction network module and the dual discriminator architecture were removed respectively to evaluate the impact of these modules on the overall effect, and the model performance after different modules were removed was compared with the complete model. The ablation experiment results are shown in Tables 3, 4, 5 and 6. Indicates that the corresponding module is retained: Table 3. Ablation experiments for learnable global cues (self-driving)

[0025] Table 4. Ablation experiments for learnable global cues (cross-driver)

[0026] Table 5. Ablation experiment (self-driven)

[0027] Table 6. Ablation experiments (across drivers)

[0028] Among the evaluation indicators, PSNR and SSIM represent peak signal-to-noise ratio and structural similarity respectively, while LPIPS measures the perceptual similarity. PSNR mainly reflects the clarity of the synthesized image. The higher the value, the better the quality of the synthesized image. SSIM takes into account the structural similarity of the synthesized image. The higher the value, the higher the degree of structural preservation of the synthesized image. LPIPS is used to measure the perceptual difference of the synthesized image. The lower the value, the better the perceptual quality of the synthesized image. LMD is used to measure the deviation between the motion amplitude of the key points of the face of the synthesized animation and the real animation. The lower the value, the more accurate the motion amplitude of the synthesized animation. AUE is used to quantitatively evaluate the lip synchronization accuracy of the synthesized animation. The lower the value, the higher the lip synchronization accuracy. Sync-C is used to evaluate the audio and video synchronization confidence of the synthesized animation. The higher the value, the better the audio and video synchronization effect. FPS is used to measure the model inference speed and rendering efficiency. The higher the value, the higher the model efficiency.

[0029] As can be seen from Table 1, in the self-driven scenario setting, the proposed method has performances of 33.51, 0.944 and 0.038 in the three indicators of PSNR, SSIM and LPIPS, respectively, and especially outstanding performance in the frame rate (FPS) indicator, reaching 167FPS. These results show that the proposed method can achieve higher quality image rendering in the self-driven scenario, while maintaining significantly higher real-time rendering efficiency than other comparison methods. Compared with other comparison methods, the proposed method has achieved a better balance between image quality and efficiency, and can significantly improve the rendering speed while ensuring that the image quality is close to the optimal level, making it more suitable for real-time interactive application scenarios. In addition, in terms of motion quality and lip synchronization accuracy, the LMD index of the proposed method is 2.681, the AUE index is 1.062, and the Sync-C index is 6.205. All indicators are at the same level as advanced 3DGS-based methods such as TalkingGaussian and GaussianTalker, indicating that the proposed method does not sacrifice motion quality and lip synchronization accuracy while maintaining high rendering efficiency.

[0030] As can be seen from Table 2, in the cross-driven scenario setting, that is, when the model is driven by audio from unseen subjects, the proposed method still shows excellent performance and robustness. On the two test sets Testset A and Testset B, the three lip synchronization related indicators of the proposed method, Sync, LMD, and AUE, are significantly better than those of 2D-driven methods such as Wav2Lip and PC-AVS, and are at a competitive level with advanced 3D-driven methods such as RAD-NERF, ER-NeRF, TalkingGaussian, and GaussianTalker. These results show that the proposed method can still maintain stable and reliable lip synchronization performance in the cross-driven scenario, reflecting the good generalization ability and robustness of the proposed method when processing unseen audio input. It is particularly worth mentioning that compared with other 3D-driven methods, the proposed method still has a significant rendering efficiency advantage (as shown in Table 1) while maintaining the same or better lip synchronization performance, which makes the proposed method more advantageous and potential in practical applications.

[0031] From the data in Table 3 in the self-driving scenario, it can be seen that after the introduction of the optimizable global prompt module, the proposed method has achieved improvements in PSNR, SSIM and Sync-C indicators. PSNR is improved to 33.48, SSIM is improved to 0.944, and Sync-C is significantly improved to 6.141, indicating that the optimizable global prompt module can effectively improve the image quality and lip synchronization accuracy in the self-driving scenario.

[0032] From the data in Table 4 in the cross-driving scenario, it can be seen that after the introduction of the optimizable global prompt module, the method has achieved more significant improvements in the two lip synchronization related indicators LMD and Sync-C. The LMD indicator remains optimal, and Sync-C is significantly improved to 4.478, indicating that the optimizable global prompt module plays a more prominent role in improving lip synchronization accuracy and enhancing model robustness in the cross-driving scenario. From the data in Tables 3 and 4 in the self-driving and cross-driving scenarios, it can be seen that on the basis of introducing the optimizable global prompt module, after further introducing the progressive conditional attribute prediction network, the image quality index and lip synchronization accuracy index of the method in both scenarios have been further significantly improved, and the frame rate remains at a very competitive level.

[0033] From the data in Table 5 in the self-driving scenario, it can be seen that after the progressive conditional attribute prediction network, spatial discriminator and temporal discriminator modules are gradually introduced, the PSNR, SSIM and LPIPS indicators of this method all show a trend of continuous improvement, indicating that each module has made a positive contribution to improving image quality. From the data in Table 6 in the cross-driving scenario, it can be seen that after the progressive conditional attribute prediction network, spatial discriminator and temporal discriminator modules are gradually introduced, the LMD, AUE and Sync-C indicators of this method still show a trend of continuous improvement, indicating that each module has also made an important contribution to improving the lip synchronization performance in the cross-driving scenario. From the comprehensive full-module ablation experimental data in Tables 5 and 6 in the self-driving and cross-driving scenarios, it can be seen that the various modules proposed in the present invention, such as the optimizable global prompt module, the progressive conditional attribute prediction network and the dual discriminator architecture, have brought continuous and significant performance improvements to the proposed method in both scenarios, which effectively verifies the effectiveness and necessity of each module.

[0034] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A high-fidelity speech-driven digital human synthesis method based on 3DGS, characterized in that: Execute the following steps S1 to S3 to generate a digital human animation driven by a voice signal: Step S1: feature encoding and static Gaussian parameter prediction are performed on the digital human to construct a static digital human model. The basic shape and appearance of the digital human are captured and rendered using 3D Gaussian Splatting software. The static digital human model is trained using a back propagation method to obtain a trained static digital human model. Step S2: construct and train a voice-driven digital human synthesis system, including an optimizable global prompt module, a progressive conditional attribute prediction network module, and a dual discriminator architecture module; wherein the optimizable global prompt module generates global prompts, the progressive conditional attribute prediction network module takes audio signals, expression parameters, viewing angle parameters, and global prompts as inputs, predicts dynamic deformation in stages, combines the dynamic deformation with the static digital human model, and obtains a dynamic digital human model, and the dual discriminator architecture module discriminates between the dynamic digital human model and the real dynamic face image; Step S3: Input the speech signal into the trained speech-driven digital human synthesis system, output the speech-driven digital human animation, and complete the synthesis of the digital human animation.

2. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 1, characterized in that: The specific steps of step S1 are as follows: Step S1.1: Position the digital human in space Input multi-resolution triplanar Encode and obtain the feature vector , where the multi-resolution triplane consists of three orthogonal 2D feature grids The shape of each 2D feature grid is , H represents the feature hidden dimension, and R represents the dimension resolution; Step S1.2: Transform the feature vector Enter a static network , the static network Based on the multi-layer perceptron, the feature vector Mapping to static Gaussian parameters , including the average position , average scale , average rotation , average spherical harmonic coefficients and the average opacity value , static Gaussian parameters Constructing a static digital human model ; Step S1.3: Using 3D Gaussian Splatting rendering software, based on static Gaussian parameters , for static digital human models Rendering to generate synthetic static face images ; Step S1.4: Synthesize a static face image Compared with real static face images Input the spatial discriminator to obtain the discrimination result output by the spatial discriminator; Step S1.5: Based on the discrimination result and synthesized static face image Compared with real static face images The static loss function between , back-propagation optimization training of the static digital human model to obtain a trained static digital human model.

3. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 2 is characterized in that: The static loss function described in step S1.5 Including fragment importance balance loss function , structural similarity D-SSIM loss function , Perceptual Similarity LPIPS Loss Function And the adversarial loss function ,in: Fragment Importance Balanced Loss Function The calculation formula is as follows: ; in, Indicates the number of pixels, Represents a synthetic static face image Middle i The value of pixels, Represents a real face image Middle i The value of pixels; Structural Similarity D-SSIM Loss Function The calculation formula is as follows: ; in, Represents the differentiable structural similarity D-SSIM function; Perceptual Similarity LPIPS Loss Function The calculation formula is as follows: ; in, Indicates the number of network layers, Represents the first j The feature extraction operation of the layer, and Respectively represent j The height and width of the layer feature map, Represents the pixel position in the feature map; Adversarial Loss Function The calculation formula is as follows: ; ; ; in, is a spatial discriminator For synthesizing static face images The judgment result of is a spatial discriminator For real static face images The judgment result of Is a real static face image Tags, It is a synthetic static face image The label of , BCE represents the binary cross entropy loss function, and MSE represents the mean square error loss function; Static loss function The calculation formula is as follows: ; in, , , and Represent the weight coefficients of each loss function respectively.

4. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 3 is characterized in that: The specific steps of step S2 are as follows: Step S2.1: Convert the audio signal a , processed facial features e , visual feature v and optimizable global hints Input progressive conditional attribute prediction network module to predict dynamic deformation in stages ; Among them is Global position offset, is the scale change, is the rotation adjustment amount, is the opacity value change, is the variation of spherical harmonic coefficients; Step S2.2: Dynamic deformation Compared with the trained static digital human model Combined to obtain dynamic Gaussian parameters , dynamic Gaussian parameters Constructing a dynamic digital human model; Step S2.3: Using 3D Gaussian Splatting software, based on dynamic Gaussian parameters Rendering to generate synthetic dynamic face image sequences ; Step S2.4: Synthesize dynamic face image sequence and real dynamic face image sequences Input the dual discriminator architecture module respectively, the dual discriminator architecture module includes a spatial discriminator and time discriminator , obtain the discrimination results output by the spatial discriminator and the temporal discriminator; Step S2.5: Based on the discrimination results and synthesized dynamic face image sequence Real dynamic face image sequence The dynamic loss function between , back-propagation optimization training of progressive conditional attribute prediction network module, optimizable global prompt module and dual discriminator architecture module to obtain a trained voice-driven digital human synthesis system.

5. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 4 is characterized in that: The specific steps of step S2.1 are as follows: Step S2.1.1: Predict the global position offset of the dynamic digital human model ; Step S2.1.2: Offset at global position Based on the prediction of the scale change of the dynamic digital human model and rotation adjustment , refine the facial geometry; Step S2.1.3: Predict the opacity change of the dynamic digital human model based on the facial geometry ; Step S2.1.4: Change the opacity Based on the above, the change of spherical harmonic coefficients of the dynamic digital human model is predicted. .

6. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 4 is characterized in that: The spatial discriminator in step S2.4 The identification includes the following steps: Step S2.4.1.1: Extract features of different scales of the input image through a convolutional neural network to obtain multi-scale features ;in It is the feature map obtained after the original resolution image passes through the convolutional neural network. It is the feature map obtained after the original image is downsampled by 1 / 2 and then passed through the convolutional neural network. It is the feature map obtained after the original image is downsampled by 1 / 4 and then passed through the convolutional neural network; Step S2.4.1.2: Multi-scale features Input to the multi-layer perceptron to obtain the spatial discriminator Output the judgment result.

7. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 4 is characterized in that: The temporal discriminator in step S2.4 The identification includes the following steps: Step S2.4.2.1: Perform 2D Fourier transform on the input synthetic dynamic face image sequence to extract frequency domain features , and then the original frame and its frequency domain features Perform splicing to obtain splicing features ; Step S2.4.2.2: Splicing features Three-dimensional convolution is used to capture temporal correlation features, and the local attention module is used to purify the obtained temporal correlation features; Step S2.4.2.3: Generate the corresponding score by combining two-dimensional convolution with multi-layer perceptron.

8. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 4 is characterized in that: The dynamic loss function described in step S2.5 Including static loss function , temporal adversarial loss function And the temporal consistency loss function ,in: Static loss function The calculation formula is the same as the static loss function in step S1.5 The calculation formula is consistent; Temporal adversarial loss function The calculation formula is the same as the adversarial loss function in step S1.5 The calculation formula is consistent; Temporal consistency loss function The calculation formula is as follows: ; in, T Indicates the number of frames in the image sequence, and Represents the synthetic dynamic face image sequence Middle t Frame and t +1 frame image, and Represents real dynamic face image sequences Middle t Frame and t +1 frame image; Dynamic loss function The calculation formula is as follows: ; in, and They represent the weight coefficients of the temporal adversarial loss function and the temporal consistency loss function respectively.

9. The high-fidelity speech-driven digital human synthesis method based on 3DGS according to claim 4 is characterized in that: The specific steps of step S3 are as follows: Step S3.1: Speech signal, expression parameter, viewing angle parameter and global prompt generated by the global prompt module can be optimized Input into the progressive conditional attribute prediction network module to predict dynamic deformation ; Step S3.2: Dynamic deformation Static Digital Human Model Combined to obtain dynamic Gaussian parameters ; Step S3.3: Using 3D Gaussian Splatting software, based on dynamic Gaussian parameters Generate speech-driven digital human animation in real time.

Citation Information

Patent Citations

  • Real-time high-fidelity voice-driven digital human system

    CN117746840A

  • Digital human head portrait generation method based on real-time audio driving

    CN119006663A

  • Digital human modeling method and device, equipment, storage medium and program product

    CN119295682A

  • High-definition digital human video generation method and system for customizing real human image, storage medium and equipment

    CN119729130A

  • Three dimensional gaussian splatting initialization based on trained neural radiance field representations

    US20240355047A1

Cited By

  • Large-attitude face animation synthesis method and system based on three-plane features

    CN121392078A