Real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning
Through the improved GE2E network and 3D digital holographic cabin technology, separable control of tone, emotions and rhythm is achieved, which solves the shortcomings of voice style and action driving in speech synthesis, improves the synchronization and immersion of the virtual human system, and is suitable for scenarios such as virtual customer service, virtual speech, etc.
Patent Information
- Application Number
- CN202510803896.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The prior art lacks the ability to control the speech style, speech adjustment rhythm and emotional changes of specific characters in speech synthesis, resulting in insufficient expressiveness, interactivity and immersion in synthesized speech, and the digital human action driver lacks the linkage mechanism of voice content and semantic rhythm, affecting the user's immersive experience.
Through the improved GE2E network, the multi-component embedding mechanism is introduced to realize the separable control of tone, emotion and rhythm. Combined with semantic-driven three-dimensional action modeling, a high-precision alignment mechanism between speech frames and digital human action frames is established, and the synchronous rendering of speech images and spatial visual output is realized through the 3D digital holographic cabin.
It realizes a more natural, more synchronous and immersive human-computer interaction experience, improves the personalization and expressiveness of speech generation, solves the problem of inconsistent timing between speech and actions, and provides real-time multimodal interaction capabilities.
Smart Images

Figure CN120318437B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction, and in particular to a real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning. Background Art
[0002] With the rapid development of artificial intelligence, deep learning and human-computer interaction technology, cutting-edge fields such as three-dimensional digital humans, speech synthesis, virtual reality and holographic imaging are constantly converging and are widely used in various digital scenarios such as virtual customer service, virtual speeches, remote meetings, education and training, cultural and tourism exhibitions, and metaverse social networking. Among them, the human-computer interaction system with three-dimensional digital humans as the core has become a research hotspot. Its goal is to build a "virtual agent" that is as close to real people as possible in terms of vision, hearing and interactive experience, so as to achieve a more natural, immersive and intelligent remote communication experience.
[0003] Among existing speech synthesis technologies, mainstream methods mainly adopt model architectures based on deep neural networks, such as the Tacotron series, FastSpeech series, VITS, and Glow-TTS. These models can generally convert input text into high-quality speech waveforms, with the advantages of natural, clear, and fluent speech. However, most traditional TTS systems tend to use general speech models and can only generate speech with neutral intonation and standard timbre. They lack the ability to control the voice style, intonation rhythm, and emotional changes of specific characters. Therefore, they cannot meet the needs of personalized speech expression or the reproduction of specific character voices.
[0004] In order to improve the personalization and authenticity of speech synthesis, researchers have gradually introduced "voice cloning" into the speech modeling process. Voice cloning technology learns the timbre characteristics of a specific speaker and can achieve high-fidelity speech replication with only a small number of voice samples without the need for a large amount of target person's voice data. The existing voice cloning models can extract the speaker's timbre embedding vector for speech synthesis. However, existing models usually only focus on the static embedding expression of the speaker's timbre and lack fine-grained control over the speaker's emotions and intonation, resulting in significant deficiencies in the generated speech in terms of expressiveness, interactivity and immersion, especially in application scenarios that require driving three-dimensional digital humans or virtual images.
[0005] On the other hand, three-dimensional digital human systems can currently achieve relatively natural movement generation through technologies such as facial expression capture and skeletal motion modeling. However, the linkage mechanism between speech and movement is still imperfect. Existing solutions mostly rely on preset rules or static mapping and lack the ability to control dynamic movements based on semantics, emotions and rhythm. This leads to significant temporal inconsistencies and emotional disconnection between synthesized speech and digital human expressions and postures, affecting the user's immersive experience.
[0006] Overall, the existing technology has obvious deficiencies in the following aspects: First, sound cloning models are mostly limited to timbre embedding modeling, lacking multi-component expression and joint optimization mechanisms for intonation rhythm and emotional changes, which limits the expressiveness of synthesized speech in emotional interaction; Second, digital human action drive still relies on single-modal input, lacks a multi-dimensional linkage mechanism from voice content, intonation and semantic rhythm, and it is difficult to achieve highly synchronized, natural and emotionally consistent voice-action joint output; Third, in terms of terminal presentation, there is a lack of 3D holographic cabin platforms that are oriented to interactive needs and have multi-modal real-time rendering capabilities, which cannot meet the immersive experience requirements of "speaking is performance" and "communication is response" in human-computer interaction scenarios. Summary of the Invention
[0007] One purpose of the present invention is to propose a real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning. The present invention realizes the separable control of timbre, emotion and rhythm by improving the GE2E network and introducing a multi-component embedding mechanism. Through the acoustic synthesis process and vector injection strategy of FastSpeech2 and HiFi-GAN, it ensures that the speech output has high fidelity and rich emotional expression. Combined with the semantic-driven three-dimensional action modeling mechanism, a high-precision alignment mechanism between speech frames and digital human action frames is established. Through the 3D digital holographic cabin, the synchronous rendering of speech images and spatial visualization output are realized, thereby improving the authenticity, synchronization and interactivity of the virtual human system.
[0008] According to an embodiment of the present invention, a real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning includes the following steps:
[0009] S1. Collect the user's facial image, body movement image and original voice signal and perform preprocessing;
[0010] S2, extracting features from the preprocessed facial image and body movement image to generate facial expression feature vectors and body movement feature vectors;
[0011] S3, modeling the original speech signal, generating a speech timbre feature vector using an improved GE2E network, encoding a preset target speech text, generating a semantic feature vector, and splicing it with the speech timbre feature vector to form speech synthesis data;
[0012] S4, inputting the speech synthesis data into a speech synthesis model to generate synthesized speech audio;
[0013] S5, mapping the facial expression feature vector to a three-dimensional facial muscle control parameter vector, and mapping the limb movement feature vector to a three-dimensional skeleton movement control parameter vector, to generate a three-dimensional digital human movement sequence;
[0014] S6. Align the timestamps of the 3D digital human action sequence and the synthesized speech audio to construct a synchronous output stream of speech drive and action control;
[0015] S7. Input the synchronized output stream into the 3D digital holographic cabin for rendering, generate three-dimensional digital human image frames synchronized with the voice frames in real time, and output stereoscopic visualization through the holographic cabin projection device to present virtual human interaction with synchronized voice and action.
[0016] Optionally, the preprocessing includes denoising, image enhancement, Fourier transform, and normalization and standardization.
[0017] Optionally, feature extraction of facial image sequences and body motion image sequences adopts a deep convolutional neural network and a human posture estimation model respectively.
[0018] Optionally, the S3 specifically includes:
[0019] S31. Perform fundamental frequency analysis on the original speech signal, extract the frame-level fundamental frequency sequence, calculate the fundamental frequency change rate between adjacent frames, and adjust the frame length and frame shift of each frame based on the fundamental frequency change rate:
[0020] ;
[0021] in, represents the frame length of the tth frame of the original speech signal, represents the frame shift of the tth frame of the original speech signal, and Represent the basic frame length and basic frame shift respectively, 、 represents the adjustment coefficient, Indicates the absolute value of the fundamental frequency change rate of the t-th frame;
[0022] S32, performing sliding window framing and window function weighting on the original speech signal based on the frame length and frame shift of each frame, performing short-time Fourier transform and then processing with a Mel filter to obtain a Mel spectrogram sequence;
[0023] S33, inputting the mel-spectrogram sequence into an improved GE2E network, wherein the improved GE2E network includes a frame-level feature encoding module, a statistical pooling module, and a multi-branch embedding module, wherein the frame-level feature encoding module includes a multi-layer fully connected network and a bidirectional gated recurrent unit, and outputs a frame embedding vector;
[0024] S34. Perform statistical pooling on the frame embedding vector through the statistical pooling module to calculate the mean vector and standard deviation vector:
[0025] ;
[0026] in, represents the mean vector of all frame embedding vectors, represents the standard deviation vector of all frame embedding vectors, represents the frame embedding vector of the t-th frame, and T represents the total number of frames;
[0027] Concatenate the mean vector and the standard deviation vector to generate a speech timbre feature vector;
[0028] S35, deconstructing the speech timbre feature vector into three vectors through the multi-branch embedding module:
[0029] ;
[0030] in, represents the speech timbre feature vector, represents the timbre component vector, represents the emotion component vector, represents the intonation rhythm component vector, Represents a splicing operation;
[0031] S36. Construct a timbre center vector based on the timbre component vector, and use a joint loss function to optimize the parameters of the GE2E network to generate an optimized speech timbre feature vector:
[0032] ;
[0033] ;
[0034] in, represents the timbre center vector of the jth speaker, M represents the number of speech samples corresponding to the speaker, represents the timbre component vector of the i-th speech sample, L represents the joint loss function, represents the main loss, and represents the loss weight of the subtask, and represent emotion-assisted loss term and intonation-assisted loss term respectively;
[0035] The main loss is the GE2E network loss:
[0036] ;
[0037] in, represents the main loss, represents the cosine similarity between the i-th speech sample and the k-th timbre center vector, represents the cosine similarity function, represents the timbre center vector of the kth speaker, w and b represent the scaling and offset parameters, M represents the number of speech samples corresponding to the speaker, represents the timbre component vector of the i-th speech sample, e represents a natural constant, represents the cosine similarity between the i-th speech sample and the j-th timbre center vector;
[0038] S37. Encode the preset target speech text through a text encoder to generate a semantic feature vector, and concatenate it with the optimized speech timbre feature vector to form speech synthesis data.
[0039] Optionally, the emotion auxiliary loss term is constructed based on a cross entropy loss function, and the prediction error between the timbre component vector and the true emotion label is used to optimize the embedding expression capability of the GE2E network.
[0040] Optionally, the intonation auxiliary loss term adopts a mean square error regression function to minimize the distance between the intonation rhythm component vector and the target intonation vector, so that the GE2E network learns the expressive ability of speaking rhythm, stress and pauses, and realizes rhythm synchronization control in multimodal driving.
[0041] Optionally, the S4 specifically includes:
[0042] S41, inputting the speech synthesis data into the acoustic module of the speech synthesis model;
[0043] S42. Encode the speech synthesis data using a text encoder to generate a phoneme-level semantic vector, input the phoneme-level semantic vector into a duration predictor to predict a frame duration for each phoneme, and expand each phoneme according to the frame duration to generate a phoneme sequence;
[0044] S43. Introduce the intonation rhythm component vector to perform frame-level proportional control on the frame duration:
[0045] ;
[0046] in, represents the frame duration after regulation, represents the rhythm adjustment factor, tanh represents the hyperbolic tangent function, represents the intonation rhythm component vector, r represents the intonation rhythm direction vector, P represents the transposition operation, Indicates the frame duration;
[0047] S44. Further process the phoneme sequence, introduce the emotion component vector as a control factor, and regulate the pitch predictor and energy predictor respectively to obtain the fundamental frequency vector and energy vector after emotion modulation:
[0048] ;
[0049] in, represents the fundamental frequency vector after emotion modulation of the t-th frame, represents the energy vector after modulation of the t-th frame, and denote the fundamental frequency vector of the pitch predictor and the energy vector of the energy predictor, respectively, and Represents the adjustment proportional coefficient, and represents the weight vector;
[0050] S45. Input the emotion-modulated fundamental frequency vector and energy vector and the phoneme sequence into the decoder, output the acoustic feature sequence, and send it to the vocoder module to reconstruct the speech waveform to obtain the synthesized speech audio.
[0051] Optionally, the speech synthesis model includes an acoustic module and a vocoder module.
[0052] Optionally, the acoustic module takes speech synthesis data as input and outputs an acoustic feature sequence, and the vocoder module converts the acoustic feature sequence into a time-domain speech signal.
[0053] Optionally, the acoustic module adopts a FastSpeech2 network, including a text encoder, a duration predictor, a pitch predictor, an energy predictor and a decoder, and the vocoder module adopts a HiFi-GAN network.
[0054] The beneficial effects of the present invention are:
[0055] First of all, the present invention proposes a real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning, which can effectively overcome the technical bottlenecks of existing technologies in voice expressiveness, voice and action linkage, and real-time spatial presentation of virtual humans, and realize a more natural, more synchronous and more immersive human-computer interaction experience. By introducing an improved GE2E network structure, the voice timbre feature vector is deconstructed into timbre component, emotion component and intonation rhythm component in the voice modeling stage, so that the voice synthesis not only has the timbre characteristics of the target person, but also can express the emotional state and language rhythm in a specific context, thereby improving the personalization and expressiveness of voice generation. In addition, the present invention optimizes the voice embedding structure through a joint loss function, so that the system has the ability to learn the emotion and intonation dimensions while learning the timbre characteristics, realizing a deep expansion of the semantic control of the voice cloning model, and enhancing the interactive responsiveness and pragmatic fit of speech synthesis.
[0056] Secondly, in terms of digital human motion driving, the present invention maps facial expression features and limb movement features into three-dimensional muscle control parameters and bone control parameters respectively, and constructs a joint motion control sequence of the face and limbs, so that the expression, posture and language content of the three-dimensional virtual human are naturally synchronized. At the same time, combined with the timestamp alignment mechanism of the synthesized voice frame and the action frame, a synchronous output stream is constructed, which effectively solves the problem of timing misalignment between voice and visual performance, and ensures the authenticity and coherence when users watch the digital human communicating.
[0057] Finally, the present invention inputs the synchronized multimodal output data into the three-dimensional holographic cabin system, converts the motion control frames into spatial image frames through a real-time rendering module, and realizes holographic stereo output in combination with audio signals, so that users can obtain an immersive voice-visual linkage experience without the need for head-mounted devices. While ensuring low-latency output, the system can also support users to provide real-time feedback through voice or action, forming a continuous interactive closed loop. This mechanism significantly improves the instant responsiveness, degree of anthropomorphism and immersion of human-computer interaction, and provides a complete and efficient solution for building highly realistic virtual humans, metaverse interactive portals or remote visual communication platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0059] Figure 1 This is a flowchart of a real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning proposed by the present invention;
[0060] Figure 2 This is a schematic diagram of the improved GE2E network structure of the real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning proposed in the present invention. DETAILED DESCRIPTION
[0061] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0062] refer to Figure 1-2 A real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning includes the following steps:
[0063] S1. Collect the user's facial image, body movement image and original voice signal and perform preprocessing;
[0064] S2, extracting features from the preprocessed facial image and body movement image to generate facial expression feature vectors and body movement feature vectors;
[0065] S3, modeling the original speech signal, generating a speech timbre feature vector using an improved GE2E network, encoding a preset target speech text, generating a semantic feature vector, and splicing it with the speech timbre feature vector to form speech synthesis data;
[0066] S4, inputting the speech synthesis data into a speech synthesis model to generate synthesized speech audio;
[0067] S5, mapping the facial expression feature vector to a three-dimensional facial muscle control parameter vector, and mapping the limb movement feature vector to a three-dimensional skeleton movement control parameter vector, to generate a three-dimensional digital human movement sequence;
[0068] S6. Align the timestamps of the 3D digital human action sequence and the synthesized speech audio to construct a synchronous output stream of speech drive and action control;
[0069] S7. Input the synchronized output stream into the 3D digital holographic cabin for rendering, generate three-dimensional digital human image frames synchronized with the voice frames in real time, and output stereoscopic visualization through the holographic cabin projection device to present virtual human interaction with synchronized voice and action.
[0070] By integrating deep learning and sound cloning technology, this invention constructs a three-dimensional digital human voice and motion collaborative synthesis method for multimodal data input, achieving high-precision real-time synchronous linkage of virtual human voice, expression and body movements, and improving the immersion and realism of digital human interaction.
[0071] In this embodiment, the preprocessing includes denoising, image enhancement, Fourier transform, and normalization and standardization. Specifically, image data is subjected to median filtering and bilateral filtering to remove noise, and histogram equalization and gamma correction are combined to enhance image contrast and detail. A two-dimensional Fourier transform is applied to extract frequency domain texture features. Speech signals are denoised using spectral subtraction, and time-frequency spectrum features are extracted using a short-time Fourier transform. All modal data are finally normalized and standardized to unify feature scales and distributions.
[0072] The present invention introduces image enhancement, Fourier transform and normalization operations in the preprocessing stage, which significantly improves the signal-to-noise ratio and feature extraction quality of multimodal input data, and provides higher robustness and generalization capability for subsequent deep network modeling.
[0073] In this embodiment, the feature extraction of facial image sequences and body movement image sequences adopts a deep convolutional neural network and a human posture estimation model respectively.
[0074] The present invention uses a deep convolutional neural network and a human posture estimation model to perform specific feature extraction on image sequences, achieving high-precision semantic encoding of facial expressions and body movements, and effectively enhancing the delicacy and naturalness of the three-dimensional virtual human movement generation.
[0075] In this embodiment, S3 specifically includes:
[0076] S31. Perform fundamental frequency analysis on the original speech signal, extract the frame-level fundamental frequency sequence, calculate the fundamental frequency change rate between adjacent frames, and adjust the frame length and frame shift of each frame based on the fundamental frequency change rate:
[0077] ;
[0078] in, represents the frame length of the tth frame of the original speech signal, represents the frame shift of the tth frame of the original speech signal, and Represent the basic frame length and basic frame shift respectively, 、 represents the adjustment coefficient, Indicates the absolute value of the fundamental frequency change rate of the t-th frame;
[0079] S32, performing sliding window framing and window function weighting on the original speech signal based on the frame length and frame shift of each frame, performing short-time Fourier transform and then processing with a Mel filter to obtain a Mel spectrogram sequence;
[0080] S33, inputting the mel-spectrogram sequence into an improved GE2E network, wherein the improved GE2E network includes a frame-level feature encoding module, a statistical pooling module, and a multi-branch embedding module, wherein the frame-level feature encoding module includes a multi-layer fully connected network and a bidirectional gated recurrent unit, and outputs a frame embedding vector;
[0081] S34. Perform statistical pooling on the frame embedding vector through the statistical pooling module to calculate the mean vector and standard deviation vector:
[0082] ;
[0083] in, represents the mean vector of all frame embedding vectors, represents the standard deviation vector of all frame embedding vectors, represents the frame embedding vector of the t-th frame, and T represents the total number of frames;
[0084] Concatenate the mean vector and the standard deviation vector to generate a speech timbre feature vector;
[0085] S35, deconstructing the speech timbre feature vector into three vectors through the multi-branch embedding module:
[0086] ;
[0087] in, represents the speech timbre feature vector, represents the timbre component vector, represents the emotion component vector, represents the intonation rhythm component vector, Represents a splicing operation;
[0088] S36. Construct a timbre center vector based on the timbre component vector, and use a joint loss function to optimize the parameters of the GE2E network to generate an optimized speech timbre feature vector:
[0089] ;
[0090] ;
[0091] in, represents the timbre center vector of the jth speaker, M represents the number of speech samples corresponding to the speaker, represents the timbre component vector of the i-th speech sample, L represents the joint loss function, represents the main loss, and represents the loss weight of the subtask, and represent emotion-assisted loss term and intonation-assisted loss term respectively;
[0092] The main loss is the GE2E network loss:
[0093] ;
[0094] in, represents the main loss, represents the cosine similarity between the i-th speech sample and the k-th timbre center vector, represents the cosine similarity function, represents the timbre center vector of the kth speaker, w and b represent the scaling and offset parameters, M represents the number of speech samples corresponding to the speaker, represents the timbre component vector of the i-th speech sample, e represents a natural constant, represents the cosine similarity between the i-th speech sample and the j-th timbre center vector;
[0095] S37. Encode the preset target speech text through a text encoder to generate a semantic feature vector, and concatenate it with the optimized speech timbre feature vector to form speech synthesis data.
[0096] This paper systematically improves the modeling process of the original speech signal, adopts a dynamic frame division mechanism and a structurally optimized GE2E network, improves the resolution of speech features and the accuracy of timbre extraction, and provides high-quality input for subsequent sound cloning and speech synthesis.
[0097] In this embodiment, the emotion auxiliary loss term is constructed based on the cross entropy loss function, and the prediction error between the timbre component vector and the true emotion label is used to optimize the embedding expression ability of the GE2E network.
[0098] In this embodiment, the intonation auxiliary loss term adopts the mean square error regression function, which enables the GE2E network to learn the expressive ability of speaking rhythm, stress and pauses by minimizing the distance between the intonation rhythm component vector and the target intonation vector, and realizes rhythm synchronization control in multimodal driving.
[0099] The present invention introduces emotion and intonation auxiliary loss functions into the improved GE2E network, which enhances the network's ability to model multi-component speech information (timbre, emotion, rhythm), making the generated speech richer and more natural in pragmatic expression and emotional expression.
[0100] In this embodiment, the S4 specifically includes:
[0101] S41, inputting the speech synthesis data into the acoustic module of the speech synthesis model;
[0102] S42. Encode the speech synthesis data using a text encoder to generate a phoneme-level semantic vector, input the phoneme-level semantic vector into a duration predictor to predict a frame duration for each phoneme, and expand each phoneme according to the frame duration to generate a phoneme sequence;
[0103] S43. Introduce the intonation rhythm component vector to perform frame-level proportional control on the frame duration:
[0104] ;
[0105] in, represents the frame duration after regulation, represents the rhythm adjustment factor, tanh represents the hyperbolic tangent function, represents the intonation rhythm component vector, r represents the intonation rhythm direction vector, P represents the transposition operation, Indicates the frame duration;
[0106] S44. Further process the phoneme sequence, introduce the emotion component vector as a control factor, and regulate the pitch predictor and energy predictor respectively to obtain the fundamental frequency vector and energy vector after emotion modulation:
[0107] ;
[0108] in, represents the fundamental frequency vector after emotion modulation of the t-th frame, represents the energy vector after modulation of the t-th frame, and denote the fundamental frequency vector of the pitch predictor and the energy vector of the energy predictor, respectively, and Represents the adjustment proportional coefficient, and represents the weight vector;
[0109] S45. Input the emotion-modulated fundamental frequency vector and energy vector and the phoneme sequence into the decoder, output the acoustic feature sequence, and send it to the vocoder module to reconstruct the speech waveform to obtain the synthesized speech audio.
[0110] The present invention enhances the speech synthesis model's ability to control speaking rhythm and emotional expression by applying the intonation rhythm component vector to the phoneme frame duration prediction and injecting the emotion component vector into the pitch and energy prediction process, making the generated speech more natural and expressive.
[0111] In this embodiment, the speech synthesis model includes an acoustic module and a vocoder module.
[0112] In this embodiment, the acoustic module takes speech synthesis data as input and outputs an acoustic feature sequence, and the vocoder module converts the acoustic feature sequence into a time-domain speech signal.
[0113] In this embodiment, the acoustic module adopts the FastSpeech2 network, including a text encoder, a duration predictor, a pitch predictor, an energy predictor and a decoder, and the vocoder module adopts the HiFi-GAN network.
[0114] Example 1:
[0115] In order to verify the feasibility of the present invention in implementation, the present invention is applied to a remote virtual conference interaction system, and a virtual human interaction experience based on voice cloning and expression synchronization control is realized through a 3D digital holographic cabin. The voice-action linkage effect, speech synthesis expressiveness, and synchronization stability of multimodal output are further tested.
[0116] The test scenario was set up in a 5m×5m×3m three-dimensional holographic cabin environment at an artificial intelligence industry research institute in Shanghai. The cabin was equipped with a 360-degree holographic projection system with a resolution of 3840×2160 and a frame rate of 60fps. It was also equipped with multi-angle cameras for collecting user facial and body images, and a high-fidelity microphone array for voice collection. The system deployed a server-side based on NVIDIA graphics cards to execute the complete process described in this invention: from voice and image acquisition, feature extraction, sound cloning, speech synthesis, to digital human rendering and holographic output.
[0117] The test subjects were three participants of different ages, both male and female, with significant individual differences in their speech expression styles. During the preprocessing phase, the system denoised and Fourier transformed the raw speech, normalized and enhanced the facial and motion images, and then applied Mel-spectrogram analysis to the improved GE2E network, generating a three-component embedding vector encompassing timbre, emotion, and rhythm. During the speech synthesis phase, the FastSpeech2 architecture receives the speech control vector and combines it with the text to generate a sequence of phoneme vectors. This is then expanded to the frame level using a duration predictor. The system further manipulates the acoustic features using the modulated fundamental frequency and energy vector, and the final speech waveform is generated using the HiFi-GAN vocoder. The motion control component, based on the extracted facial expressions and skeletal features, drives the 3D digital human model to generate continuous motion frames, which are output synchronously with the speech audio.
[0118] To verify the improvement effects of the present invention in the three key dimensions of "sound-action synchronization", "speech expression authenticity" and "interaction delay", the following experiment was designed: the method of the present invention was compared with the traditional TTS+animation synthesis method, and 30 rounds of interaction tasks were carried out in the same scenario. Each round included three parts: user questions, system responses, and digital human feedback expressions and actions. The content covered four typical virtual interaction commands: greetings, emotional expressions, time queries, and scene simulations.
[0119] Table 1 Comparative test results of the present invention and the traditional system in digital human interaction tasks
[0120] ;
[0121] In terms of speech synthesis response, the method of the present invention is based on the neural speech synthesis model constructed by FastSpeech2 and HiFi-GAN. While introducing a triple embedding control mechanism of timbre, emotion and intonation, it can still control the average speech generation time to 482.8 milliseconds. Compared with the 608 milliseconds of the traditional TTS method, the overall delay is reduced by about 20%, providing sufficient processing redundancy for real-time interaction. In the digital human action performance part, the present invention effectively controls the action-driven delay by mapping facial expression features and skeletal action features into three-dimensional control parameters respectively, and combining the alignment mechanism of speech frames and action frames. The average delay is only 110 milliseconds, which is significantly improved compared with the 264 milliseconds of the traditional system, effectively avoiding the sense of "asynchronous" speech-action disconnection.
[0122] In terms of user subjective perception, all three key scoring indicators significantly outperformed traditional methods. The naturalness score was 4.72, significantly higher than the 3.90 of the traditional system, indicating that the improved GE2E network can more accurately reconstruct personalized timbre and maintain high fidelity and coherence during speech synthesis. In terms of emotional expression matching, the proposed method scored 4.64, a significant improvement over the 3.20 of the traditional method. This is due to the independent modeling of emotional components during the embedding process and optimization with auxiliary losses, making the synthesized speech more expressive. The coordination of expression and speech scored 4.82, far higher than the 3.10 of the traditional method, reflecting the precision and model linkage advantages of the proposed method in speech-driven expression synchronization control.
[0123] In addition, in terms of synchronization error control between voice frames and action frames, the average error of the present invention is ±36 milliseconds, which is nearly three times lower than the ±112 milliseconds of the traditional method. From a technical perspective, it ensures the alignment consistency of hearing and vision at the user perception layer. At the same time, in the holographic projection frame rate stability test, the present invention can continuously output 60FPS stable image frames, while the traditional method only maintains 47FPS under multimodal scheduling, indicating that the rendering architecture of the present invention is more suitable for real-time interactive scenarios and avoids screen jitter and frame drops.
[0124] The total interactive response time measured after integrating the three links of speech generation, action linkage, and image rendering is 673 milliseconds, which is significantly better than the 1014 milliseconds of the traditional method. The overall interaction delay is reduced by more than 34%. This means that the present invention not only has significant advantages in speech synthesis quality and expression, but also demonstrates efficient response and strong interactive capabilities in engineering implementation. It is suitable for real-time human-computer interaction scenarios such as remote virtual meetings, virtual explanations, and intelligent customer service that require a high degree of consistency in "speaking-acting-seeing".
[0125] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning, characterized in that: The steps include: S1. Collect the user's facial image, body movement image and original voice signal and perform preprocessing; S2, extracting features from the preprocessed facial image and body movement image to generate facial expression feature vectors and body movement feature vectors; S3, modeling the original speech signal, generating a speech timbre feature vector using an improved GE2E network, encoding a preset target speech text, generating a semantic feature vector, and splicing it with the speech timbre feature vector to form speech synthesis data; The improved GE2E network includes a frame-level feature encoding module, a statistical pooling module and a multi-branch embedding module. The frame-level feature encoding module includes a multi-layer fully connected network and a bidirectional gated recurrent unit; S4, inputting the speech synthesis data into a speech synthesis model to generate synthesized speech audio; S5, mapping the facial expression feature vector to a three-dimensional facial muscle control parameter vector, and mapping the limb movement feature vector to a three-dimensional skeleton movement control parameter vector, to generate a three-dimensional digital human movement sequence; S6. Align the timestamps of the 3D digital human action sequence and the synthesized speech audio to construct a synchronous output stream of speech drive and action control; S7, inputting the synchronized output stream into the 3D digital holographic cabin for rendering, generating three-dimensional digital human image frames synchronized with the voice frames in real time, and outputting the three-dimensional visualization through the holographic cabin projection device, presenting virtual human interaction with synchronized voice and action; The S4 specifically includes: S41, inputting the speech synthesis data into the acoustic module of the speech synthesis model; S42. Encode the speech synthesis data using a text encoder to generate a phoneme-level semantic vector, input the phoneme-level semantic vector into a duration predictor to predict a frame duration for each phoneme, and expand each phoneme according to the frame duration to generate a phoneme sequence; S43. Introduce the intonation rhythm component vector to perform frame-level proportional control on the frame duration: ; in, represents the frame duration after regulation, represents the rhythm regulation factor, represents the hyperbolic tangent function, represents the intonation rhythm component vector, represents the intonation rhythm direction vector, represents the transpose operation, Indicates the frame duration; S44. Further process the phoneme sequence, introduce the emotion component vector as a control factor, and regulate the pitch predictor and energy predictor respectively to obtain the fundamental frequency vector and energy vector after emotion modulation: ; in, Indicates the The fundamental frequency vector after frame emotion modulation, Indicates the The energy vector after frame modulation, and denote the fundamental frequency vector of the pitch predictor and the energy vector of the energy predictor, respectively, and Represents the adjustment proportional coefficient, and represents the weight vector, represents the emotion component vector; S45. Input the emotion-modulated fundamental frequency vector and energy vector and the phoneme sequence into the decoder, output the acoustic feature sequence, and send it to the vocoder module to reconstruct the speech waveform to obtain the synthesized speech audio.
2. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 1, characterized in that: The preprocessing includes denoising, image enhancement, Fourier transform, and normalization and standardization.
3. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 1, characterized in that: The feature extraction of facial image sequences and body movement image sequences uses deep convolutional neural networks and human posture estimation models respectively.
4. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 1, characterized in that: The S3 specifically includes: S31. Perform fundamental frequency analysis on the original speech signal, extract the frame-level fundamental frequency sequence, calculate the fundamental frequency change rate between adjacent frames, and adjust the frame length and frame shift of each frame based on the fundamental frequency change rate: ; in, Represents the original speech signal The frame length of the frame, Represents the original speech signal Frame shift of frames, and Represent the basic frame length and basic frame shift respectively, 、 represents the adjustment coefficient, Indicates the The absolute value of the frame's fundamental frequency change rate; S32, performing sliding window framing and window function weighting on the original speech signal based on the frame length and frame shift of each frame, performing short-time Fourier transform and then processing with a Mel filter to obtain a Mel spectrogram sequence; S33, inputting the mel-spectrogram sequence into an improved GE2E network, wherein the improved GE2E network includes a frame-level feature encoding module, a statistical pooling module, and a multi-branch embedding module, wherein the frame-level feature encoding module includes a multi-layer fully connected network and a bidirectional gated recurrent unit, and outputs a frame embedding vector; S34, performing statistical pooling on the frame embedding vector through a statistical pooling module to calculate the mean vector and the standard deviation vector; Concatenate the mean vector and the standard deviation vector to generate a speech timbre feature vector; S35, deconstructing the speech timbre feature vector into three vectors through the multi-branch embedding module: ; in, represents the speech timbre feature vector, represents the timbre component vector, represents the emotion component vector, represents the intonation rhythm component vector, Represents a splicing operation; S36. Construct a timbre center vector based on the timbre component vector, and use a joint loss function to optimize the parameters of the GE2E network to generate an optimized speech timbre feature vector: ; in, represents the joint loss function, represents the main loss, and represents the loss weight of the subtask, and represent emotion-assisted loss term and intonation-assisted loss term respectively; The main loss is the GE2E network loss: ; in, represents the main loss, Indicates the The speech samples and The cosine similarity of the timbre center vectors, represents the cosine similarity function, Indicates the The timbre center vector of each speaker, and represents the scaling and offset parameters, Indicates the number of speech samples corresponding to the speaker, Indicates the The timbre component vector of the speech samples, represents a natural constant, Indicates the The speech samples and The cosine similarity of the timbre center vectors; S37. Encode the preset target speech text through a text encoder to generate a semantic feature vector, and concatenate it with the optimized speech timbre feature vector to form speech synthesis data.
5. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 4, characterized in that: The emotion auxiliary loss term is constructed based on the cross entropy loss function, and the prediction error between the timbre component vector and the true emotion label is used to optimize the embedding expression ability of the GE2E network.
6. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 4, characterized in that: The intonation-assisted loss term adopts a mean square error regression function. By minimizing the distance between the intonation rhythm component vector and the target intonation vector, the GE2E network learns the expressiveness of speech rhythm, stress and pauses, and realizes rhythm synchronization control in multimodal driving.
7. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 1, characterized in that: The speech synthesis model includes an acoustic module and a vocoder module.
8. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 7, characterized in that: The acoustic module takes speech synthesis data as input and outputs an acoustic feature sequence, and the vocoder module converts the acoustic feature sequence into a time-domain speech signal.
9. The method of real-time interactive 3D digital holographic cabin based on deep learning and sound cloning according to claim 8, characterized in that: The acoustic module adopts the FastSpeech2 network, including a text encoder, a duration predictor, a pitch predictor, an energy predictor and a decoder, and the vocoder module adopts the HiFi-GAN network.
Citation Information
Patent Citations
Interactive digital human synthesis method based on voice-driven artificial intelligence
CN118969009A
Voiceprint recognition, identity confirmation and dialogue implementation method applied to sentiment analysis
CN119993168A