A method and system for sound replication and low-delay streaming speech synthesis based on ultra-short samples

By using an ultra-short sample voice replication module and a low-latency streaming speech synthesis engine, the problems of homogeneous timbre, high sample size, long latency, and poor environmental adaptability in traditional TTS systems are solved, achieving efficient, personalized, and low-latency speech synthesis, thus improving the fluency of intelligent voice interaction and user experience.

CN120748417BActive Publication Date: 2026-05-01GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG CHAOTENG INFORMATION TECHNOLOGY CO LTD
Filing Date
2025-07-17
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, traditional TTS systems suffer from problems such as homogenized timbre, high sample size requirements, high customization costs, long latency, and poor environmental adaptability. This makes it difficult to achieve rapid customization of personalized timbre and ultra-low latency speech synthesis, affecting the smoothness of intelligent voice interaction and user experience.

Method used

Employing an ultra-short sample sound replication module and a low-latency streaming speech synthesis engine, it rapidly extracts and separates timbre, style, and environmental features through deep multi-scale coding and adversarial decoupling techniques. Combined with low-latency bidirectional streaming output, it achieves efficient generation and real-time adaptation of personalized timbres.

Benefits of technology

It achieves high-quality sound replication based on ultra-short samples, greatly reducing the technical threshold and cost of timbre customization, providing an ultra-low latency real-time interactive experience, enhancing the personalization and environmental adaptability of speech synthesis, and improving the fluency and user experience of intelligent voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748417B_ABST
    Figure CN120748417B_ABST
Patent Text Reader

Abstract

The application provides a sound replication based on ultra-short samples and a low-delay streaming voice synthesis method and system, relates to the technical field of artificial intelligence, is suitable for intelligent interaction, outbound service and multi-modal communication scenarios, and realizes deep personalized customization and ultra-low-delay real-time generation of intelligent voice interaction through an innovative sound replication module and a voice synthesis engine, supports two-way streaming interaction at the system level, thereby improving the fluency and response speed of the dialogue; the ultra-short sample sound replication module specially designed for processing ultra-short audio samples and the voice synthesis engine with ultra-low-delay and two-way streaming output capability are integrated, the capabilities of the two are applied to the real-time and interactive intelligent voice interaction process, a complete and efficient solution is formed, and the important commercial application value is directly targeted to solve the industry pain points.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for sound reproduction and low-latency streaming speech synthesis based on ultra-short samples Technical Field

[0001] This invention relates to the field of artificial intelligence, and more particularly to a method and system for sound replication and low-latency streaming speech synthesis based on ultra-short samples. Background Technology

[0002] Currently, the continuous advancements in artificial intelligence technologies, particularly ASR (Automatic Speech Recognition), NLP (Natural Language Processing), and LLM (Large Language Modeling), have driven the widespread application of AI-powered intelligent voice interaction systems, which have become key tools for enterprises to improve efficiency and optimize customer interactions.

[0003] However, existing technologies have the following shortcomings: First, traditional TTS systems rely on pre-trained standard voice models, resulting in homogeneous and unpersonalized voices. This makes it difficult to meet enterprises' needs for personalized synthesized speech. It is also difficult to use specific voices or adjust the voice style according to business scenarios, leading to a stiff intelligent voice interaction experience that lacks trust and friendliness. If it were possible to analyze subtle changes in speech (such as intonation and rhythm) to perceive the user's emotional state and adjust the response accordingly, it would not only enhance the realism of human-computer interaction but also provide users with a more considerate service experience.

[0004] Secondly, the target speaker typically needs to provide a large amount of high-quality recording data (tens of minutes to several hours) for model training. The high sample size requirement leads to high customization costs and long cycles, making it difficult to apply to business scenarios that require frequent changes or rapid generation of specific voice timbres. Thirdly, in real-time AI intelligent voice interaction, the cumulative latency throughout the entire process (from user voice input to voice playback), especially the excessively long TTS synthesis stage (generally greater than 500 milliseconds), can cause untimely robot responses and dialogue stuttering, affecting call fluency and user acceptance. Existing TTS technology struggles to achieve extremely low end-to-end latency while ensuring high-quality synthesized speech, particularly the time from text input to the output of the first playable audio data packet.

[0005] Finally, traditional voice replication and synthesis technologies are mostly trained and optimized in ideal, noise-free environments, while the actual intelligent voice interaction environment is complex, including background noise, channel distortion, etc. The voice generated in an ideal environment may have problems such as timbre distortion, reduced clarity, or unnatural sound in actual interaction.

[0006] Given the limitations of the existing technologies, there is an urgent need for a new technical solution to quickly and accurately replicate high-quality personalized timbres based on a very small number of samples, apply the replicated timbre model to a system that can achieve ultra-low latency, two-way streaming speech synthesis, and ultimately integrate it into an intelligent interactive process to provide a more intelligent, efficient, and user-friendly AI interactive communication service. Summary of the Invention

[0007] The purpose of this invention is to provide a method and system for voice replication and low-latency streaming speech synthesis based on ultra-short samples, applicable to intelligent interaction, communication services, and multimodal communication scenarios, to solve the problems mentioned in the background art, such as homogenized timbre, high requirements for voice replication samples, large real-time interaction latency, and poor environmental adaptability. This invention, through an innovative voice replication module and speech synthesis engine, achieves deep personalization of synthesized speech and ultra-low latency real-time generation, and supports bidirectional streaming interaction at the system level, thereby improving the fluency and response speed of dialogue.

[0008] To achieve the above objectives, the core feature of this invention is the integration of an ultra-short sample sound replication module specifically designed for processing ultra-short audio samples and a speech synthesis engine with ultra-low latency and bidirectional streaming output capabilities, and the application of the capabilities of both to a real-time, interactive intelligent interactive process.

[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0010] In a first aspect of the present invention, a method for sound replication and low-latency streaming speech synthesis based on ultra-short samples is provided, comprising the following steps:

[0011] Users or system administrators can upload a target speaker audio sample of less than 15 seconds in length through the management interface and set style guidance parameters and acoustic environment modeling instructions.

[0012] The audio samples are processed using a pre-trained sound replication model to generate personalized timbre data containing timbre, style and environmental features. The generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID.

[0013] The current timbre ID is determined according to the preset requirements of the scene, and the corresponding personalized timbre data is retrieved from the timbre repository according to the timbre ID. Acoustic features are synthesized according to the text content, personalized timbre data and emotion control parameters. The acoustic features are converted into audio data blocks through a vocoder, and then the audio blocks are pushed to the player for playback.

[0014] Based on the user's voice input, subsequent response text is generated through LLM, and real-time emotion control parameters are dynamically generated and updated according to the progress of the conversation with the user and the analysis results. The latest audio data block is generated based on the context of the processed text and the retrieved personalized timbre data, and continuously pushed to the player for playback to achieve continuous voice playback.

[0015] Preferably, the style guidance parameters represent parameters that guide the replication of a specific speaking style. The style guidance parameters include preset style tags, which include standard customer service, dynamic marketing, and personalized reference audio. The acoustic environment modeling instructions are used to reproduce background noise and channel characteristics in the sample and to simulate the human voice effect.

[0016] Preferably, the personalized timbre data includes timbre, style, and environmental features, wherein the environmental features are those that simulate intelligent voice interaction.

[0017] Preferably, the timbre ID represents an identifier for personalized timbre data used in synthesis. The corresponding timbre embedding, style representation, and acoustic environment feature model are retrieved from the timbre repository using the timbre ID through a low-latency bidirectional streaming speech synthesis engine.

[0018] Preferably, the emotion control parameters are used to dynamically adjust the synthesis parameters of the generated speech. The emotion control parameters include speech rate control parameters, volume control parameters, emotion intensity parameters, and pause control commands. The speech rate control parameters are set to 50-200 words / minute, and the emotion intensity coefficient is set to 0.8-1.5 times the reference value.

[0019] Preferably, the acoustic features are scored and evaluated according to a speech expression effectiveness evaluation function, which is as follows:

[0020] ,

[0021] in, This represents the speech performance score for input audio sample i. Indicates the speech rate deviation coefficient. This represents the volume deviation coefficient. Indicates the emotional intensity deviation coefficient. This represents the speech rate of the input audio sample i. This indicates the number of words spoken per unit of time. This represents the volume of the input audio sample i. This indicates the standard volume, with a value of 50 dB. Indicates the intensity of emotion;

[0022] When the speech performance score is lower than a preset threshold, the target speaker's audio sample needs to be uploaded again and the acoustic features regenerated. When the speech performance score is higher than the preset threshold, no action is required.

[0023] Preferably, the step of updating the real-time emotion control parameters includes:

[0024] The low-latency bidirectional streaming speech synthesis engine receives the text analysis results, emotion control parameters, and target acoustic environment simulation parameters from the user's speech input;

[0025] When generating acoustic features using a vocoder model, the fundamental frequency of the speech is dynamically adjusted, including adjusting pitch, loudness or stress, and speech rate and pauses; and speech that conforms to a specified prosody and emotional tendency is generated through spectral characteristics.

[0026] In a second aspect of the invention, a system for sound replication and low-latency streaming speech synthesis based on ultra-short samples is provided, comprising:

[0027] The audio sample input module is used to input a target speaker audio sample with a duration of less than 15 seconds, and to set style guidance parameters and acoustic environment modeling instructions;

[0028] The ultra-short sample sound replication module is used to process audio samples using a pre-trained sound replication model to generate personalized timbre data containing timbre, style and environmental features. The generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID.

[0029] A low-latency bidirectional streaming speech synthesis engine is used to determine the current timbre ID according to the preset requirements of the scene, and retrieve the corresponding personalized timbre data in the timbre repository according to the timbre ID. It synthesizes acoustic features based on text content, personalized timbre data and emotion control parameters, converts the acoustic features into audio data blocks through a vocoder, and then pushes the audio blocks to the player for playback.

[0030] Preferably, the low-latency bidirectional streaming speech synthesis engine further includes a block-aware causal streaming generation submodule, which is used to generate subsequent response text through LLM based on the user's voice input, dynamically generate and update real-time emotion control parameters based on the progress of the dialogue with the user and the analysis results, generate the latest audio data block based on the context of the processed text and the retrieved personalized timbre data, and continuously push it to the player for playback to achieve continuous playback of the speech.

[0031] Compared with the prior art, the beneficial effects of the present invention include the following:

[0032] 1. Revolutionary sample efficiency: High-quality sound replication is achieved based on ultra-short samples (no more than 15 seconds), which greatly reduces the technical threshold, time and economic cost of timbre customization.

[0033] 2. Ultimate Real-Time Interactive Experience and Fast Turn-by-Turn Response: The ultra-low latency bidirectional streaming speech synthesis engine keeps the text-to-speech conversion latency to an extremely low level (less than 200 milliseconds TTFB) and supports receiving incremental text and starting output immediately. This capability, combined with system-level pipelined processing, significantly shortens the response time between dialogue turns, providing a smooth and efficient interactive experience.

[0034] 3. High fidelity and deep personalization: Even with extremely short input samples, the innovative deep multi-scale coding and adversarial decoupling techniques, which are pre-trained, are used to effectively capture and replicate the unique timbre, style and accent of the target speaker, resulting in high-quality synthesized speech with significant personalized features.

[0035] 4. Enhanced acoustic environment adaptability and simulation capabilities: The voice replication module has the ability to model and control the acoustic environment. The speech synthesis engine can use this information or simulate specific acoustic environments according to instructions, so that the replicated timbre or synthesized speech can better adapt to the complex actual call environment, sounding more realistic and stable.

[0036] 5. Intelligent and dynamic expressive ability: The synthesis engine combines text analysis and real-time interactive status to dynamically adjust the output speech rate, volume, rhythm and emotion, making the robot's speech expression richer, more natural and more infectious.

[0037] 6. Systematic Integration Value: The above-mentioned innovative technologies are seamlessly integrated into the core processes of the intelligent voice interaction system, forming a complete and efficient solution that directly addresses industry pain points and has significant commercial application value. Attached Figure Description

[0038] Figure 1 is a flowchart of a method for sound replication and low-latency streaming speech synthesis based on ultra-short samples according to the present invention.

[0039] Figure 2 is a framework diagram of a sound replication and low-latency streaming speech synthesis system based on ultra-short samples according to the present invention. Detailed Implementation

[0040] Please refer to Figure 1. In a first aspect, the present invention relates to a method for sound replication and low-latency streaming speech synthesis based on ultra-short samples, comprising the following steps:

[0041] Users or system administrators can upload a target speaker audio sample of less than 15 seconds in length through the management interface and set style guidance parameters and acoustic environment modeling instructions.

[0042] The ultra-short sample sound replication module efficiently and effectively captures and models key timbre, speaking style, accent, and acoustic environment features, generating structured and personalized timbre data that can be used for subsequent speech synthesis.

[0043] The audio samples are processed using a pre-trained sound replication model to generate personalized timbre data containing timbre, style and environmental features. The generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID.

[0044] This module relies on a powerful deep learning model that has been thoroughly trained offline with a large amount of diverse speech data, learning the ability to extract, decouple, and represent key acoustic features from speech. When processing ultra-short sample audio samples uploaded by users, this pre-trained model is used to perform efficient feature extraction and modeling, generating personalized timbre data.

[0045] Input and output of deep learning models:

[0046] enter:

[0047] Target timbre audio sample: A clear recording of a single speaker, lasting between 10 and 15 seconds.

[0048] Style control parameters: Parameters used to guide the replication of a specific speaking style, using preset style tags ("standard customer service", "energetic marketing", etc.) or a reference audio with a representative style.

[0049] Acoustic environment modeling instruction: instructs the module to learn and reproduce background noise or channel characteristics in the sample (simulating telephone line effects).

[0050] Output:

[0051] Personalized timbre data: encoded target speaker timbre embedding, style representation, and acoustic environment feature model. This data is stored in the system's internal timbre repository and associated with a unique timbre ID.

[0052] Timbre ID: A unique identifier for retrieving this personalized timbre data.

[0053] The specific procedures and issues to be addressed are as follows:

[0054] a) Solving the problem of robust coding of voiceprint and style features under ultra-short sample conditions:

[0055] Problem: Ultra-short audio samples have limited information content and are easily interfered with. The key challenge is how to extract stable, high-quality speaker identity (timbre) and speaking style features that are unaffected by the content of the speech.

[0056] Solution: Preprocessed (noise reduction, normalization) ultra-short audio samples are input into a pre-trained deep multi-scale robust speaker and style encoder. This encoder employs a carefully designed neural network architecture, whose core processing steps include:

[0057] ① Initial feature extraction: First, the audio signal is converted into a time-frequency representation (Mel spectrogram), which is an acoustic feature map that simulates human auditory perception, facilitating subsequent network processing.

[0058] ② Multi-scale parallel analysis: The encoder contains parallel or multi-layered processing paths trained to analyze spectrograms at different time lengths and frequency granularities. Some layers use small-scale convolutional kernels (an operation that extracts features in local regions of data) to capture short-term speech details, such as the formant structure of vowels or the articulation features of consonants; other layers use larger-scale convolutional kernels or sequence models (recurrent neural networks (RNNs) or Transformer networks, two types of neural networks adept at handling sequential data and capturing long-term dependencies) to capture features over longer time ranges, analyzing speech rate, intonation curves, energy envelopes, etc. This multi-scale analysis ensures that both the micro-features of vocalization and the macro-features of speech flow can be effectively captured.

[0059] ③ Attention Focus: During feature extraction, the encoder integrates an attention mechanism. An attention mechanism is a network technique that allows the model to automatically focus on the most important parts of the input information. It allows the model to dynamically allocate more computational resources or weights to the regions of the signal that contain the richest speaker identity or style information, such as a clear pronunciation segment or a typical intonation change, based on the characteristics of the input signal. This enables the encoder to "find" and prioritize the most representative parts in ultra-short samples, effectively suppressing noise or irrelevant transient information, and improving the accuracy and robustness of feature extraction.

[0060] ④ Feature Fusion and Encoding: Features at different scales and after attention weighting are finally fused and compressed into compact vector representations through subsequent layers of the network, namely timbre embedding (a feature representing the speaker's unique vocal quality) and style representation (a feature representing the speaker's habitual speech rate, intonation patterns, etc.). These vectors are trained to be as independent of the speech content as possible.

[0061] b) Solving the problem of the tight coupling and separation of timbre, content, and style:

[0062] Problem: In any speech segment, timbre, the specific content of the speech (text / phonemes), and speaking style (pace, emotion, stress) are naturally blended together. To independently synthesize arbitrary text using the extracted timbre and style, these three elements must be effectively separated (decoupled) from the original audio signal. For very short samples with limited information, accurate decoupling is even more challenging.

[0063] Solution: Leveraging the powerful timbre-content-style decoupling mechanism of a pre-trained model. The core capability of this mechanism is learned during offline training by constructing complex adversarial training tasks. The model employs a framework based on Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs) combined with adversarial training. During training, the model is trained to decompose input audio into at least three independent latent representations: timbre representation, content representation (typically associated with the input text or phoneme sequence), and style representation. Multiple discriminators are introduced for adversarial training. One discriminator attempts to determine the content of the original audio based solely on the timbre representation, while the training objective is to make the timbre representations produced by the generator (decoder) and encoder impossible for this discriminator, thus forcing the timbre representations to lack content information.

[0064] Similarly, training also ensures decoupling between timbre and style, and between content and style. When dealing with very short samples, the pre-trained model applies this decoupling capability to refine and separate the purest possible timbre embeddings and style representations from the initial features generated by the encoder, striving to reduce the interference of the original speech content or transient emotions on these representations.

[0065] c) Addressing the impact of the acoustic environment on replication results and simulation requirements:

[0066] Problem: The acoustic environment (background noise, reverberation, channel effects) of the original recording sample will mix into the speech signal. Direct replication will include these ambient sounds, affecting the purity of the timbre. At the same time, in scenarios such as intelligent voice interaction and AI outbound calls, it is sometimes necessary to add specific environmental effects (simulating telephone channels) during synthesis to increase realism.

[0067] Solution: Utilizes the adaptive acoustic environment modeling capabilities inherent in a pre-trained model. During offline training, this model learns to recognize and represent non-speaker acoustic features in speech signals. When processing ultra-short samples, a specific branch or mechanism within the model analyzes the audio, extracting feature representations of its acoustic environment (modeling the statistical characteristics of background noise, frequency response characteristics of the telephone channel, etc.). This solution transforms sound replication from simple audio copying into intelligent perception, modeling, purification, and control of the acoustic environment, significantly improving the applicability and auditory realism of the generated personalized timbre in complex real-world environments. The system processes the extracted environmental features according to user-defined acoustic environment modeling instructions.

[0068] ①Learn and store environmental features: If the instruction requires the reproduction of the sample environment, the extracted environmental features will be stored as part of the personalized timbre data and used to reproduce a similar acoustic environment during subsequent synthesis.

[0069] ② Environmental cleansing: If the instruction requires a pure timbre, the model will apply the environmental robustness or cleansing techniques learned during training to try to remove the identified environmental influences from the timbre embedding and generate a timbre representation that is closer to the speaker's pure voice.

[0070] ③Environmental simulation control: Even if a pure tone is extracted from a sample, the ability of the environment modeling branch can be used to inject artificial or preset environmental effects into the pure tone according to instructions (analog telephone channel) during the synthesis stage to increase realism.

[0071] Workflow (sound replication and data storage):

[0072] Sample Upload and Parameter Settings: Users or administrators can upload a target speaker audio sample with a duration of ≤ 15 seconds through the system interface, and set style guidance parameters and acoustic environment modeling instructions.

[0073] Preprocessing: The uploaded audio is normalized (resampled to 16kHz, volume normalized) and basic noise reduction is performed.

[0074] Feature extraction and data generation: The preprocessed audio is fed into a pre-trained ultra-short sample sound replication model.

[0075] Model Execution: Preliminary features are extracted using a deep multi-scale robust voiceprint and style encoder; a timbre-content-style decoupling mechanism based on adversarial training is applied to separate clean timbre embeddings and style representations; adaptive acoustic environment modeling is performed according to instructions, extracting or processing environmental features. Finally, personalized timbre data containing timbre, style, and environmental features is generated.

[0076] Data storage and ID assignment: The generated personalized timbre data is stored in the system's timbre database or repository. The system assigns a unique timbre ID to it and returns this ID to the user or system for use in subsequent speech synthesis.

[0077] Subsequently, a low-latency bidirectional streaming speech synthesis engine receives the text stream to be synthesized and the specified personalized voice ID from the upper-layer application (NLP or LLM output). By retrieving the corresponding personalized voice data, it generates high-quality, natural speech with the specified voice and style in a bidirectional streaming manner with extremely low end-to-end latency (the time from receiving text to outputting the first audio data packet, target < 200 milliseconds). Here, "bidirectional streaming" means that the TTS engine can receive incremental text input and immediately start streaming audio output, thereby supporting system-level pipelined processing and faster dialogue turn responses.

[0078] Input and output of a low-latency bidirectional streaming speech synthesis engine:

[0079] enter:

[0080] Text stream to be synthesized: Response statements or paragraphs of text provided by the system (dialogue management or LLM module) in incremental, chunked, or streaming form. The engine is configured to receive and process partial or continuously appended text input without waiting for complete sentences or paragraphs.

[0081] Personalized Timbre ID: Specifies an identifier for the personalized timbre data used in synthesis. The engine uses this ID to retrieve the corresponding timbre embedding, style representation, and acoustic environment feature model from the timbre repository.

[0082] Real-time control parameter stream: Signals from the dialogue management or sentiment analysis module, used to dynamically adjust the synthesis parameters of subsequent speech blocks. This includes, but is not limited to: speech rate control parameters, volume control parameters, emotion intensity parameters, and pause control commands. These parameters can be dynamically adjusted based on the dialogue context or real-time user feedback; wherein the speech rate control parameter is set at 50-200 words / minute, and the emotion intensity coefficient is set at 0.8-1.5 times the baseline value.

[0083] Target acoustic environment simulation parameters: used to superimpose specific acoustic environment effects during the synthesis stage, simulating telephone channel noise, reverberation, or sounds of specific scenes.

[0084] Output:

[0085] Real-time audio stream: A continuous stream of digital audio data (preferably in PCM format with a 16kHz sampling rate) output in small data blocks, with specified timbre, style, emotion, and rhythm, and adjusted speech rate, volume, and simulated ambient sounds according to emotion control parameters.

[0086] Core algorithms and workflow of a low-latency bidirectional streaming speech synthesis engine (real-time inference and generation):

[0087] The low-latency bidirectional streaming speech synthesis engine comprises a pre-trained speech synthesis model and a pre-trained, highly efficient vocoder. Speech synthesis is an online, real-time inference process that deeply leverages personalized timbre data and real-time emotion control parameters.

[0088] a) Solving the problem of generating low-latency streaming acoustic features from incremental text and personalized data:

[0089] Problem: Traditional TTS models require complete text input to begin synthesis, and their internal processing can introduce significant latency. To achieve low latency and system-level bidirectional streaming interaction, a model is needed that can receive incomplete text and immediately begin generating acoustic features, while incorporating personalized timbre and style, and ensuring the naturalness and fluency of the generated speech, especially in the transitions between text blocks.

[0090] Solution: A pre-trained block-aware causal streaming generation model is used. This model is a complex deep neural network whose core design is for efficiently processing sequential data and providing streaming output capabilities. This model achieves low latency and streaming output in the text-to-acoustic feature conversion process, effectively utilizing incremental text and personalized timbre data. Through causal processing and inter-block optimization, it helps ensure the quality and fluency of synthesized speech, directly supporting the needs of AI intelligent voice interaction systems for rapid response and bidirectional streaming interaction. The specific working mechanism includes:

[0091] ① Incremental Text Reception and Chunking Processing: The engine continuously receives the text stream to be synthesized from the upstream system. The text is automatically or segmented into small text blocks based on punctuation and semantic units. The engine processes the received text blocks immediately without waiting for the end of the entire sentence or paragraph.

[0092] ② Personalized Condition Loading and Application: The engine quickly retrieves corresponding personalized timbre data (timbre embedding, style representation, environmental features, etc.) from the timbre repository based on the input timbre ID, and inputs this data as conditional information into the generative model. The model is trained to generate acoustic features with the target timbre and style based on these conditions.

[0093] ③ Causal Acoustic Feature Generation: The generation model processes text blocks as units. When generating acoustic features (e.g., Mel spectrogram blocks) corresponding to the current text block, the model's calculation strictly follows the principle of causality, relying only on the content of the current text block, the input personalized timbre data, real-time emotion control parameters, and the contextual information of previously processed text blocks. The model maintains a state internally to remember the contextual information of previous blocks, thereby ensuring that acoustic feature blocks can be generated and output immediately in text order, achieving a delay as low as the first byte (block).

[0094] ④ Efficient Intra-Block Generation Mechanism: The model employs flow matching or consistency models internally. These techniques allow the model to quickly generate a complete acoustic feature block from the latent space (controlled by conditions such as text, timbre, and style) in a more parallel or efficient manner than traditional autoregressive models (e.g., prediction per audio sample). This efficient intra-block generation is key to achieving overall low latency.

[0095] ⑤ Natural inter-block transitions: The model is optimized during training to take into account the contextual information of block boundaries when generating acoustic feature blocks, ensuring a smooth transition of acoustic features generated from different text blocks at the boundaries. Furthermore, signal processing techniques such as overlap-addition can be applied to the acoustic feature domain to further optimize the connections between blocks, helping to ensure the fluency and natural rhythm of the final synthesized speech.

[0096] b) Solve the problem of high-efficiency, high-quality conversion of acoustic features to speech waveforms:

[0097] Problem: Converting acoustic features into a final audible speech waveform is the last step in the speech synthesis pipeline. Traditional vocoders can be a computational bottleneck, introducing additional latency. A high-efficiency vocoder that is fast enough while maintaining sound quality is needed.

[0098] Solution: A pre-trained, lightweight, and highly efficient vocoder is used. This vocoder is a deep neural network specifically trained to rapidly convert acoustic features into raw audio waveforms (PCM data blocks). The system uses HiFi-GAN or a vocoder model with similar non-autoregressive, highly efficient generation capabilities. These vocoders are designed to generate waveforms in parallel in a non-autoregressive manner, meaning they do not need to generate audio samples one by one like traditional vocoders, but can generate an audio block at once or in fewer steps, resulting in extremely high computational efficiency. The vocoder receives the acoustic feature blocks output by the synthesis model and efficiently converts them into the final audio data blocks. This invention employs a highly efficient vocoder optimized for real-time inference scenarios, aiming to ensure that the conversion process from acoustic features to the final speech waveform matches the speed of the acoustic feature generation process, minimizing its role as a bottleneck in the entire low-latency streaming synthesis chain, which is key to achieving end-to-end low latency.

[0099] c) Solving the problems of dynamic adjustment of speech expression and acoustic environment simulation:

[0100] Problem: In order to provide a more natural and intelligent AI-powered voice interaction experience, synthesized speech needs to be able to dynamically adjust its speech rate, volume, rhythm and emotion according to the dialogue content, context and real-time control signals, and be able to simulate specific acoustic environments (such as telephone channel effects) as needed.

[0101] Solution: These control capabilities are integrated into a block-aware causal streaming generation model and / or a high-efficiency vocoder, driven by the input real-time emotion control parameter stream and target acoustic environment simulation parameters. This deeply integrates dynamic control of speech rate, volume, prosody, and emotion, as well as acoustic environment simulation capabilities, into a low-latency streaming synthesis pipeline, making it subject to real-time parameter control. This significantly enhances the expressiveness, intelligence, and environmental adaptability of AI-powered intelligent voice interaction, providing a more human-like and immersive user experience.

[0102] ① Dynamic Prosody and Emotion Adjustment: The engine receives text analysis results (prosodic tags, sentiment tendencies) and real-time emotion control parameters (emotional intensity, pause instructions, etc.) from the upstream system. This information serves as additional conditional input to the generative model. The model is trained to dynamically adjust the fundamental frequency (determining pitch), energy (determining loudness / stress), and duration (determining speech rate and pauses) of the speech when generating acoustic features, based on these conditions, thereby generating speech that conforms to the specified prosody and emotional tendency. Real-time emotion control parameters allow for dynamic changes in the performance of subsequent speech during the synthesis process based on external signals.

[0103] ② Adaptive speech rate adjustment: The engine receives speech rate control parameters. This can be achieved in several ways: by adjusting the model's internal predictions of phoneme or syllable durations when generating acoustic features; or by adjusting the temporal scaling ratio of the acoustic features. The pre-trained synthetic model is trained to support speech rate variations within a certain range while maintaining the naturalness of the speech. This adjustment can be adaptive based on the importance of the text content or real-time instructions.

[0104] ③ Adaptive Volume Adjustment: The engine receives volume control parameters. This can be achieved in the acoustic feature domain (adjusting the energy of the acoustic features) or the waveform domain (adjusting the amplitude of the final waveform). The synthesis model or vocoder can adjust the loudness of the corresponding parts based on text analysis (e.g., emphasizing words) or real-time volume parameters. High-efficiency vocoders can also be controlled to proportionally adjust the loudness of the output waveform. This adjustment can be adaptive based on text content or real-time commands.

[0105] ④ Acoustic Environment Simulation: The engine receives simulation parameters for the target acoustic environment. The system can recreate the original environment using stored environmental features retrieved from personalized timbre data (if the environment was modeled during replication), or apply preset or learned environment transitions based on simulation parameters (e.g., a parameter representing a "telephone channel"). This can be achieved by influencing spectral characteristics during the acoustic feature generation stage, or by superimposing specific noise spectra and applying channel filters by the vocoder during waveform generation. This functionality makes the synthesized speech sound more natural and realistic in specific call environments.

[0106] The current timbre ID is determined according to the preset requirements of the scene, and the corresponding personalized timbre data is retrieved from the timbre repository according to the timbre ID. Acoustic features are synthesized according to the text content, personalized timbre data and emotion control parameters. The acoustic features are converted into audio data blocks through a vocoder, and then the audio blocks are pushed to the player for playback.

[0107] The first segment of text is sent to the low-latency bidirectional streaming speech synthesis engine in either streaming or chunked form, along with a specified personalized timbre ID and an initial stream of real-time emotion control parameters (e.g., default speech rate, volume). The low-latency bidirectional streaming speech synthesis engine receives the timbre ID, retrieves the corresponding personalized timbre data from its timbre repository, loads the pre-trained speech synthesis model and vocoder, and immediately begins processing. The low-latency bidirectional streaming speech synthesis engine's chunk-aware causal streaming generation model rapidly synthesizes acoustic features based on text chunks, personalized timbre data, and emotion control parameters, and converts these acoustic features into audio data chunks using a lightweight and efficient vocoder. The first audio chunk is generated and pushed to FreeSWITCH for playback in less than 200 milliseconds.

[0108] Based on the user's voice input, subsequent response text is generated through LLM, and real-time emotion control parameters are dynamically generated and updated according to the progress of the conversation with the user and the analysis results. Based on the subsequent response text and the real-time emotion control parameters, the latest audio data block is generated based on the context of the processed text and the retrieved personalized timbre data, and continuously pushed to the player for playback to achieve continuous playback of the voice.

[0109] While the TTS engine plays the first audio segment, the ASR module begins listening to the user's voice input, and the NLP / LLM module may also be processing user input or preparing subsequent response text. When the LLM generates the first part or a text block of the subsequent response text, the system immediately sends it to the TTS engine incrementally. Simultaneously, based on the dialogue progress and the analysis results (e.g., emotional changes), the system dynamically generates and updates the real-time emotion control parameter stream (e.g., if the user's speech rate increases, the system may increase the robot's speech rate parameter; if the system detects that the user is emotionally agitated, it may adjust the emotion intensity parameter and decrease the volume).

[0110] The acoustic features are scored and evaluated according to a speech expression effectiveness evaluation function, which is as follows:

[0111] ,

[0112] in, This represents the speech performance score for input audio sample i. Indicates the speech rate deviation coefficient. This represents the volume deviation coefficient. Indicates the emotional intensity deviation coefficient. This represents the speech rate of the input audio sample i. This indicates the number of words spoken per unit of time. This represents the volume of the input audio sample i. This indicates the standard volume, with a value of 50 dB. Indicates the intensity of emotion;

[0113] When the speech expression effect score is less than the preset threshold, the target speaker's audio sample needs to be uploaded again and the acoustic features need to be regenerated. When the speech expression effect score is greater than the preset threshold, no operation is required.

[0114] The dynamic weight calculation method for the speech rate compensation mechanism (α term):

[0115] Speech rate feature extraction: The original speech signal is preprocessed, including framing and windowing, and then the speech rate-related feature parameters, such as fundamental frequency and energy, are obtained by using a feature extraction algorithm based on the vocal tract model (such as LPC or MFCC) as the original basis for measuring speech rate.

[0116] Speech rate difference calculation: The extracted raw speech rate feature values ​​are compared with the set standard speech rate template (obtained through statistical analysis of a large amount of speech data), and the degree of difference is calculated to obtain the speech rate difference value. The difference value can be calculated using methods such as Euclidean distance and cosine similarity to quantify the degree of deviation between the current speech rate and the standard speech rate.

[0117] Dynamic weight adjustment: Based on the speech rate difference and combined with a pre-set weight adjustment function, the dynamic weight α of the speech rate compensation mechanism is calculated in real time. The weight adjustment function is usually designed as a non-linear function, such as the sigmoid function, to ensure that when the speech rate difference is large, the weight can increase rapidly, thereby significantly compensating for the speech rate of the synthesized speech; while when the speech rate difference is small, the weight changes relatively slowly to ensure the naturalness of the synthesized speech.

[0118] Spectrum reconstruction algorithm for environmental noise modeling method (β term):

[0119] Noise classification and feature extraction: First, different types of environmental noise (such as white noise, traffic noise, workshop noise, etc.) are classified and collected. Then, short-time Fourier transform (STFT) is performed on each type of noise sample for frequency domain analysis to extract its key feature parameters, including the noise's spectral amplitude, phase information, and power spectral density, and a noise feature library is constructed.

[0120] Target speech spectrum estimation: For the noisy speech signal to be processed, STFT transformation is also performed to obtain the mixed spectrum. Using a deep learning-based spectrum separation algorithm (such as U-Net or RNN), combined with a noise feature library, the target speech spectrum and noise spectrum in the mixed spectrum are initially separated and estimated. By training a deep learning model, it can automatically learn the differences between the speech spectrum and noise spectrum under different noise environments, thereby achieving more accurate spectrum separation.

[0121] Spectrum Reconstruction and Optimization: After obtaining the initial target speech spectrum, an iterative optimization-based spectrum reconstruction algorithm is used for further optimization. Based on the separation results of the target speech spectrum and the noise spectrum, and combined with prior knowledge of the speech signal (such as the spectral smoothness and continuity of speech), the spectral parameters are iteratively updated, continuously adjusting the spectral amplitude and phase information to make the reconstructed speech spectrum more consistent with the spectral characteristics of real speech, while minimizing noise residue. During the iteration process, appropriate constraints (such as minimizing total variation, sparsity constraints, etc.) can be introduced to ensure the stability and accuracy of the reconstructed spectrum, ultimately obtaining a relatively clean and clear target speech spectrum.

[0122] Multimodal fusion strategy for quantifying emotional intensity (λ term)

[0123] Multimodal data preprocessing: Preprocessing is performed on multimodal data such as speech, text, and vision. For speech data, voice activity detection (VAD) and endpoint detection are performed to extract speech segments and calculate feature parameters such as pitch, timbre, and speech rate. For text data, word segmentation, part-of-speech tagging, and semantic analysis are performed to extract sentiment keywords and sentiment information. For visual data (such as facial expressions and body movements in videos), computer vision techniques (such as CNNs) are used for face detection, expression recognition, and body movement analysis to extract emotion-related visual feature parameters.

[0124] Intramodal feature fusion: Within each modality, multiple extracted features are fused. In the speech modality, methods such as weighted averaging and feature concatenation can be used to fuse features like pitch, timbre, and speech rate into a comprehensive speech emotion feature vector. In the text modality, semantic analysis results and the emotion weights of sentiment keywords are combined to construct a text emotion feature vector. In the visual modality, facial expression features and body language features are fused temporally (e.g., using LSTM networks) to obtain a visual emotion feature vector. The purpose of intramodal feature fusion is to fully utilize multiple pieces of information within the same modality and improve the modality's ability to represent emotional intensity.

[0125] Intermodal Feature Fusion and Weight Calculation: The emotional feature vectors from speech, text, and vision modalities are integrated, and an attention mechanism is employed to automatically learn the weight of each modality in the current emotional expression. A multilayer perceptron (MLP) network is constructed, taking the emotional feature vectors from each modality as input. After nonlinear transformation by the network, the weight coefficients corresponding to each modality are output. The core idea of ​​the attention mechanism is to allow the model to automatically focus on the modal features most representative of emotional intensity judgment, dynamically allocating weights according to the importance of different modalities in specific emotional scenarios, thereby achieving adaptive fusion of multimodal features.

[0126] Emotional intensity quantification and output: Based on the calculated modal weight coefficients, the integrated emotional feature vector after intermodal fusion is weighted and summed to obtain the final emotional intensity quantification value λ. This quantification value can comprehensively reflect the emotional intensity jointly expressed by speech, text, and visual multimodal information, providing a more accurate basis for emotional expression for speech synthesis systems, enabling synthesized speech to convey different emotional semantics more vividly and realistically.

[0127] Throughout the synthesis and playback process, users will hear the robot's speech rate, volume, tone, etc., change naturally as the conversation progresses or external commands are given. For example, the speech rate may slow down when broadcasting important information, the tone may be calmer when the user is emotional, or there may always be a simulated telephone line ambient sound.

[0128] End call: The call process terminates when the conversation achieves its business objectives, the user hangs up, or other termination conditions are met.

[0129] The following is a comparison of latency and audio reproduction quality between the present invention and traditional technologies and competing products. Compared with traditional technologies and competing products, the present invention has significant improvements in both latency and audio reproduction quality.

[0130] Table 1: End-to-End Delay Comparison

[0131]

[0132] Table 2: Replication Quality of Ultra-Short Samples

[0133]

[0134] Traditional technologies suffer from high latency and inaccurate timbre reproduction. This invention, however, significantly reduces response latency through optimized algorithms and a highly efficient processing architecture, effectively improving response speed, particularly in multi-turn dialogue scenarios. In audio reproduction, it more accurately restores timbre details, enhancing similarity to the original audio in key indicators such as pitch and timbre, resulting in superior timbre reproduction quality and a more natural and high-quality voice interaction experience for users in various scenarios.

[0135] This invention can be integrated into existing ASR / TTS architectures, enabling upgrades to streaming synthesis capabilities through API interfaces.

[0136] The technical solutions of the present invention are further illustrated by the following embodiments, but are not limited to these embodiments.

[0137] Example 1: Replicating the gentle customer service voice and rapid dialogue rounds in a financial debt collection scenario:

[0138] Scenario: A bank needs to make outbound calls to remind customers of overdue payments and wants to use a specially trained, gentle, and friendly female customer service voice to alleviate customer resistance. The bank used an internal employee with a sweet voice, "Ms. Wang," to record approximately 12 seconds of sample speech with a gentle tone for voice replication. Simultaneously, the system needs to ensure that after a simple customer response, the chatbot can quickly respond without noticeable pauses and appropriately emphasize the text content.

[0139] Voice reproduction stage (offline):

[0140] ① Upload the 12-second audio sample of Ms. Wang to the sound replication module management interface of the system of this invention.

[0141] ② Set "Gentle Customer Service Style" as the style guidance parameter, and set "Analog Telephone Channel" as the environment modeling instruction.

[0142] ③ Activate the ultra-short sample voice replication module. The module uses a pre-trained voice replication model to process audio samples and generate personalized voice data that includes Ms. Wang's timbre, gentle style, and simulated telephone channel characteristics.

[0143] ④ Store the generated personalized tone data in the system's tone repository and assign a unique tone ID, such as "WangXiaojie_001".

[0144] Outbound calling application stage (online):

[0145] ① In the configuration of collection outbound calls for overdue customers, specify the use of voice ID "WangXiaojie_001".

[0146] ② The system initiates a call. After the user answers, the system sends the first part of the opening text, "Hello, this is XX Bank customer service. We have important information to share regarding your current bill. Are you Mr. Zhang?" ("Hello, this is XX Bank customer service. We have important information to share regarding your current bill"), to the low-latency bidirectional streaming speech synthesis engine in a streaming manner. Simultaneously, it sends the voice ID "WangXiaojie_001" and the initial real-time emotion control parameters. Text analysis indicates that "current bill" is the key point, and the corresponding volume control parameters are set slightly higher.

[0147] ③ The engine receives the timbre ID, text, and emotion control parameters, retrieves the corresponding personalized timbre data from the repository, and loads a pre-trained speech synthesis model and vocoder. The engine's block-aware causal streaming generation model quickly synthesizes acoustic features based on these inputs and applies slightly higher energy parameters to "this period's bill." This is then converted into audio data blocks using a lightweight, high-efficiency vocoder. The first audio block is generated within 170ms and pushed to FreeSWITCH for playback. The user hears natural speech with Ms. Wang's unique timbre and gentle tone, with "this period's bill" sounding slightly emphasized and featuring simulated telephone tone effects.

[0148] ④ While the engine is synthesizing and playing the first segment of text, the LLM module may have already determined the remaining text, "Are you Mr. Zhang?", based on business logic and sent it incrementally to the TTS engine. The corresponding real-time emotion control parameters indicate the use of standard volume and speech rate. After receiving this text and parameters, the engine seamlessly continues synthesis, ensuring a smooth transition with the previous segment of speech.

[0149] ⑤ After hearing "Are you Mr. Zhang?", the user responds "Yes". ASR recognizes the user's response.

[0150] ⑥ Based on the user's confirmation, the LLM quickly generates the next reply text, "Okay, Mr. Zhang, your bill for this period is overdue. Please handle it as soon as possible." This is then sent to the TTS engine, along with the voice ID and real-time emotion control parameters, either streaming or in chunks. Assume the system sets the robot's speech rate to slightly faster based on the user's response speed.

[0151] ⑦ After receiving the text, voice ID, and updated speech rate parameters, the TTS engine again utilizes its low-latency streaming synthesis capabilities to quickly generate and play the response speech, with a speech rate slightly faster than the opening remarks. Due to the engine's fast response and streaming processing, the interval between the user finishing speaking and the robot starting to speak the next sentence is extremely short, and the user perceives the dialogue turn response as very fast.

[0152] Example 2: Smooth information delivery and style switching in a news summary outbound call scenario:

[0153] Scenario: A media company wants to use AI to call subscribers and deliver summaries of the day's important news. The news summaries may be lengthy, requiring the robot to deliver them at a natural, fluent pace and tone, and different types of summaries may require different styles or voices.

[0154] System Configuration: The system uses the technology of this invention to pre-generate and store multiple timbre data by uploading a small number of samples and setting style and environmental parameters, and assigns corresponding timbre IDs, including a standard male announcer timbre ID (e.g., "NewsAnchor_Male_002", set not to simulate telephone channels and maintain a clear timbre), a lively female timbre ID (e.g., "Entertainment_Female_003", set not to simulate telephone channels), and a parameter configuration that simulates a live interview environment (e.g., "LiveInterview_Env_Param").

[0155] Outbound calling process:

[0156] ① The system dials the user's phone number and selects an appropriate voice ID (e.g., using "NewsAnchor_Male_002") and target acoustic environment simulation parameters (e.g., not simulating ambient sounds) based on the user's subscription preferences and news type. The initial real-time emotion control parameters are set to standard broadcast speed and volume.

[0157] ② The system sends the news summary text in segments, by sentence or paragraph, to the low-latency bidirectional streaming speech synthesis engine, while also sending the selected timbre ID, initial emotion control parameters, and environmental parameters.

[0158] ③ The engine loads the specified timbre data and model, receives and processes the first text block, and the block-aware causal streaming generation model quickly synthesizes the corresponding acoustic features. It is then rapidly converted into audio data blocks by a lightweight and efficient vocoder and pushed to the communication platform for playback.

[0159] ④ While playing the first audio segment, the system continues to send subsequent text segments to the TTS engine. Upon receiving a new text segment, the engine immediately begins processing it, utilizing the contextual information of the already processed text to ensure a smooth transition in rhythm and pace between the current segment and the previous one. Real-time emotion control parameters can dynamically fine-tune the pace and pauses based on punctuation or importance of the text.

[0160] ⑤ If a segment requires a change in style or simulation environment (e.g., broadcasting an entertainment news summary), the system will send the new text block along with the updated timbre ID (switching to "Entertainment_Female_003"), style parameters (lively), and possible target acoustic environment simulation parameters (e.g., "LiveInterview_Env_Param") to the engine.

[0161] ⑥ After receiving the updated parameters, the engine smoothly switches to a new female voice and lively style when synthesizing subsequent text blocks, and overlays simulated live interview ambient sound effects. This switching and simulation are completed seamlessly during low-latency streaming.

[0162] ⑦ The entire process continues, and the user hears a voice broadcast generated from a specified timbre, with a natural speaking speed and tone, no obvious pauses between blocks, and smooth switching between different styles and ambient sounds. This smooth broadcast experience is thanks to the TTS engine's low-latency bidirectional streaming processing capabilities, rapid response to incremental text and dynamic control commands, and integrated environment simulation capabilities.

[0163] Example 3, In-vehicle navigation scenario:

[0164] The low-latency bidirectional streaming speech synthesis engine of this invention is cleverly integrated into the in-vehicle system, enabling it to work seamlessly with other functional modules of the in-vehicle system. This integration method neither interferes with the original functions of the in-vehicle system nor fails to fully leverage the advantages of the synthesis engine, providing drivers with a superior voice navigation experience.

[0165] Noise Environment Model Construction: A dedicated noise environment model for vehicles was established to address the unique characteristics of the in-vehicle environment. This model comprehensively considers various noise factors generated during vehicle operation, such as engine noise, wind noise, and tire noise, as well as noise variations under different vehicle speeds and road conditions. Through the collection, analysis, and processing of extensive in-vehicle environmental noise data, a model that accurately reflects the characteristics of in-vehicle noise was constructed, providing a foundation for subsequent speech synthesis optimization.

[0166] Real-time loudness adjustment: When the synthesis engine generates navigation voice, the system monitors the noise level in the vehicle environment in real time and dynamically adjusts the loudness of the synthesized voice according to preset loudness adjustment rules. For example, when the vehicle is traveling at high speed and the noise level is high, the voice loudness is appropriately increased to ensure that the driver can hear the navigation instructions clearly; while when the vehicle is traveling at low speed or stationary, the voice loudness is reduced to avoid unnecessary interference to the driver, and at the same time, it can save energy consumption of the in-vehicle audio system.

[0167] Noise-resistant coding strategy optimization: Based on the vehicle noise environment model, an advanced noise-resistant coding algorithm is used to process the synthesized speech. This coding strategy can minimize the impact of noise on speech transmission and playback while ensuring speech quality. By adjusting the noise-resistant coding parameters in real time, the synthesized speech maintains high clarity and intelligibility in various complex vehicle noise environments, ensuring that drivers can accurately obtain navigation information and improving driving safety and navigation accuracy.

[0168] Example 4, Virtual Assistant Scenario:

[0169] This invention is fully deployed in smartphone voice assistants, making it one of the core functions of the voice assistant. Deep integration with the smartphone's operating system and related hardware ensures that the voice assistant can quickly respond to user voice commands and smoothly perform speech synthesis and playback. Simultaneously, the system's resource allocation and scheduling mechanisms are optimized to ensure that the voice assistant does not significantly impact the performance of other phone functions during operation, providing users with a personalized voice service experience.

[0170] This embodiment supports users uploading 15 seconds of voice as samples for customizing the assistant's voice. The system analyzes and processes the uploaded voice, extracting its voice features, including key parameters such as pitch, timbre, speech rate, and intonation. These parameters are then used to train and adjust the synthesis engine, generating a voice model highly similar to the user's voice. In subsequent voice interactions, the voice assistant will use this custom voice to further enhance user acceptance and usage frequency.

[0171] For multi-turn dialogue scenarios, the speech synthesis module has been deeply optimized to enable rapid response. During the dialogue, the voice assistant can understand the user's voice commands in real time and quickly generate corresponding voice responses. By employing efficient speech recognition, semantic understanding, and speech synthesis algorithms, as well as optimizing the system architecture, the dialogue response time has been significantly shortened, making the communication between the user and the voice assistant more natural and fluent, just like talking to a real person. No matter whether the user's question involves multiple fields or requires complex logical reasoning, the voice assistant can provide accurate and reasonable answers in a short time, improving user efficiency and satisfaction.

[0172] Example 5, Medical Consultation Scenario:

[0173] This invention is applied to a medical consultation system, with specialized adaptation and customization. Considering the professional and specific nature of medical consultations, the synthesis engine has been optimized to accurately synthesize the speech of medical terminology and related vocabulary, ensuring accuracy and standardization of the speech. Simultaneously, it interfaces with the hospital's information system to achieve patient data sharing and interaction, enabling the voice assistant to provide personalized medical consultation services based on the patient's specific condition and medical history.

[0174] Gentle Voice Replication: By collecting a large amount of speech data with gentle voice characteristics, the synthesis engine was trained and optimized to successfully replicate a gentle and friendly voice for medical consultation scenarios. This voice can bring patients a sense of security and trust, alleviate their anxiety during medical treatment, and enhance the interactivity and affinity between patients and the medical consultation system. During the speech synthesis process, the system dynamically adjusts the intensity and expression of the gentle voice according to the content and emotional needs of the dialogue, making the speech more natural, vivid, and close to the tone and intonation of a real doctor.

[0175] To address the numerous technical terms involved in medical consultations, the system employs specialized speech synthesis technology and processing algorithms to ensure these terms are clearly and accurately pronounced. Through phoneme decomposition, pitch adjustment, and speech rate control of technical terms, patients can easily understand doctors' diagnoses and treatment recommendations. Furthermore, the system can supplement and expand upon explanations of technical terms based on patient feedback and needs, thereby improving patients' understanding and acceptance of medical information, facilitating smoother doctor-patient communication, and enhancing the quality of medical services.

[0176] Example 6, Educational Tutoring Scenario:

[0177] By being deployed on online education platforms, educators can easily upload 15-second audio recordings, which are then used to generate personalized teaching assistants based on their unique voices. When explaining knowledge, this assistant uses intelligent algorithms to process teaching content in real time, accurately extracting key points and difficulties, and combining this with a lively tone and moderate speaking speed to generate easy-to-understand audio explanations, quickly responding to student needs. Simultaneously, it integrates with the platform's teaching resource database, linking and expanding knowledge points, and pushing audio explanations according to students' learning progress and comprehension abilities, achieving personalized teaching. In interactive Q&A sessions, the assistant analyzes student questions in real time, quickly providing accurate answers, and can also simulate online communication scenarios to guide student thinking, stimulate learning interest, and improve learning outcomes.

[0178] Please refer to Figure 2. In a second aspect of the present invention, a system for sound replication and low-latency streaming speech synthesis based on ultra-short samples is provided, comprising:

[0179] The audio sample input module is used to input a target speaker audio sample with a duration of less than 15 seconds, and to set style guidance parameters and acoustic environment modeling instructions;

[0180] The ultra-short sample sound replication module is used to process audio samples using a pre-trained sound replication model to generate personalized timbre data containing timbre, style and environmental features. The generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID.

[0181] A low-latency bidirectional streaming speech synthesis engine is used to determine the current timbre ID according to the preset requirements of the scenario. Preferably, the low-latency bidirectional streaming speech synthesis engine also includes a block-aware causal streaming generation submodule, which is used to generate subsequent response text through LLM based on the user's voice input, and dynamically generate and update real-time emotion control parameters based on the progress of the dialogue with the user and the analysis results. Based on the subsequent response text and the real-time emotion control parameters, the latest audio data block is generated based on the context of the processed text and the retrieved personalized timbre data, and continuously pushed to the player for playback to achieve continuous playback of the speech.

[0182] Furthermore, all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with authorization from the owner of the relevant device.

[0183] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for sound replication and low-latency streaming speech synthesis based on ultra-short samples, characterized in that, Includes the following steps: Users or system administrators can upload a target speaker audio sample of less than 15 seconds in length through the management interface, and set style guidance parameters and acoustic environment modeling instructions; the audio sample is processed using a pre-trained sound replication model to generate personalized timbre data containing timbre, style and environmental features, and the generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID; The current timbre ID is determined according to the preset requirements of the scenario. Based on this timbre ID, corresponding personalized timbre data is retrieved from the timbre repository. Acoustic features are synthesized based on the text content, personalized timbre data, and emotion control parameters. These acoustic features are converted into audio data blocks using a vocoder, and then the audio blocks are pushed to the player for playback. Based on the analysis results of the user's voice input feedback, subsequent response text is generated using LLM. Real-time emotion control parameters are dynamically generated and updated based on the progress of the dialogue with the user and the analysis results. The latest audio data blocks are generated based on the context of the processed text and the retrieved personalized timbre data, and are continuously pushed to the player for continuous voice playback. The acoustic features are scored and evaluated according to a voice expression effect evaluation function, which is as follows: ,in, This represents the speech performance score for input audio sample i. Indicates the speech rate deviation coefficient. This represents the volume deviation coefficient. Indicates the emotional intensity deviation coefficient. This represents the speech rate of the input audio sample i. This indicates the number of words spoken per unit of time. This represents the volume of the input audio sample i. This indicates the standard volume, with a value of 50 dB. The speech expression effect score indicates the intensity of emotion. When the score is less than a preset threshold, the target speaker's audio sample needs to be re-uploaded and the acoustic features regenerated. When the score is greater than the preset threshold, no action is required. The steps for updating the real-time emotion control parameters include: the low-latency bidirectional streaming speech synthesis engine receiving the text analysis results, emotion control parameters, and target acoustic environment simulation parameters of the user's speech input; dynamically adjusting the fundamental frequency of the speech when generating acoustic features through the vocoder model, including adjusting the pitch, loudness, or stress, and adjusting the speech rate and pauses; and generating speech that conforms to the specified prosody and emotional tendency through spectral characteristics.

2. The method for sound replication and low-latency streaming speech synthesis based on ultra-short samples according to claim 1, characterized in that, The style guidance parameters represent parameters that guide the replication of a specific speaking style. The style guidance parameters include preset style tags, which include standard customer service, dynamic marketing, and personalized reference audio. The acoustic environment modeling instructions are used to reproduce the background noise and channel characteristics in the sample, and to simulate and display the human voice effect.

3. The method for sound replication and low-latency streaming speech synthesis based on ultra-short samples according to claim 1, characterized in that, The personalized timbre data includes timbre, style, and environmental features, wherein the environmental features are those that simulate intelligent voice interaction.

4. The method for sound replication and low-latency streaming speech synthesis based on ultra-short samples according to claim 1, characterized in that, The timbre ID represents an identifier for personalized timbre data used in synthesis. The corresponding timbre embedding, style representation, and acoustic environment feature model are retrieved from the timbre repository using the timbre ID through a low-latency bidirectional streaming speech synthesis engine.

5. The method for sound replication and low-latency streaming speech synthesis based on ultra-short samples according to claim 1, characterized in that, The emotion control parameters are used to dynamically adjust the synthesis parameters of the generated speech. The emotion control parameters include speech rate control parameters, volume control parameters, emotion intensity parameters, and pause control commands. The speech rate control parameters are set at 50-200 words / minute, and the emotion intensity parameters are set at 0.8-1.5 times the baseline value.

6. A system for sound replication and low-latency streaming speech synthesis based on ultra-short samples, applied to a method for sound replication and low-latency streaming speech synthesis based on ultra-short samples as described in any one of claims 1-5, characterized in that, include: An audio sample input module is used to input a target speaker audio sample with a duration of less than 15 seconds and set style guidance parameters and acoustic environment modeling instructions. An ultra-short sample voice replication module is used to process the audio sample using a pre-trained voice replication model, generating personalized timbre data containing timbre, style, and environmental features. The generated personalized timbre data is stored in the system's timbre repository and associated with a unique timbre ID. A low-latency bidirectional streaming speech synthesis engine is used to determine the current timbre ID according to preset scene requirements, retrieve the corresponding personalized timbre data from the timbre repository based on the timbre ID, synthesize acoustic features based on text content, personalized timbre data, and emotion control parameters, convert the acoustic features into audio data blocks using a vocoder, and then push the audio blocks to a player for playback. The acoustic features are scored and evaluated according to a speech expression effect evaluation function, which is as follows: ,in, This represents the speech performance score for input audio sample i. Indicates the speech rate deviation coefficient. This represents the volume deviation coefficient. Indicates the emotional intensity deviation coefficient. This represents the speech rate of the input audio sample i. This indicates the number of words spoken per unit of time. This represents the volume of the input audio sample i. This indicates the standard volume, with a value of 50 dB. The speech expression effect score indicates the intensity of emotion. When the score is less than a preset threshold, the target speaker's audio sample needs to be re-uploaded and the acoustic features regenerated. When the score is greater than the preset threshold, no action is required. The steps for updating the real-time emotion control parameters include: the low-latency bidirectional streaming speech synthesis engine receiving the text analysis results, emotion control parameters, and target acoustic environment simulation parameters of the user's speech input; dynamically adjusting the fundamental frequency of the speech when generating acoustic features through the vocoder model, including adjusting the pitch, loudness, or stress, and adjusting the speech rate and pauses; and generating speech that conforms to the specified prosody and emotional tendency through spectral characteristics.

7. The sound replication and low-latency streaming speech synthesis system based on ultra-short samples according to claim 6, characterized in that, The low-latency bidirectional streaming speech synthesis engine also includes a block-aware causal streaming generation submodule, which is used to generate subsequent response text through LLM based on the user's voice input, dynamically generate and update real-time emotion control parameters based on the progress of the dialogue with the user and the analysis results, generate the latest audio data block based on the context of the processed text and the retrieved personalized timbre data, and continuously push it to the player for playback to achieve continuous playback of the speech.

Citation Information

Patent Citations

  • Speech synthesis method and related equipment

    CN108962217A