Speech synthesis system capable of realizing controllable background removal and retention based on environment perception

By introducing an environment perception mechanism and a controllable mask speech prediction strategy in the speech synthesis system, the existing system's difficulties in handling noise and retaining acoustic background are solved, and a higher quality and more flexible background processing effect is achieved.

CN119943028AActive Publication Date: 2025-05-06SHANGHAI JIAOTONG UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510121489.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

When existing zero-sample text-to-speech systems deal with low-quality prompt speech or contaminated by environmental noise, reverb or interference speakers, it is difficult to simultaneously remove unnecessary background sounds and retain important acoustic background information, resulting in lower speech quality and loss of background information.

Method used

A speech synthesis system for controlled background removal and retention of environment-aware environment-aware controllable background removal and retention is proposed. Through a time predictor, acoustic model and a dual prompt speech encoder, combined with a stream matching algorithm and a controllable mask speech prediction training strategy, precise control of background information can be achieved, and environmental information can be flexibly removed or retained in different scenarios.

Benefits of technology

It significantly improves the system's robustness in a noisy environment, ensures that the output speech maintains high quality under different background noise conditions, improves speech clarity and nature, and maintains tone matching and background consistency when the acoustic background is required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943028A_ABST
    Figure CN119943028A_ABST
Patent Text Reader

Abstract

The invention discloses an environment-aware controllable background removal and retention speech synthesis system, relates to the field of speech, and provides a speech synthesis system capable of perceiving an acoustic environment according to noisy prompt speech so as to carry out controllable background removal and retention. According to the method, control signals related to texts, prompt voices and tasks are used as input, a time length predictor, an acoustic model and a double prompt voice encoder are included, a controllable mask voice prediction training strategy is further provided on the basis of a stream matching algorithm on the basis of a training strategy, and controllable background removal and retention are achieved by providing noisy prompt voices. According to the invention, the robustness and controllability of the system for processing the prompt voice with noise, reverberation and interference to the speaker are improved, the removal and retention of the background contained in the prompt voice can be effectively controlled when the voice is generated, and the higher generated voice quality and the more similar acoustic background are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech, and in particular to an environment-aware speech synthesis system capable of controllable background removal and retention. Background Art

[0002] Text-to-Speech (TTS) technology is an assistive technology that converts written text into spoken language. Zero-shot Text-to-Speech (ZSTTS) technology aims to generate natural speech from an unseen speaker conditioned on a prompt speech. Although recent advances in deep learning techniques have improved the quality of synthesized speech, these zero-shot text-to-speech systems still have difficulty processing low-quality prompt speech or speech that is contaminated by environmental noise, reverberation, or interfering speakers.

[0003] Acoustic background plays a vital role in natural speech, providing environmental context that helps listeners understand their surroundings, but a strong background can make it difficult for listeners to hear what is being said. How to deal with these acoustic backgrounds depends on the specific situation: while in some cases it is necessary to remove the background to ensure the clarity of the speech, sometimes retaining the background is crucial to maintain the contextual integrity of the speech.

[0004] There are two difficulties in processing these noisy prompt speech. First, the zero-shot text-to-speech system should remove the unnecessary background sounds in the prompt speech to generate clean speech with the timbre of the target speaker. Second, when the acoustic background information in the prompt speech is important, the zero-shot text-to-speech system should not only imitate the timbre of the target speaker, but also generate speech with a consistent acoustic background. The present invention defines this dual requirement as a background removal and preservation task. The former aims to generate clean speech with the same timbre as the target speaker in the prompt speech, while the latter requires the generation of speech with a consistent speaker timbre and similar acoustic background as the prompt speech.

[0005] Most existing technical research focuses on the background removal task, which aims to enhance the noise robustness of zero-shot text-to-speech systems to noisy prompt speech. The most direct approach is to apply speech enhancement (SE) technology to the noisy prompt speech, and then input the enhanced prompt speech into the text-to-speech system. However, even the most advanced speech enhancement models inevitably introduce processing artifacts, which reduce the quality of speech generated by the text-to-speech system. Existing research mainly focuses on enhancing the robustness of text-to-speech models to generate clean speech from noisy prompt speech. Recently, some studies have improved the noise robustness of stream matching models through the Masked Speech Denoising (MSD) strategy, which uses the contextual prompt speech x when training the acoustic model. ctx Noise is added to the target speech for enhancement and the original clean speech target is predicted. However, while this approach is effective in scenarios involving additional noise, it loses the background preservation capability inherent in traditional stream matching models, limiting its application in scenarios where the acoustic environment of the prompt needs to be preserved. In addition, existing systems reduce the timbre similarity between the synthesized speech and the prompt speech compared to the prompt speech, especially when dealing with prompt speech with reverberation or interfering speakers.

[0006] In contrast, less attention has been paid to the task of background preservation. Some existing systems, such as VoiceLDM and Audiobox, control the speech synthesis of the acoustic background through text descriptions, but many acoustic backgrounds cannot be accurately and specifically described in words. EATTS achieves background preservation for prompt speech with reverberation, but the system is only effective in scenarios with reverberation, a relatively stable acoustic feature. Recent zero-shot text-to-speech models such as VoiceBox and VALL-E, respectively, have made use of stream matching models and language models to initially achieve the ability to preserve background, demonstrating the potential of generating speech that reflects the acoustic environment of the input prompt speech. However, systems such as Voicebox and VALL-E lack a control mechanism, resulting in unstable speech synthesis results under different prompt speech. For the VoiceBox system, the timbre and background in the prompt speech are intertwined, resulting in the acoustic background in the generated speech only existing when there is a speech segment, and disappearing in the silent part. In addition, these existing systems cannot remove or preserve environmental information at the same time, because background removal and preservation tasks are inherently conflicting and may affect their effectiveness.

[0007] Therefore, those skilled in the art have devoted themselves to developing a system that can remove or retain environmental information at the same time. Summary of the invention

[0008] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to develop a system that can remove or retain environmental information at the same time.

[0009] To achieve the above-mentioned purpose, the present invention provides an environment-aware speech synthesis system with controllable background removal and retention, characterized in that the present invention proposes a speech synthesis system that can perceive the acoustic environment according to noisy prompt speech, thereby performing controllable background removal and retention, taking text, prompt speech and task-related control signals as input, comprising a duration predictor, an acoustic model and a dual prompt speech encoder, and in terms of training strategy, based on a stream matching algorithm, further proposing a controllable masked speech prediction training strategy, and achieving controllable background removal and retention by providing noisy prompt speech.

[0010] Furthermore, the text is used to control the content of the synthesized speech, the prompt speech is used to control the speaker timbre and acoustic background of the synthesized speech, and the task-related control signal is used to control the synthesized speech to be clean speech (or the synthesized speech has an acoustic background similar to the prompt speech.

[0011] Furthermore, the acoustic model models the conditional distribution qxz,x ctx , according to the phoneme sequence z and the prompt x ctx Generate a Mel-spectrogram x, the acoustic model is used as a vector field estimator, and at each time stamp t∈0,1, a randomly selected mask part is predicted At the same time, the visible part x ctx =x⊙1mask is regarded as a prompt; the training goal of the acoustic model is As described above, the output of network θ is v t w,x ctx ,z;θ, the flow of step t is denoted by w, while x0 is Gaussian noise sampled from a normal distribution, σ min is a hyperparameter that controls the flow matching bias.

[0012] Furthermore, the duration predictor and the acoustic model both use a classifier-free guidance strategy CFG during training and testing to balance pattern coverage and sample fidelity; during the training of the acoustic model, the acoustic prompt x ctx and phoneme sequence z with probability p uncond is randomly discarded; during the inference process, the acoustic model first samples a Gaussian noise x0 from the normal distribution p0 and uses the ODE solver to evaluate the flow w; the modified vector field under the CFG strategy becomes the formula In Instead of v t w,x ctx,z;θ, where α is a hyperparameter that controls the strength of the guidance.

[0013] Furthermore, the duration predictor models the conditional distribution qyl,y ctx , where y is the predicted duration sequence of each phoneme, given the input phoneme index sequence l and the prompt y ctx Therefore, the training objectives and reasoning process of the duration predictor are respectively and formula

[0015] Furthermore, the acoustic model proposes a new controllable masked speech prediction strategy, which unifies the background removal and retention tasks into a mask prediction problem. By introducing task-related control signals, the system is guided to accurately switch between background removal and retention, thereby achieving dual-objective control.

[0016] Furthermore, in the controllable masked speech prediction strategy, the acoustic model is obtained by probability P N , P R and P IS Model different types of noise and simulate enhanced noisy speech x by adding environmental noise, reverberation, and interfering speech to clean training samples x. aug ; The probability of not adding background, that is, x aug = x, denoted as P C , part of the speech is randomly masked and needs to be predicted, recorded as before and after enhancement and The unmasked part is used as a hint and is recorded as A multi-task learning loss is designed based on the conventional flow matching loss, as shown in the formula:

[0017]

[0018] The acoustic model The output is v t w,x ctx ,z,c;θ, the flow of step t is denoted by w, x0 is Gaussian noise sampled from a normal distribution, σ min is a hyperparameter that controls the flow matching bias.

[0019] Furthermore, two identical encoders are introduced into the acoustic model, which process the noisy prompt speech independently.

[0020] Furthermore, a control mechanism guided by a binary control signal c decides which speaker encoder branch to activate; this architecture design can minimize the model interference between the removal and retention tasks while further improving the performance of background retention.

[0021] Furthermore, in the inference process, given a control signal c for a given task, the inference process first generates a phoneme sequence z from a given transcription by the duration predictor; then, the noise cue x ctx , phoneme sequence z, Gaussian noise x0 and control signal c are input into the acoustic model; the modified CFG strategy is adopted, and the acoustic model predicts the final Mel spectrogram through the ordinary differential equation solver and guided by the defined vector field, as shown in the formula Finally, the HiFiGAN vocoder converts the mel-spectrogram into a waveform.

[0022] The present invention has the following technical effects:

[0023] Through the dual background processing mechanism, the robustness of the system in a noisy environment is significantly enhanced. Compared with the traditional system, the background removal module of the present invention can filter out unnecessary background sounds more accurately, ensuring that the output speech maintains a high quality under various background noise conditions. Preliminary experimental data show that the speech clarity and naturalness of the system in different noise scenarios are significantly improved, especially in scenarios with complex acoustic environments.

[0024] For scenarios where the consistency of the acoustic background of the prompt speech needs to be maintained, the present invention achieves high standards in both timbre matching and background consistency, further improving the realism and presence of the generated speech. Compared with other methods, the timbre similarity index (such as timbre matching score) and background consistency score of the present invention in the background preservation task are better than the existing technology, and have industrial application value.

[0025] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a system overview diagram of a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0028] In the drawings, components with the same structure are indicated by the same numerical reference numerals, and components with similar structures or functions are indicated by similar numerical reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of the components is appropriately exaggerated in some places in the drawings.

[0029] The present invention proposes a speech synthesis system that can perceive the acoustic environment based on noisy prompt speech, thereby performing controllable background removal and retention. The system takes text, prompt speech and task-related control signals as input, wherein the text is used to control the content of the synthesized speech, the prompt speech is used to control the speaker's timbre and acoustic background of the synthesized speech, and the task-related control signal is used to control the synthesized speech to be clean speech (without acoustic background) or the synthesized speech has an acoustic background similar to the prompt speech. The entire system is structurally as follows: Figure 1 , including a duration predictor, an acoustic model and a dual prompt speech encoder. In terms of training strategy, the present invention further proposes a controllable masked speech prediction training strategy based on a stream matching algorithm, and realizes controllable background removal and retention by providing noisy prompt speech.

[0030] 1. Duration Predictor and Acoustic Model Based on Stream Matching

[0031] Flow Matching (FM) is a simulation-free training CNF method. The relationship between the vector field and the flow is defined by an ordinary differential equation (ODE). The text-to-speech system of the present invention consists of an acoustic model based on flow matching and a duration predictor.

[0032] The acoustic model models the conditional distribution qxz,x ctx Specifically, it is based on the phoneme sequence z and the prompt x ctx Generate a Mel spectrogram x. The model acts as a vector field estimator and predicts a randomly selected mask part at each time stamp t∈0,1 At the same time, the visible part x ctx =x⊙1-mask is regarded as a prompt. The training objective of the acoustic model is as follows: As described above, the output of network θ is v t w,x ctx ,z;θ, the flow of t steps is denoted by w, and x0 is the Gaussian noise sampled from a normal distribution. min is a hyperparameter that controls the flow matching bias.

[0033] The duration predictor and acoustic model both use the Classifier Free Guidance (CFG) strategy during training and testing to balance pattern coverage and sample fidelity. During the training of the acoustic model, the acoustic cue x ctx and phoneme sequence z with probability p uncond are randomly discarded. During inference, the model first samples a Gaussian noise x0 from a normal distribution p0 and evaluates the flow w using an ODE solver. The modified vector field under the CFG strategy becomes the formula In Instead of v t w,x ctx ,z;θ, where α is a hyperparameter that controls the strength of the guidance.

[0034] The duration predictor is similar to the acoustic model, which models the conditional distribution qyl,y ctx , where y is the predicted duration sequence of each phoneme, given the input phoneme index sequence l and the prompt y ctx Therefore, the training objective and inference process of the duration predictor are respectively and formula

[0035]

[0036] 2. Controllable Masked Speech Prediction Training Strategy

[0037] Although the masked speech denoising strategy effectively estimates clean audio from noise cues, it cannot selectively preserve the acoustic environment. To address this problem, this paper proposes a new controllable masked speech prediction (CMSP) strategy for acoustic models, unifying the background removal and preservation tasks into a mask prediction problem. This enables the model to perform dual mask prediction within the same training framework, ensuring that the model can flexibly adapt to different application scenarios while maintaining consistent speaker characteristics.

[0038] Specifically, by probability P N , P R and P IS Different types of noise are modeled. The present invention simulates the enhanced noisy speech x by adding environmental noise, reverberation and interfering speech to the clean training sample x respectively. aug The probability of not adding background, that is, x aug = x, denoted as P C A part of the speech is randomly masked and needs to be predicted, which are recorded as before and after enhancement. and The unmasked part is used as a hint and is recorded as The present invention designs a multi-task learning loss based on the conventional flow matching loss, as shown in Formula 5. The output of the acoustic model θ is v t w,x ctx ,z,c;θ, the flow of t steps is denoted by w, and x0 is the Gaussian noise sampled from a normal distribution. min is a hyperparameter of the control flow matching deviation. Other parameters are consistent with Formula 1.

[0039] By introducing the binary control signal cc0, c1, the model is guided to focus on the removal (corresponding to the control signal c0) or retention (corresponding to the control signal c1) of background elements, thereby achieving precise control of the acoustic environment of the generated speech.

[0040]

[0041] 3. Dual Prompt Speech Encoder

[0042] Previous research has shown that it is very important to keep the synthesized speech in the speaker's voice when processing noisy cues. However, other systems, such as VoiceBox, may limit their ability to accurately extract speaker information in complex acoustic conditions by simply splicing noise cues into the model input. To address this limitation, the present invention proposes a dual-cue speech encoder specifically designed to better handle noisy cues. Figure 1 As shown in the figure, in addition to splicing the noise prompt into the input, the present invention introduces two identical encoders into the model, which process the noisy prompt speech independently, and decide which speaker encoder branch to activate through a control mechanism guided by a binary control signal c. This architecture design can minimize the model interference between the removal and retention tasks, while further improving the performance of background retention.

[0043] During inference, given a task-specific control signal c, the inference process first generates a phoneme sequence z from a given transcription by the duration predictor. Then, the noise cue x ctx , phoneme sequence z, Gaussian noise x0 and control signal c are input into the acoustic model. Correspondingly, the present invention adopts a modified CFG strategy. The acoustic model predicts the final Mel spectrogram through the ODE solver and guided by the defined vector field, as shown in the formula Finally, the HiFiGAN vocoder converts the mel-spectrogram into a waveform.

[0044] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. An environment-aware speech synthesis system with controllable background removal and preservation, characterized in that: The present invention proposes a speech synthesis system that can perceive the acoustic environment according to noisy prompt speech, thereby performing controllable background removal and retention. The system takes text, prompt speech and task-related control signals as input, and includes a duration predictor, an acoustic model and a dual prompt speech encoder. In terms of training strategy, based on a stream matching algorithm, a controllable masked speech prediction training strategy is further proposed, which achieves controllable background removal and retention by providing noisy prompt speech.

2. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 1, characterized in that: The text is used to control the content of the synthesized speech, the prompt speech is used to control the speaker timbre and acoustic background of the synthesized speech, and the task-related control signal is used to control the synthesized speech to be clean speech (or the synthesized speech has an acoustic background similar to the prompt speech.

3. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 2, characterized in that: The acoustic model models the conditional distribution q(x|z,x ctx ), according to the phoneme sequence z and the prompt x ctx Generate a Mel-spectrogram x, the acoustic model is used as a vector field estimator, and at each timestamp t∈[0,1], a randomly selected mask part is predicted At the same time, the visible part x ctx =x⊙(1-mask) is regarded as a prompt; the training goal of the acoustic model is As described above, the output of network θ is v t (w,x ctx ,z;θ), the flow of t steps is denoted by w, and x0 is the Gaussian noise sampled from a normal distribution, σ min is a hyperparameter that controls the flow matching bias.

4. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 3, characterized in that: The duration predictor and the acoustic model both use the classifier-free guidance strategy CFG during training and testing to balance pattern coverage and sample fidelity; during the training of the acoustic model, the acoustic prompt x ctx and phoneme sequence z with probability p uncond is randomly discarded; during the inference process, the acoustic model first samples a Gaussian noise x0 from the normal distribution p0 and uses the ODE solver to evaluate the flow w; the modified vector field under the CFG strategy becomes the formula In Instead of v t (w,x ctx ,z;θ), where α is a hyperparameter that controls the strength of the guidance.

5. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 4, characterized in that: The duration predictor models the conditional distribution q(y|l,y ctx ), where y is the predicted duration sequence of each phoneme, given the input phoneme index sequence l and the prompt y ctx Therefore, the training objectives and reasoning process of the duration predictor are respectively and formula 6. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 5, characterized in that: The acoustic model proposes a new controllable masked speech prediction strategy, which unifies the background removal and retention tasks into a mask prediction problem. By introducing task-related control signals, the system is guided to accurately switch between background removal and retention, thereby achieving dual-objective control.

7. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 6, characterized in that: In the controllable masked speech prediction strategy, the acoustic model is expressed by probability P N , P R and P IS Model different types of noise and simulate enhanced noisy speech x by adding environmental noise, reverberation, and interfering speech to clean training samples x aug ; The probability of not adding background, that is, x aug =x, denoted as P C , part of the speech is randomly masked and needs to be predicted, recorded as before and after enhancement and The unmasked part is used as a hint and is recorded as A multi-task learning loss is designed based on the conventional flow matching loss, as shown in the formula: The acoustic model The output is v t (w,x ctx ,z,c;θ), the flow of step t is denoted by w, x0 is Gaussian noise sampled from a normal distribution, σ min is a hyperparameter that controls the flow matching bias.

8. The environment-aware controllable background removal and preservation speech synthesis system as claimed in claim 7, characterized in that: Two identical encoders are introduced into the acoustic model, which process the noisy prompt speech independently.

9. The environment-aware controllable background removal and retention speech synthesis system as claimed in claim 8, characterized in that: A control mechanism guided by a binary control signal c decides which speaker encoder branch to activate; this architecture design can minimize the model interference between the removal and preservation tasks while further improving the performance of background preservation.

10. The environment-aware controllable background removal and preservation speech synthesis system as claimed in claim 9, characterized in that: In the inference process, given a control signal c for a given task, the inference process first generates a phoneme sequence z from the given transcription by the duration predictor; then, the noise cue x ctx , phoneme sequence z, Gaussian noise x0 and control signal c are input into the acoustic model; the modified CFG strategy is adopted, and the acoustic model predicts the final Mel spectrogram through the ordinary differential equation solver and guided by the defined vector field, as shown in the formula Finally, the HiFiGAN vocoder converts the mel-spectrogram into a waveform.

Citation Information

Patent Citations

  • Timbre-controllable video sound synthesis model and construction method, device and application thereof

    CN115359777A

  • Speech synthesis method and device, electronic equipment and readable storage medium

    CN117854470A

  • Software and hardware combined music perception product for hearing-impaired children based on posture recognition technology

    CN118116553A

  • Speaker generation method based on text expression driving

    CN118865941A

  • Method for voice conversation and system therefor

    JP2001142483A