Noise Compensation in Speaker-Adaptive Acoustic Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-to-speech systems struggle to maintain voice quality when adapting acoustic models from noisy speech samples, as existing methods fail to effectively separate noise effects from speaker characteristics, leading to poor matching between input and output voices.
Innovation Solution
A method and apparatus that utilize a deep neural network to map noisy speech parameters to clean speech parameters, employing noise characterization and adaptation to isolate speaker transforms, ensuring high-quality voice synthesis regardless of noise levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If acoustic models are adapted from noisy speech samples, then the system can work in noisy environments, but the voice quality and speaker matching deteriorate
Solution Approach 1:
The patent segments the speech signal into clean speech components and noise components separately. By dividing the noisy speech into distinct elements (clean speech signal and noise signal), the system can process each independently and recombine them to achieve both noise environment adaptability and high voice quality.
Solution Approach 2:
The patent extracts the noise characteristics from the noisy speech samples and separates them from the clean speech signal. This extraction allows the system to remove noise effects while preserving speaker characteristics, resolving the contradiction between working in noisy environments and maintaining voice quality.
2Ease of operation
If existing adaptation methods are used on noisy speech, then processing can proceed without special handling, but speaker characteristics become corrupted by noise effects
Solution Approach 1:
The patent introduces an intermediary noise compensation module that acts as a mediator between the noisy speech input and the acoustic model adaptation. This intermediary processes the noisy signal to extract and remove noise effects before the clean speech signal is used for adaptation, ensuring reliable speaker matching while maintaining operational simplicity.
3Measurement precision
If noise-free samples are required for high-quality adaptation, then voice synthesis quality improves, but the system cannot handle noisy input environments
Solution Approach 1:
The patent converts the harmful noise in the input speech into a beneficial feature by using noise characterization to identify and remove noise effects. The noise that would normally degrade quality is instead used as information to guide the compensation process, allowing the system to achieve high voice synthesis quality directly from noisy inputs without requiring separate noise-free samples.
Data Source
AI summary
An acoustic model is adapted, relating acoustic units to speech vectors. The acoustic model comprises a set of acoustic model parameters related to a given speech factor. The acoustic model parameters enable the acoustic model to output speech vectors with different values of the speech factor. The method comprises inputting a sample of speech which is corrupted by noise; determining values of the set of acoustic model parameters which enable the acoustic model to output speech with said first value of the speech factor; and employing said determined values of the set of speech factor parameters in said acoustic model. The acoustic model parameters are obtained by obtaining corrupted speech factor parameters using the sample of speech, and mapping the corrupted speech factor parameters to clean acoustic model parameters using noise characterization paramaters characterizing the noise.


