Noise Compensation in Speaker-Adaptive Acoustic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech systems struggle to maintain voice quality when adapting acoustic models from noisy speech samples, as existing methods fail to effectively separate noise effects from speaker characteristics, leading to poor matching between input and output voices.

Innovation Solution

A method and apparatus that utilize a deep neural network to map noisy speech parameters to clean speech parameters, employing noise characterization and adaptation to isolate speaker transforms, ensuring high-quality voice synthesis regardless of noise levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If acoustic models are adapted from noisy speech samples, then the system can work in noisy environments, but the voice quality and speaker matching deteriorate

Engineering Contradiction:
Improvenoise environment adaptabilityVSAvoidvoice quality
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech signal into clean speech components and noise components separately. By dividing the noisy speech into distinct elements (clean speech signal and noise signal), the system can process each independently and recombine them to achieve both noise environment adaptability and high voice quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the noise characteristics from the noisy speech samples and separates them from the clean speech signal. This extraction allows the system to remove noise effects while preserving speaker characteristics, resolving the contradiction between working in noisy environments and maintaining voice quality.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of operation

If existing adaptation methods are used on noisy speech, then processing can proceed without special handling, but speaker characteristics become corrupted by noise effects

Engineering Contradiction:
Improveprocessing simplicityVSAvoidspeaker matching accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces an intermediary noise compensation module that acts as a mediator between the noisy speech input and the acoustic model adaptation. This intermediary processes the noisy signal to extract and remove noise effects before the clean speech signal is used for adaptation, ensuring reliable speaker matching while maintaining operational simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If noise-free samples are required for high-quality adaptation, then voice synthesis quality improves, but the system cannot handle noisy input environments

Engineering Contradiction:
Improvevoice synthesis qualityVSAvoidnoise environment handling
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent converts the harmful noise in the input speech into a beneficial feature by using noise characterization to identify and remove noise effects. The noise that would normally degrade quality is instead used as information to guide the compensation process, allowing the system to achieve high voice synthesis quality directly from noisy inputs without requiring separate noise-free samples.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS10373604B2Noise compensation in speaker-adaptive systems
Publication Date: 2019.08.06 KK TOSHIBA
  • US10373604B2 patent drawing
  • US10373604B2 patent drawing
  • US10373604B2 patent drawing

AI summary

An acoustic model is adapted, relating acoustic units to speech vectors. The acoustic model comprises a set of acoustic model parameters related to a given speech factor. The acoustic model parameters enable the acoustic model to output speech vectors with different values of the speech factor. The method comprises inputting a sample of speech which is corrupted by noise; determining values of the set of acoustic model parameters which enable the acoustic model to output speech with said first value of the speech factor; and employing said determined values of the set of speech factor parameters in said acoustic model. The acoustic model parameters are obtained by obtaining corrupted speech factor parameters using the sample of speech, and mapping the corrupted speech factor parameters to clean acoustic model parameters using noise characterization paramaters characterizing the noise.