Speech Processing Voice Characteristic Compensation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing technologies fail to address the unnatural sound and reduced intelligibility caused by speakers adjusting their voice in noisy environments, even after noise suppression, as these techniques do not account for the speaker's adaptive changes in voice characteristics.

Innovation Solution

A method and apparatus that detect input voice characteristics in noise-suppressed voice signals and compare them to reference voice characteristics from a noise-free environment, modifying the signal when differences exceed a threshold to create a more natural-sounding voice signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If noise suppression is applied to remove or reduce background noise, then the background noise level is reduced, but the speech becomes unnatural and intelligibility is reduced due to the speaker's voice adaptation to background noise

Engineering Contradiction:
Improvebackground noise levelVSAvoidspeech naturalness and intelligibility
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The system detects voice characteristics during noise-free periods before the noise suppression process, storing these as reference characteristics. This preliminary detection allows the subsequent noise-suppressed speech to be corrected by comparing against these pre-established reference characteristics, thereby maintaining naturalness while removing noise.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors voice characteristics in the noise-suppressed signal and compares them against reference characteristics. When deviations are detected (indicating unnatural speech), the system applies corrective adjustments. This closed-loop feedback mechanism ensures that speech naturalness is maintained throughout the noise suppression process.

Inventive Principle:
Principle #23Feedback

2Object-affected harmful factors

If better noise suppression techniques are used, then the background noise is more effectively removed, but the unnatural sound and reduced intelligibility become more noticeable and disturbing

Engineering Contradiction:
Improvebackground noise removal effectivenessVSAvoidunnatural sound and reduced intelligibility
Core Design Contradiction:
Object-affected harmful factorsVSObject-generated harmful factors

Solution Approach 1:

The system implements continuous monitoring of voice characteristics in the noise-suppressed signal, comparing them against reference characteristics obtained during noise-free periods. When deviations indicating unnatural speech are detected, corrective adjustments are applied. This feedback mechanism ensures that higher-quality noise suppression does not lead to more noticeable unnatural artifacts.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts speech parameters (such as pitch, energy, or spectral characteristics) based on the detected deviation from reference characteristics. By modifying these parameters in response to detected unnaturalness, the system maintains natural speech quality even when aggressive noise suppression is applied.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9530427B2Speech processing
Publication Date: 2016.12.27 NOKIA TECHNOLOGIES OY
  • US9530427B2 patent drawing
  • US9530427B2 patent drawing
  • US9530427B2 patent drawing

AI summary

A technique for enhancing speech signal captured in a noisy environment is provided. According an example embodiment, the technique comprises obtaining a current time frame of a noise-suppressed voice signal, derived on basis of a current time frame of a source audio signal comprising a source voice signal, detecting input voice characteristics for the current time frame of noise-suppressed voice signal, obtaining reference voice characteristics for said current time frame, said reference voice characteristics being descriptive of the source voice signal in noise-free or low-noise environment, and creating a current time frame of a modified voice signal by modifying said current time frame of the noise-suppressed voice signal in response to a difference between the detected input voice characteristic and the reference voice characteristics exceeding a predetermined threshold.