Speech Denoising via Power Spectrum Iteration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for processing noisy speech struggle to effectively reduce environmental noise, leading to degraded speech quality, particularly due to limitations in tracing changes in speech and noise over time.

Innovation Solution

A method and apparatus that utilize a power spectrum iteration factor and moving average power spectrum to enhance signal-to-noise ratio (SNR) by obtaining noise from quiet periods in noisy speech, calculating a power spectrum iteration factor, and determining a moving average power spectrum, which reduces noise and improves speech quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If a short-term spectral estimation algorithm is adopted to reduce environmental noise, then the noise reduction capability is improved, but the speech quality is degraded due to inability to trace changes in speech and noise over time

Engineering Contradiction:
Improveenvironmental noiseVSAvoidspeech quality
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent applies dynamics by making the power spectrum estimation adaptive over time through iterative updating. The algorithm transitions from static short-term estimation to dynamic long-term tracking by continuously updating the power spectrum using a moving average approach with iteration factor, enabling the system to adapt to changing speech and noise characteristics while maintaining speech quality

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements feedback by using the previously estimated power spectrum to inform the current estimation. The iterative update mechanism feeds back the (m-1)th frame power spectrum information into the mth frame estimation process, creating a closed-loop system that continuously refines noise estimation and improves speech quality over time

Inventive Principle:
Principle #23Feedback

2Device complexity

If conventional spectral estimation algorithms are used, then the processing complexity is reduced, but the signal-to-noise ratio (SNR) is insufficient due to inability to accurately trace speech and noise changes

Engineering Contradiction:
Improveprocessing complexityVSAvoidsignal-to-noise ratio
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies partial action by selectively updating only the necessary components of the power spectrum estimation using a moving average approach. Instead of complete re-estimation, the algorithm performs partial updates using the iteration factor, reducing computational complexity while maintaining improved SNR through cumulative learning from previous frames

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements preliminary action by pre-calculating and storing the power spectrum from previous frames. This preliminary information is then reused in subsequent estimations through the iterative update process, avoiding redundant computations and reducing overall processing complexity while improving measurement precision

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9978391B2Method, apparatus and server for processing noisy speech
Publication Date: 2018.05.22 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US9978391B2 patent drawing
  • US9978391B2 patent drawing
  • US9978391B2 patent drawing

AI summary

According to an embodiment, a power spectrum iteration factor is determined according to a noisy speech and a background noise, and a moving average power spectrum of the speech is obtained according to the power spectrum iteration factor. A server is able to trace the noisy speech according to the power spectrum iteration factor.