Speech Signal Separation Using Adaptive Voice Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing methods struggle to effectively separate speech signals from background noise in noisy acoustic environments, particularly in real-world scenarios with multiple noise sources and reverberation, due to limitations in adaptability and computational complexity.

Innovation Solution

A robust speech separation process utilizing a voice activity detector and independent component analysis (ICA) with adaptive learning and post-processing stages, which adjusts signal separation based on detected voice activity and environmental conditions, and includes mechanisms for wind noise management and scaling to optimize signal quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If simple filtering processes with predetermined noise characteristics are used, then processing speed is fast enough for real-time, but speech signal degradation occurs and adaptability to different environments is poor

Engineering Contradiction:
Improveprocessing speedVSAvoidspeech signal quality
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent implements dynamic adaptation by continuously updating noise characteristics based on environmental conditions. The system transitions from static predetermined filtering to dynamic filtering that adjusts to changing acoustic environments, thereby maintaining speech signal quality while enabling real-time processing across varied conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes filtering parameters adaptively based on detected noise characteristics. By monitoring the acoustic environment and adjusting filter parameters dynamically, the system achieves both real-time processing speed and high speech signal quality without degradation.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If predetermined noise filtering assumptions are applied, then processing is simple and fast, but noise removal is over-inclusive or under-inclusive causing speech degradation

Engineering Contradiction:
Improveprocessing complexityVSAvoidnoise identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system performs self-adjustment by automatically learning and adapting to the specific acoustic environment. Instead of relying on predetermined assumptions, the filtering system serves itself by continuously optimizing its parameters based on real-time environmental feedback, achieving precise noise identification without excessive complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where the system monitors the results of noise filtering and adjusts its parameters accordingly. This closed-loop approach ensures accurate noise identification by continuously refining the filtering process based on actual performance, avoiding both over-inclusive and under-inclusive removal.

Inventive Principle:
Principle #23Feedback

3Reliability

If robust speech separation is achieved through adaptive learning, then speech quality is high, but computational power requirements increase

Engineering Contradiction:
Improvespeech signal qualityVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial adaptation by focusing computational resources on the most critical aspects of speech separation. Rather than performing exhaustive adaptive learning on all parameters, the system selectively adapts only the necessary components, achieving high speech quality while reducing computational energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary processing steps that prepare the signal for more efficient adaptive learning. By pre-processing the input signal to extract key features and reduce dimensionality beforehand, the subsequent adaptive learning requires less computational power while maintaining high speech signal quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7464029B2Robust separation of speech signals in a noisy environment
Publication Date: 2008.12.09 QUALCOMM INC
  • US7464029B2 patent drawing
  • US7464029B2 patent drawing
  • US7464029B2 patent drawing

AI summary

A method for improving the quality of a speech signal extracted from a noisy acoustic environment is provided. In one approach, a signal separation process is associated with a voice activity detector. The voice activity detector is a two-channel detector, which enables a particularly robust and accurate detection of voice activity. When speech is detected, the voice activity detector generates a control signal. The control signal is used to activate, adjust, or control signal separation processes or post-processing operations to improve the quality of the resulting speech signal. In another approach, a signal separation process is provided as a learning stage and an output stage. The learning stage aggressively adjusts to current acoustic conditions, and passes coefficients to the output stage. The output stage adapts more slowly, and generates a speech-content signal and a noise dominant signal. When the learning stage becomes unstable, only the learning stage is reset, allowing the output stage to continue outputting a high quality speech signal.