Speech Signal Processing Using Directional Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement systems face challenges in extracting a desired speech signal from noisy environments due to large noise interference and signal leakage, especially when the number of sound sources exceeds the number of microphones, leading to poor stability and quality of the separated speech.

Innovation Solution

The method involves acquiring sound source position information and using it to suppress noise signals from the sound source direction, obtaining a noise reference signal, and then removing residual noise from the speech signal to enhance the speech quality, employing techniques like adaptive filtering and blind source separation with direction constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech enhancement systems process noisy speech signals in environments with multiple sound sources, then speech quality can be improved, but noise interference and signal leakage increase when the number of sound sources exceeds the number of microphones

Engineering Contradiction:
Improvespeech qualityVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the noisy speech signal into multiple components by applying blind source separation with direction constraints. The microphone array captures signals from multiple directions, and the system separates these signals by estimating the direction of arrival (DOA) of different sound sources, effectively segmenting the mixed signal into individual source components including speech and noise.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by implementing direction-specific processing for different spatial locations. The system estimates DOA for each sound source and applies targeted enhancement and noise suppression filters based on the specific direction and characteristics of each source, rather than uniform processing across all directions.

Inventive Principle:
Principle #3Local quality

2Device complexity

If conventional speech enhancement methods are used, then processing can be simplified, but stability and quality of separated speech deteriorate when sound sources exceed microphones

Engineering Contradiction:
Improveprocessing complexityVSAvoidstability of separated speech
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent performs preliminary action by first estimating the direction of arrival (DOA) of sound sources before applying blind source separation. This preliminary directional information is used to constrain the separation process, improving stability. The system also pre-processes the microphone signals to estimate the covariance matrix and other statistical parameters needed for robust separation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the estimated DOA and separation results are continuously refined. The system uses the output of the separation process to update the DOA estimates and adjust the separation filters iteratively, improving the stability and accuracy of the separated speech signals through feedback loops.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11817112B2Method, device, computer readable storage medium and electronic apparatus for speech signal processing
Publication Date: 2023.11.14 BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
  • US11817112B2 patent drawing
  • US11817112B2 patent drawing
  • US11817112B2 patent drawing

AI summary

A method, a device, a computer readable medium and an electronic apparatus for speech signal processing are disclosed. The method comprises: acquiring sound source position information and at least two channels of sound signals from a microphone array; suppressing, according to the sound source position information, a sound signal from the sound source direction in the at least two channels of sound signals, to obtain a noise reference signal of the microphone array; acquiring, according to the sound source position information, a sound signal from the sound source direction in the at least two channels of sound signals, to obtain a speech reference signal; removing, based on the noise reference signal, a residual noise signal in the speech reference signal to obtain a desired speech signal.