Speech Enhancement Using Directional Detection for Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech enhancement techniques fail to effectively utilize direction information to differentiate between speech and noise, leading to inefficient noise removal and reduced performance in devices like mobile phones, TVs, and wearable devices.
Innovation Solution
A speech enhancement apparatus and method that incorporates a sensor unit, speech detection unit, direction estimation unit, and speech enhancement unit to detect speech and estimate the speaker's direction, allowing for targeted noise reduction by distinguishing between speech and non-speech sections, thereby improving recognition rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech enhancement is performed continuously without speech detection, then noise removal may be maintained, but computing power is wasted during non-speech sections
Solution Approach 1:
The patent implements periodic speech detection followed by conditional speech enhancement only during detected speech sections. The speech detection unit periodically analyzes audio input to identify speech sections, and the speech enhancement unit is activated only during these sections, creating a periodic on-demand enhancement pattern that avoids continuous processing during non-speech periods.
2Measurement precision
If direction estimation is performed continuously, then speaker direction accuracy is maintained, but processing efficiency decreases
Solution Approach 1:
The patent performs speech detection as a preliminary action before direction estimation. The speech detection unit first identifies whether a speech section is present, and only when speech is detected does the direction estimation unit activate to estimate speaker direction. This preliminary filtering prevents unnecessary direction estimation processing during non-speech sections.
3Reliability
If speech enhancement is applied to all audio sections, then consistent noise suppression is achieved, but speech quality may be degraded in non-speech sections
Solution Approach 1:
The patent applies speech enhancement with different characteristics to different audio sections based on speech detection results. During speech sections, full speech enhancement processing is applied to remove noise. During non-speech sections, enhancement is suspended or applied with reduced intensity, preserving the natural characteristics of background audio and avoiding quality degradation from inappropriate processing.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A speech enhancement method including: estimating a direction of a speaker by using an input signal, generating direction information indicating the estimated direction, detecting speech of a speaker based on a result of the estimating a direction, and enhancing the speech of the speaker by using the direction information of the estimating of a direction based on a result of the detecting of speech.