Multi-Microphone Gain Control for Close-Talk Speech Leveling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic gain control (AGC) systems for sound signals struggle to distinguish between intentional and unintentional level variations, often requiring synchronized microphones, known positions, and are computationally complex, failing to effectively equalize close-talk level variations and handle simultaneously active talkers.
Innovation Solution
A sound signal processing apparatus using multiple microphones to estimate power measures and determine a gain factor based on the ratio between power measures from microphones at different distances from the target source, incorporating band-limited power measures and probabilities to differentiate between intentional and unintentional signal variations, allowing for robust and efficient automatic gain control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional automatic gain control systems are used to equalize level variations, then unintentional variations due to distance fluctuations can be equalized, but intentional natural dynamic changes of speech signals are also altered
Solution Approach 1:
The patent segments the level variations into two distinct components: intentional variations (natural dynamic changes) and unintentional variations (distance fluctuations). By separating these two types of variations, the system can selectively equalize only the unintentional ones while preserving the intentional natural dynamic changes of speech signals.
Solution Approach 2:
The patent introduces acoustic source localization (ASL) as an intermediary mechanism to detect and classify the cause of level variations. The ASL system acts as a mediator that identifies whether a level variation is intentional or unintentional, enabling the AGC system to make informed decisions about which variations to equalize and which to preserve.
2Measurement precision
If acoustic source localization methods are used to distinguish intentional and unintentional variations, then close-talk level variations can be equalized, but the methods require synchronized microphones with known positions
Solution Approach 1:
The patent extracts and eliminates the requirement for microphone synchronization and known positions from the acoustic source localization method. By removing these complex prerequisites, the system can perform close-talk level equalization using a simpler, more flexible microphone configuration that does not require precise synchronization or position information.
Solution Approach 2:
The patent replaces the complex synchronized microphone system with a simpler, more economical microphone configuration. The new approach uses basic microphone elements without requiring expensive synchronization hardware or complex positioning systems, making the solution more accessible and easier to implement.
3Measurement precision
If conventional acoustic source localization methods are used, then level variations can be detected, but the computational complexity increases significantly
Solution Approach 1:
The patent replaces complex computational acoustic source localization algorithms with a simpler, more efficient detection mechanism. By substituting the heavy computational approach with a lighter alternative, the system maintains level variation detection capability while significantly reducing computational complexity and processing requirements.
4Measurement precision
If acoustic source localization is used to equalize level variations, then distance fluctuations can be compensated, but simultaneously active talkers cannot be handled
Solution Approach 1:
The patent enhances the system's universality by enabling it to handle multiple functions: distance fluctuation compensation and simultaneous multi-talker processing. The improved acoustic source localization method can identify and track multiple sound sources independently, allowing the system to compensate for distance variations while managing simultaneously active talkers without interference.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a sound signal processing apparatus (100) for enhancing a sound signal from a target source. The sound signal processing apparatus (100) comprises a plurality of microphones (101a-f), wherein each microphone (101a-f) is configured to receive the sound signal from the target source; an estimator (103) configured to estimate a first power measure on the basis of the sound signal from the target source received by a first microphone (101a-f) of the plurality of microphones (101a-f) and a second power measure on the basis of the sound signal from the target source received by at least a second microphone (101a-f) of the plurality of microphones (101a-f), which is located more distant from the target source than the first microphone (101a-f), wherein the estimator (103) is further configured to determine a gain factor on the basis of a ratio between the second power measure and the first power measure; and an amplifier (105) configured to apply the gain factor to the sound signal from the target source received by the first microphone (101a-f).