Audio Signal Processing Apparatus for Clear Talker Utterance Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in accurately capturing the beginning of a talker's utterance while maintaining clarity, as current gain control methods either introduce time lag or reduce clarity due to voice leakage across microphones.
Innovation Solution
An audio signal processing apparatus that combines gating type and gain sharing type gain control by selecting a channel group based on volume levels and forming multiple sound collection beams, which narrows down the number of channels to improve clarity and capture the start of a talker's utterance effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If gating type gain control is used to block non-talker audio signals, then clarity is improved, but time lag occurs from when a talker is changed to when the microphone gain increases
Solution Approach 1:
The system performs preliminary gain adjustment before the talker switch is complete. By anticipating the talker change and pre-adjusting gains, the system eliminates the time lag that would otherwise occur when switching between talkers, while maintaining clarity through controlled gain distribution.
Solution Approach 2:
The gain values are dynamically adjusted based on real-time talker detection and voice activity. Instead of static gating, the system continuously modifies gain parameters to smoothly transition between talkers, eliminating time lag while preserving clarity through adaptive control.
2Loss of time
If gain sharing type control is used to set gain according to each audio signal level, then time lag is reduced, but clarity is reduced when voice leaks across multiple microphones
Solution Approach 1:
Different gain values are assigned to different microphone channels based on their specific characteristics and spatial relationships. The system identifies which microphones are capturing talker voice versus leakage, and applies localized gain control to each channel, maintaining overall responsiveness while preserving clarity by preventing leakage from dominating.
Solution Approach 2:
The system dynamically changes gain parameters based on detected voice activity and talker identification. By adjusting gain values in response to real-time conditions, the system maintains responsiveness (reducing time lag) while preserving clarity through context-aware parameter modification.
3Adaptability or versatility
If multiple microphones are used to capture talker voice, then coverage is improved, but clarity is reduced due to voice leakage across microphones
Solution Approach 1:
The audio signal processing is segmented by microphone channel, with each channel receiving individualized gain control based on its specific capture characteristics. This allows the system to utilize multiple microphones for broad coverage while maintaining clarity by processing and controlling each microphone's contribution separately.
Solution Approach 2:
The system applies gain control selectively to microphone channels, giving partial action to microphones that are capturing leakage versus those capturing primary talker voice. By using excessive gain on optimal microphones and suppressing others, the system achieves both broad coverage and high clarity.
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
An audio signal processing method includes selecting a channel group of at least two channels according to a predetermined reference, from among audio signals of at least three channels, and controlling a gain of the audio signal of each channel of selected channel group, according to a volume level of the audio signal of each channel of the channel group.