Sound Acquisition Device Volume Control for Multi-Speaker Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sound acquisition methods struggle to automatically adjust the volume of speech sounds from multiple speakers at different levels, leading to inappropriate volume reproduction due to delays in setting appropriate amplification factors.
Innovation Solution
A sound acquisition method and apparatus that detect utterance periods, determine sound source positions, convert signals to frequency domain, calculate covariance matrices, and use filter coefficients to adjust the volume of each sound source to a desired level, ensuring appropriate volume control for each speaker.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single microphone is used to acquire speech sounds from multiple participants at different positions, then the device complexity is reduced, but the volume reproduction of different speakers becomes inappropriate and difficult to distinguish
Solution Approach 1:
The patent divides the acoustic space into multiple regions and assigns different microphones to each region. By segmenting the microphone array into multiple channels positioned at different locations, the system can independently capture speech sounds from different speakers at different positions, thereby resolving the volume reproduction issue while maintaining manageable device complexity
Solution Approach 2:
Each microphone channel is optimized for its specific local position and acoustic environment. The patent applies local quality by adjusting the characteristics of each microphone channel according to its specific location and the speech sounds it captures, enabling appropriate volume reproduction for each speaker while keeping the overall system design simple
2Extent of automation
If the amplification factor is determined based on long-time mean power, then the volume is automatically adjusted, but a delay of several to tens of seconds develops in setting appropriate amplification factor
Solution Approach 1:
The patent performs preliminary actions by pre-positioning multiple microphones in specific locations before speech sounds are captured. This preliminary arrangement of microphone channels allows the system to immediately identify and adjust amplification factors for different speakers without delay, while maintaining automatic volume adjustment functionality
Solution Approach 2:
The system continuously monitors the acoustic signals from multiple microphone channels and provides feedback to dynamically adjust amplification factors in real-time. This feedback mechanism eliminates the delay inherent in long-time mean power calculations by enabling immediate response to changes in speech sound levels
3Adaptability or versatility
If multiple speakers are present and their speech sounds are acquired at different levels, then the adaptability of the system is improved, but the speech sounds are reproduced at inappropriate volumes
Solution Approach 1:
The patent segments the acoustic space into multiple regions with dedicated microphones for each region. This segmentation enables the system to adapt to multiple speakers at different positions while maintaining precise volume control for each speaker, as each microphone channel is optimized for its specific local environment
Solution Approach 2:
The system dynamically adjusts the amplification factors of different microphone channels based on the detected speech sounds and their spatial positions. This dynamic adjustment capability allows the system to adapt to varying speaker positions and volumes while maintaining accurate volume reproduction for each speaker
Data Source
AI summary
Upon detecting an utterance period by a state decision part 14, a sound source position detecting part 15 detects the positions of sound sources 91 to 9K are detected by a sound source position detecting part 15, then covariance matrix of acquired signals are calculated by a covariance matrix calculating part 18 in correspondence to the respective sound sources, and stored in a covariance matrix storage part 18 in correspondence to the respective sound sources. The acquired sound level for each sound source is estimated by an acquired sound level estimating part 19 from the stored covariance matrix, and filter coefficients are determined by a filter coefficient calculating part 21 from the estimated acquired sound levels and the covariance matrices, and the filter coefficients are set in filters 121 to 12M. Acquired signals from the respective microphones are filtered by the filters, then the filtered outputs are added together by an adder 13, and the added output is provided as a send signal; by this, it is possible to generate send signals of desired levels irrespective of the positions of sound sources.


