Multitalker Beamforming via Spatial Analysis and Historical Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional beamforming technologies in voice conferencing with multiple talkers often result in poor audio quality due to noise and reverberation, particularly when multiple speakers are close or far from the microphones, leading to unstable beam steering and reduced intelligibility.
Innovation Solution
The method involves conducting spatial analysis and feature extraction to determine the relative location of sound objects, aggregating historical data to adjust beamforming patterns, and selectively applying beamforming based on the direct-to-reverberation ratio to optimize audio reception, thereby reducing noise and improving capture quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming is applied to capture the voice of an individual talker, then the signal to noise ratio is improved, but the quality deteriorates when multiple talkers speak simultaneously
Solution Approach 1:
The beamforming system dynamically adapts its beam patterns and width based on the number of active talkers detected. When multiple talkers are speaking simultaneously, the system adjusts the beamforming parameters to capture a wider angular range, ensuring all talkers remain audible. This dynamic adaptation resolves the contradiction by maintaining audio quality across varying conference conditions.
Solution Approach 2:
The system changes key beamforming parameters including beam width, steering angles, and null placement based on the detected conference state. When multiple talkers are present, the beam width is increased and steering angles are adjusted to cover multiple directions. This parameter adaptation allows the system to maintain both signal to noise ratio and audio quality simultaneously.
2Measurement precision
If the beamformer steers directionally towards the most salient talker, then the clarity of pick up is improved, but other talkers become inaudible
Solution Approach 1:
The system segments the audio capture into multiple simultaneous beam patterns, each targeting a different talker. Instead of a single directional beam, multiple beams are formed in parallel, each optimized for a specific talker's direction. This segmentation allows all talkers to be captured with clarity simultaneously, resolving the contradiction between pick up clarity and talker audibility.
Solution Approach 2:
The beamforming system performs multiple functions simultaneously by maintaining several active beam patterns that cover different angular ranges. Each beam pattern serves the function of capturing a specific talker, and collectively they provide universal coverage for all participants. This multi-functionality ensures no talker becomes inaudible while maintaining clarity for each.
3Measurement precision
If beamforming is applied in reverberant environments to isolate desired sound, then the isolation of desired sound is improved, but the complexity of the system increases
Solution Approach 1:
The system performs preliminary spatial analysis and identifies talker positions before applying complex beamforming patterns. By pre-determining the angular locations of active talkers and preparing appropriate beam patterns in advance, the system reduces computational complexity during real-time operation. This preliminary action maintains sound isolation quality while reducing overall system complexity.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
A method of processing a series of microphone inputs of an audio conference, the method including the steps of: (a) conducting a spatial analysis and feature extraction of the audio conference based on current audio activity; (b) aggregating historical information to obtain information about the approximate relative location of recent sound objects relative to the series of microphone inputs; (c) utilising the relative location or distance of the sound objects from the series of microphone inputs to determine if beam forming should be utilised to enhance the audio reception from recent sound objects.