Conference Audio Filter Coefficients for Speaker Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current conference systems struggle to effectively separate and enhance the intelligibility and association of spoken utterances between multiple participants in video or telephone conferences, especially in environments with multiple participants where audio signals overlap.
Innovation Solution
A method and device that utilize different filter coefficients for each participant to create a virtual acoustic space, allowing for spatial separation and positioning of audio signals using FIR filters and convolution techniques, enabling improved acoustic separation and identification of speakers through the use of identifiers and head tracking for dynamic adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple participants speak simultaneously in a conference system, then the quantity of information transmitted increases, but the intelligibility and association of spoken utterances deteriorates due to audio signal overlap
Solution Approach 1:
The patent segments the mixed audio signal into individual participant signals by generating unique identifiers for each participant and applying specific filter coefficients to separate their utterances. This segmentation allows the system to process and present each participant's speech independently, resolving the overlap problem while maintaining the quantity of information from all participants.
Solution Approach 2:
The patent applies different filter coefficients to different audio signals based on participant identifiers, creating localized acoustic characteristics for each participant. This local quality approach enhances the association of utterances with specific participants by giving each signal a distinct acoustic signature, thereby improving intelligibility without reducing information quantity.
2Device complexity
If traditional audio filtering is used in conference systems, then the device complexity remains low, but the acoustic separation and identification of speakers is insufficient
Solution Approach 1:
The patent changes the parameters of the filtering mechanism by introducing participant-specific filter coefficients that are dynamically adjusted based on identified speakers. This parameter change enables sophisticated acoustic separation and speaker identification without requiring complex hardware, as the enhancement is achieved through adaptive software-based filter coefficient modification.
3Ease of operation
If no spatial separation is implemented in audio output, then the ease of operation is high, but the association of utterances with specific participants deteriorates
Solution Approach 1:
The patent adds a spatial dimension to the audio output by creating a virtual acoustic space where participants are positioned at different locations. This dimensional enhancement allows listeners to associate utterances with specific participants through spatial cues without complicating the basic operation of the system, as the spatial separation is implemented through signal processing rather than physical reconfiguration.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Significantly improves the intelligibility and association of spoken utterances by creating a virtual acoustic space that allows for clear differentiation of participants, even in noisy or complex audio environments, enhancing the overall conference experience.
Implementation Method 1
the use of convolution in acoustics is known from 'Convolution: Faltung in der Studiopraxis'
Implementation Method 2
One implementation of a filter is achieved using convolution. This type of filter is called a Finite Impulse Response (FIR) filter.
Data Source
AI summary
A device for a conference system and method for operation thereof is provided. The device is configured to receive a first audio signal and a first identifier associated with a first participant. The device is further configured to receive a second audio signal and a second identifier associated with a second participant. The device includes a filter configured to filter the received first audio signal and the received second audio signal and to output a filtered signal to a number of electroacoustic transducers. The device includes a control unit connected to the filter. The control unit is configured to control one or more first filter coefficients based on the first identifier and to control one or more second filter coefficients based on the second identifier. Preferably the device comprises a headtracker function for changing the first and second filter coefficients depending on tracking of head's position.


