Conference Audio Filter Coefficients for Speaker Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current conference systems struggle to effectively separate and enhance the intelligibility and association of spoken utterances between multiple participants in video or telephone conferences, especially in environments with multiple participants where audio signals overlap.

Innovation Solution

A method and device that utilize different filter coefficients for each participant to create a virtual acoustic space, allowing for spatial separation and positioning of audio signals using FIR filters and convolution techniques, enabling improved acoustic separation and identification of speakers through the use of identifiers and head tracking for dynamic adjustments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple participants speak simultaneously in a conference system, then the quantity of information transmitted increases, but the intelligibility and association of spoken utterances deteriorates due to audio signal overlap

Engineering Contradiction:
Improvequantity of informationVSAvoidintelligibility of spoken utterances
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments the mixed audio signal into individual participant signals by generating unique identifiers for each participant and applying specific filter coefficients to separate their utterances. This segmentation allows the system to process and present each participant's speech independently, resolving the overlap problem while maintaining the quantity of information from all participants.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different filter coefficients to different audio signals based on participant identifiers, creating localized acoustic characteristics for each participant. This local quality approach enhances the association of utterances with specific participants by giving each signal a distinct acoustic signature, thereby improving intelligibility without reducing information quantity.

Inventive Principle:
Principle #3Local quality

2Device complexity

If traditional audio filtering is used in conference systems, then the device complexity remains low, but the acoustic separation and identification of speakers is insufficient

Engineering Contradiction:
Improvefiltering mechanism complexityVSAvoidacoustic separation of participants
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of the filtering mechanism by introducing participant-specific filter coefficients that are dynamically adjusted based on identified speakers. This parameter change enables sophisticated acoustic separation and speaker identification without requiring complex hardware, as the enhancement is achieved through adaptive software-based filter coefficient modification.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If no spatial separation is implemented in audio output, then the ease of operation is high, but the association of utterances with specific participants deteriorates

Engineering Contradiction:
Improveaudio output simplicityVSAvoidassociation of utterances with participants
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent adds a spatial dimension to the audio output by creating a virtual acoustic space where participants are positioned at different locations. This dimensional enhancement allows listeners to associate utterances with specific participants through spatial cues without complicating the basic operation of the system, as the spatial separation is implemented through signal processing rather than physical reconfiguration.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Significantly improves the intelligibility and association of spoken utterances by creating a virtual acoustic space that allows for clear differentiation of participants, even in noisy or complex audio environments, enhancing the overall conference experience.

Implementation Method 1

the use of convolution in acoustics is known from 'Convolution: Faltung in der Studiopraxis'

Methodology Applied
Scientific EffectConvolution:

Implementation Method 2

One implementation of a filter is achieved using convolution. This type of filter is called a Finite Impulse Response (FIR) filter.

Methodology Applied
Scientific EffectFinite Impulse Response (FIR) filter: Filter (electronic)

Data Source

PatentUS9049339B2Method for operating a conference system and device for a conference system
Publication Date: 2015.06.02 APPLE INC
  • US9049339B2 patent drawing
  • US9049339B2 patent drawing
  • US9049339B2 patent drawing

AI summary

A device for a conference system and method for operation thereof is provided. The device is configured to receive a first audio signal and a first identifier associated with a first participant. The device is further configured to receive a second audio signal and a second identifier associated with a second participant. The device includes a filter configured to filter the received first audio signal and the received second audio signal and to output a filtered signal to a number of electroacoustic transducers. The device includes a control unit connected to the filter. The control unit is configured to control one or more first filter coefficients based on the first identifier and to control one or more second filter coefficients based on the second identifier. Preferably the device comprises a headtracker function for changing the first and second filter coefficients depending on tracking of head's position.