Audio Signal Processing Apparatus for Dynamic Spatial Arrangement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual spatial audio conference systems face challenges in improving speech intelligibility, particularly when the target speaker is unknown or variable, as they rely on a priori knowledge and ideal time-frequency binary masks that are not practical in dynamic multi-party settings.

Innovation Solution

An audio signal processing apparatus and method that selects an optimal spatial arrangement of virtual audio sources based on both audio signal spectra and directional information, using transfer functions like HRTFs or BRTFs, to enhance speech intelligibility by filtering audio signals to simulate distinct virtual positions, thereby improving the separation of target and masker voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ideal time-frequency binary masks are used to separate target speaker from maskers, then speech intelligibility is improved, but the system requires a priori knowledge of target speaker and masker signals which is not available in dynamic multi-party settings

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidadaptability to dynamic multi-party settings
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system uses the audio signals themselves to automatically determine spatial arrangements and identify target speakers without requiring external a priori knowledge. The audio signals provide information about spectral content and spatial position, enabling the system to self-organize and adapt to dynamic multi-party settings autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The spatial arrangement of virtual audio sources is made dynamic and adjustable based on the actual audio signal characteristics. The system continuously adapts the spatial configuration to match the current speech situation, allowing it to handle dynamic multi-party settings where target speakers change over time

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If fixed spatial arrangements of virtual audio sources are used, then the system complexity is reduced, but speech intelligibility cannot be optimized for different speech situations and target speakers

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidspatial arrangement selection mechanism
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes spatial parameters (positions of virtual audio sources) based on audio signal characteristics. By adjusting these parameters dynamically, the system optimizes speech intelligibility for different situations without requiring a completely complex reconfiguration mechanism

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses feedback from audio signal analysis (spectral content, spatial distribution) to automatically adjust spatial arrangements. This feedback mechanism enables optimization of speech intelligibility while keeping the control mechanism relatively simple and rule-based

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly enhances speech intelligibility by up to 12-13 dB by dynamically adjusting the spatial arrangement of audio sources, improving the naturalness and clarity of speech in multi-party audio conferences, even when the target speaker is not known a priori.

Implementation Method 1

spatial filters derived from head-related impulse responses (HRIR) or their corresponding frequency-domain representations, i.e. head-related transfer functions (HRTFs)

Methodology Applied
Scientific EffectHead-Related Transfer Functions (HRTF):

Implementation Method 2

binaural room impulse responses (BRIR) or their corresponding frequency-domain representations, i.e. binaural room transfer functions (BRTF)

Methodology Applied
Scientific EffectBinaural Room Transfer Functions (BRTF):

Implementation Method 3

These filters encode the auditory cues humans use for spatial sound perception, namely interaural time difference (ITD), interaural level difference (ILD)

Methodology Applied
Scientific EffectInteraural Time Difference (ITD):

Implementation Method 4

These filters encode the auditory cues humans use for spatial sound perception, namely interaural time difference (ITD), interaural level difference (ILD)

Methodology Applied
Scientific EffectInteraural Level Difference (ILD):

Implementation Method 5

this psychoacoustic effect, scientifically known as spatial release from masking, can improve speech intelligibility by up to 12-13 dB when a target speaker and competing speakers, typically referred to as maskers, are virtually spatially separated

Methodology Applied
Scientific EffectSpatial Release from Masking:

Data Source

PatentUS10412226B2Audio signal processing apparatus and method
Publication Date: 2019.09.10 HUAWEI TECH CO LTD
  • US10412226B2 patent drawing
  • US10412226B2 patent drawing
  • US10412226B2 patent drawing

AI summary

The disclosure relates to an audio signal processing apparatus for processing a plurality of audio signals defining a plurality of audio signal spectra, the audio signals to be transmitted to a listener in such a way that the listener perceives the audio signals to originate from virtual positions of a plurality of audio signal sources. The audio signal processing apparatus comprises a selector configured to select a spatial arrangement of the virtual positions of the audio signal sources relative to the listener from a plurality of possible spatial arrangements, and a filter configured to filter the plurality of audio signals on the basis of the selected spatial arrangement.