Audio Signal Processing Apparatus for Dynamic Spatial Arrangement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual spatial audio conference systems face challenges in improving speech intelligibility, particularly when the target speaker is unknown or variable, as they rely on a priori knowledge and ideal time-frequency binary masks that are not practical in dynamic multi-party settings.
Innovation Solution
An audio signal processing apparatus and method that selects an optimal spatial arrangement of virtual audio sources based on both audio signal spectra and directional information, using transfer functions like HRTFs or BRTFs, to enhance speech intelligibility by filtering audio signals to simulate distinct virtual positions, thereby improving the separation of target and masker voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If ideal time-frequency binary masks are used to separate target speaker from maskers, then speech intelligibility is improved, but the system requires a priori knowledge of target speaker and masker signals which is not available in dynamic multi-party settings
Solution Approach 1:
The system uses the audio signals themselves to automatically determine spatial arrangements and identify target speakers without requiring external a priori knowledge. The audio signals provide information about spectral content and spatial position, enabling the system to self-organize and adapt to dynamic multi-party settings autonomously
Solution Approach 2:
The spatial arrangement of virtual audio sources is made dynamic and adjustable based on the actual audio signal characteristics. The system continuously adapts the spatial configuration to match the current speech situation, allowing it to handle dynamic multi-party settings where target speakers change over time
2Measurement precision
If fixed spatial arrangements of virtual audio sources are used, then the system complexity is reduced, but speech intelligibility cannot be optimized for different speech situations and target speakers
Solution Approach 1:
The system changes spatial parameters (positions of virtual audio sources) based on audio signal characteristics. By adjusting these parameters dynamically, the system optimizes speech intelligibility for different situations without requiring a completely complex reconfiguration mechanism
Solution Approach 2:
The system uses feedback from audio signal analysis (spectral content, spatial distribution) to automatically adjust spatial arrangements. This feedback mechanism enables optimization of speech intelligibility while keeping the control mechanism relatively simple and rule-based
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly enhances speech intelligibility by up to 12-13 dB by dynamically adjusting the spatial arrangement of audio sources, improving the naturalness and clarity of speech in multi-party audio conferences, even when the target speaker is not known a priori.
Implementation Method 1
spatial filters derived from head-related impulse responses (HRIR) or their corresponding frequency-domain representations, i.e. head-related transfer functions (HRTFs)
Implementation Method 2
binaural room impulse responses (BRIR) or their corresponding frequency-domain representations, i.e. binaural room transfer functions (BRTF)
Implementation Method 3
These filters encode the auditory cues humans use for spatial sound perception, namely interaural time difference (ITD), interaural level difference (ILD)
Implementation Method 4
These filters encode the auditory cues humans use for spatial sound perception, namely interaural time difference (ITD), interaural level difference (ILD)
Implementation Method 5
this psychoacoustic effect, scientifically known as spatial release from masking, can improve speech intelligibility by up to 12-13 dB when a target speaker and competing speakers, typically referred to as maskers, are virtually spatially separated
Data Source
AI summary
The disclosure relates to an audio signal processing apparatus for processing a plurality of audio signals defining a plurality of audio signal spectra, the audio signals to be transmitted to a listener in such a way that the listener perceives the audio signals to originate from virtual positions of a plurality of audio signal sources. The audio signal processing apparatus comprises a selector configured to select a spatial arrangement of the virtual positions of the audio signal sources relative to the listener from a plurality of possible spatial arrangements, and a filter configured to filter the plurality of audio signals on the basis of the selected spatial arrangement.


