HRTF Filtered Conference Device for Monaural Audio Spatialization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In non-homogeneous communication systems, the varying voice transmission quality across different networks and technologies diminishes the intelligibility and distinguishability of conference participants, particularly in conference systems where connections often rely on low-bandwidth mobile radio networks.
Innovation Solution
A conference device equipped with monaural HRTF filters that simulate the sound signal adjustment experienced by the human ear, allowing for virtual positioning of participants in different directions within the median plane, enhancing differentiation and intelligibility by using individually tailored HRTF filter coefficient sets, and a converter to integrate binaural audio signals into monaural systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If monaural audio transmission is used in non-homogeneous communication systems, then bandwidth consumption is reduced and system compatibility is improved, but voice transmission quality and participant distinguishability deteriorate
Solution Approach 1:
The patent introduces HRTF filters as an intermediary processing stage between the monaural audio signal and the listener's ear. These filters simulate the acoustic path from different spatial positions to the ear, creating virtual spatial cues that help distinguish between multiple participants. The filter acts as a mediator that adds spatial information to the monaural signal without requiring additional bandwidth.
Solution Approach 2:
The patent applies parameter changes by modifying the spectral characteristics of the monaural audio signal through HRTF filtering. The filter coefficients are adjusted based on the desired virtual position of the speaker, changing frequency-dependent parameters to create directional auditory impressions. This allows the system to convey spatial information through parameter manipulation rather than additional signal channels.
2Device complexity
If conventional audio filtering is used, then system complexity is kept low, but participant distinguishability and spatial perception deteriorate
Solution Approach 1:
The patent implements preliminary action by pre-calculating and storing HRTF filter coefficient sets for various spatial positions before the conference takes place. During the conference, the system simply selects and applies the appropriate pre-computed filter coefficients based on the participant's position, rather than performing complex real-time spatial calculations. This reduces computational complexity while maintaining distinguishability.
3Measurement precision
If binaural HRTF filtering is implemented, then spatial auditory impression and participant differentiation are improved, but system complexity and bandwidth requirements increase
Solution Approach 1:
The patent extracts only the essential spatial filtering characteristics needed for participant differentiation, rather than implementing full binaural processing. It takes out the key spectral cues from HRTF data and applies them to monaural signals, discarding redundant information. This extraction approach maintains spatial perception benefits while reducing system complexity and bandwidth requirements.
Data Source
AI summary
The conference device (EMCU) according to the invention has several monaural HRTF filters (HRTF1, . . . , HRTFN), each of which is to be allocated to a conference participant. The abbreviation HRTF stands for “Head Related Transfer Function”. A corresponding HRTF filter (HRTF1, . . . , HRTFN) is used for filtering a monaural audio signal coming from the conference participant to whom it is allocated. An individual, monaural HRTF filter coefficient set is responsible in this case for defining the filter characteristics of a particular HRTF filter. The conference device (EMCU) also has a conference mixing device (MP) coupled to the HRTF filters (HRTF1, . . . , HRTFN), for mixing the individually filtered audio signals from various conference participants and for transferring the mixed audio signals to conference participants.

