Personalized Spatial Audio Positioning for Mobile Voice Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems fail to seamlessly continue music or other audio uninterrupted during phone calls and effectively distinguish between audio sources, especially in mobile phone scenarios where multichannel speaker setups are not available, due to limitations in directionality and distance perception using non-individualized Binaural Room Impulse Responses (BRIRs).
Innovation Solution
The use of individualized BRIRs derived from listener-specific properties to position audio streams at distinct spatial locations, allowing music and voice communications to be directed to separate positions, enhancing the listener's ability to differentiate between sources by simulating head, torso, and ear effects, and adjusting distances based on priority signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If non-individualized Binaural Room Impulse Responses (BRIRs) are used to position audio streams, then the system complexity is reduced and ease of manufacture is improved, but the listener's ability to distinguish between audio sources and perceive spatial positions accurately deteriorates
Solution Approach 1:
The system performs preliminary measurement of individual listener properties (head, torso, and ear characteristics) and pre-computes personalized BRIRs before actual audio playback. This allows the system to have accurate spatial audio transfer functions ready in advance, resolving the contradiction by preparing precise individualized data beforehand rather than computing it in real-time during audio playback.
Solution Approach 2:
The system creates virtual copies of individual listener's acoustic characteristics by measuring their head, torso, and ear properties and generating corresponding BRIRs. These copied acoustic signatures enable accurate spatial audio positioning without requiring complex real-time processing of actual listener anatomy, thus maintaining ease of implementation while achieving high localization accuracy.
2Loss of information
If audio streams are positioned at distinct spatial locations using individualized BRIRs, then the listener's ability to differentiate between audio sources is improved, but the device complexity increases due to personalized processing requirements
Solution Approach 1:
The system extracts only the essential individualized acoustic characteristics (head, torso, and ear properties) from each listener and uses these extracted features to generate BRIRs. By taking out only the necessary personalization parameters rather than processing entire acoustic profiles, the system reduces processing complexity while maintaining the ability to differentiate audio sources effectively.
Solution Approach 2:
The system changes the parameter representation from complex full acoustic field measurements to simplified individualized BRIR parameters based on head, torso, and ear geometry. This parameter transformation maintains audio source distinguishability while reducing the computational complexity of personalized spatial audio processing.
3Productivity
If music continues uninterrupted during phone calls, then the productivity of audio playback is improved, but the ability to distinguish between multiple audio sources deteriorates without spatial separation
Solution Approach 1:
The system adds spatial dimension (three-dimensional positioning) to audio playback, allowing music and phone calls to occupy different spatial locations simultaneously. By transitioning from two-dimensional volume mixing to three-dimensional spatial separation, the system maintains music continuity while enabling clear distinction between multiple audio sources through their different positional information.
Solution Approach 2:
The system introduces personalized BRIRs as an intermediary that processes and positions audio streams in spatial space. This intermediary enables music and phone calls to coexist by mediating their spatial placement, allowing uninterrupted music playback while maintaining distinguishability through spatial separation facilitated by the BRIR processing layer.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An audio rendering system includes a processor that combines audio input signals with personalized spatial audio transfer functions preferably including room responses. The personalized spatial audio transfer functions are selected from a database having a plurality of candidate transfer function datasets derived from in-ear microphone measurements for a plurality of individuals. Alternatively, the personalized transfer function datasets are derived from actual in-ear measurements of the listener. Foreground and background positions are designated and matched with transfer function pairs from the selected dataset for the foreground and background direction and distance. Two channels of input audio such as voice and music are processed. When a voice communication such as a phone call is accepted the music being rendered is moved from a foreground to a background channel corresponding to a background spatial audio position using the personalized transfer functions. The voice call is simultaneously transferred to the foreground channel.