Spatialized Audio Positioning in Wearable Communication Sessions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio output systems for communication sessions, such as video calls, lack the ability to spatially separate audio sources, leading to a less immersive and less realistic listening experience, as all participants' voices are perceived to come from a single point in space.
Innovation Solution
The method involves displaying dynamic visual representations of participants in a communication session on a user interface and outputting audio from each participant to maintain a simulated spatial location relative to the communication session's frame of reference, independent of the wearable audio output device's position.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional audio output modes (stereo/mono) are used, then the audio output device can provide audio output, but the listening experience is less immersive and less realistic because all participants' voices are perceived to come from one overlapping point in space
Solution Approach 1:
The patent applies spatial audio technology to add a spatial dimension to audio output, transforming conventional 2D stereo/mono audio into 3D spatial audio. Each participant's voice is assigned to a specific spatial location corresponding to their visual position on the screen, creating a realistic soundstage that matches the visual layout and provides an immersive listening experience.
2Reliability
If spatial audio output mode is implemented, then the listening experience becomes more immersive and realistic, but the device complexity increases
Solution Approach 1:
The system integrates spatial audio processing into the existing audio output framework, allowing the same audio processing pipeline to handle both conventional and spatial audio modes. The spatial audio functionality is implemented as an extension of the existing audio stack, sharing common components such as audio decoding, mixing, and output drivers, thereby reducing the incremental complexity.
3Productivity
If conventional communication session methods are used, then the system can function, but it takes longer and requires more user interaction (constant pauses when participants interrupt), resulting in increased user mistakes and wasted energy
Solution Approach 1:
The system implements spatial audio feedback that provides users with directional information about which participant is speaking and from where. This spatial feedback helps users quickly identify and focus on relevant speakers, reducing the need for constant visual attention and minimizing interruptions, thereby improving communication efficiency and reducing time loss.
4Reliability
If spatial audio output mode is used in battery-operated devices, then the immersive listening experience is provided, but power consumption increases
Solution Approach 1:
The system dynamically adjusts spatial audio processing based on operational conditions such as device state, audio content, and user context. Spatial audio effects are applied selectively rather than continuously, and processing intensity is modulated according to the number of active participants and their spatial positions, optimizing power consumption while maintaining the immersive experience when needed.
Data Source
AI summary
A first wearable audio output device that is in communication with a second wearable audio output device. While the first wearable audio output devices and the second set of wearable audio output device is in an audio communication session, outputting, via the first wearable audio output device, audio from the second wearable audio output device, including, as the first wearable audio output device is moved relative to the second wearable audio output devices. Adjusting the audio so as to position the audio at a simulated spatial location relative to the first wearable audio output devices that is determined based on a respective position of the second wearable audio output devices relative to the first wearable audio output devices. Adjusting an output property other than a simulated spatial location of the audio based on a distance of the second wearable audio output devices from the first wearable audio output devices.


