Head-Tracked Audio Stream Control for Virtual Meeting Focus
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual meeting systems struggle to effectively manage audio focus based on user attention, leading to distractions and difficulty in hearing the intended speaker amidst ambient noise and multiple audio streams.
Innovation Solution
Implementing client applications that track user head and eye movements to determine focus, adjusting audio stream volumes accordingly, and providing visual indications of focus and volume changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple audio streams are provided simultaneously in virtual meetings, then participants can hear all speakers, but ambient noise increases and users have difficulty hearing the intended speaker
Solution Approach 1:
The patent applies local quality by adjusting audio volume based on the user's focal point. Different audio streams receive different volume levels depending on which speaker or region the user is looking at, rather than treating all audio equally. This creates localized audio enhancement where the focused speaker's voice is amplified while other audio streams are attenuated, directly resolving the contradiction between hearing all speakers and filtering ambient noise.
Solution Approach 2:
The system dynamically adjusts audio stream volumes in real-time based on the user's head position and eye gaze tracking data. The audio mixing parameters change continuously as the user moves their attention between different speakers, allowing the system to adapt to the user's current focus and maintain audio clarity without manual intervention.
2Reliability
If audio volume is increased for the focused speaker, then audio clarity improves, but other audio streams become harder to hear when needed
Solution Approach 1:
The audio system dynamically switches between different speakers based on real-time tracking of the user's head position and eye gaze. When the user looks at a different speaker, the system automatically adjusts which audio stream is emphasized, providing seamless adaptation without requiring manual audio switching. This maintains both audio clarity for the current focus and the ability to quickly transition to other speakers.
Solution Approach 2:
The system uses feedback from eye tracking and head position sensors to continuously monitor user attention and adjust audio mixing accordingly. This closed-loop control ensures that the audio output always matches the user's current focus, allowing them to hear the intended speaker clearly while maintaining the ability to switch focus to other speakers as needed.
3Ease of operation
If manual audio adjustment controls are provided, then users can control audio focus, but user interaction complexity increases
Solution Approach 1:
The system performs audio focusing automatically without requiring user interaction. By using eye tracking and head position data, the system self-determines which speaker the user wants to hear and adjusts audio volumes accordingly. This eliminates the need for manual audio controls while maintaining ease of operation, as the system adapts to user preferences automatically.
Solution Approach 2:
The patent replaces manual mechanical audio controls with an automated optical tracking system. Instead of requiring users to manually adjust audio sliders or buttons, the system uses eye and head tracking technology to automatically determine user focus and adjust audio mixing, substituting a more sophisticated sensing system for simple manual controls.
Data Source
AI summary
One example method includes joining, from a client device, a virtual meeting hosted by a virtual meeting provider, the virtual meeting comprising a plurality of participants, displaying the virtual meeting on the display of the client device, receiving a head tracking signal from an head tracking sensor, the head tracking signal associated with a first user of the client device, determining, based at least in part on the head tracking signal, a location and a position of the of user's head in relation to the display on which the first user is focused, and varying the volume of a first audio stream of a plurality of audio streams associated with the virtual meeting based at least in part on the position of the user's head.


