Headset Local Audio Processing for Avatar Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtual reality (VR) and augmented reality (AR) systems, there is a delay in synchronizing avatars with audio, causing distractions and reducing communication frequency among users due to the lack of synchronization between avatar movements and audio detection within the same local area.
Innovation Solution
A headset captures audio from a local area, extracts characteristics, and identifies the user generating the audio, allowing it to update the avatar based on local audio information rather than relying on audio from the user's headset or server, thereby reducing latency and improving synchronization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio is transmitted from user headsets to a server for processing and distribution, then audio can be centrally managed and distributed to all users, but latency increases causing avatar updates to occur after users hear the audio locally
Solution Approach 1:
The system segments audio distribution into two paths: local area audio is captured and processed by the local headset without server intervention, while remote audio continues to use the server-based path. This segmentation allows low-latency local avatar updates while maintaining centralized management for remote participants.
Solution Approach 2:
The headset preliminarily captures and processes audio from the local area environment before server processing is needed. By having the headset already possess local audio characteristics and user identifiers, the system can immediately update avatars when local audio events occur, eliminating the wait time for server round-trip processing.
2Loss of time
If the headset captures and processes local audio independently, then avatar update latency is reduced, but the system complexity increases
Solution Approach 1:
The headset is designed with multi-functionality, serving both as a conventional audio communication device and as a local audio processing unit capable of capturing, analyzing, and acting on local audio events. This universal design allows the headset to handle both local and remote audio paths without requiring separate dedicated hardware systems.
Solution Approach 2:
The headset performs self-service by independently capturing local audio, identifying the speaking user through its sensors and processors, and updating its own display of local avatars without requiring external server intervention. This self-service capability reduces latency while the headset maintains standard connectivity for remote communication.
3Stability of the object's composition
If avatars are updated based on server-distributed audio, then synchronization across all users is maintained, but local users experience delayed avatar updates compared to their actual audio hearing
Solution Approach 1:
The system applies local quality by treating local area audio differently from remote audio. For local users, the headset captures and processes audio locally to achieve immediate avatar updates. For remote users, the conventional server-based synchronization is maintained. This differentiated approach optimizes latency for local interactions while preserving overall system synchronization stability.
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
A local area includes multiple users each using a headset to communicate with other users in the local area, as well as with additional users in a remote area. The headset includes one or more acoustic sensors, a transducer array, one or more external capturing sensors (e.g., cameras) capturing information describing the local area. A headset of a user in the local area local detects audio from an additional user in the local area. In response to detecting the audio, an avatar representing the additional user that is displayed by the user's headset is modified to appear to be synchronized with the audio from the additional user that is heard by the user. Additionally, each headset provides respective face tracking through one or more internal imaging devices that is provided with and captured audio to a server which provides the information users in the remote area.