XR Shared-Space Audio Control for Selective Voice Cancellation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing extended reality (XR) systems struggle to provide an immersive audio experience by effectively distinguishing and canceling non-participant voices while allowing participants to hear each other, which can lead to distractions and impaired communication.
Innovation Solution
Implementing a voice activity detection system that determines participant voices and generates an antinoise signal to cancel non-participant voices, while allowing participant voices to be heard through directional sound processing and wireless transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If active noise cancellation is applied to all audio activity in XR systems, then non-participant voices are effectively canceled, but participant voices are also canceled leading to impaired communication
Solution Approach 1:
The audio signal is segmented into different categories: participant voices and non-participant voices. The system processes these segments differently, applying noise cancellation only to non-participant voices while preserving participant voices for clear communication within the virtual space.
Solution Approach 2:
The system dynamically adjusts audio processing based on the speaker's location and participant status. When a speaker is identified as a participant within the virtual space, the system dynamically switches to preserve their voice. When a speaker is identified as a non-participant, the system dynamically applies noise cancellation.
2Extent of automation
If voice activity detection is applied without participant identification, then all voices are treated equally, but this fails to distinguish between participants and non-participants for selective cancellation
Solution Approach 1:
The system uses feedback from microphone signals, camera feeds, and sensor data to continuously identify speaker locations and determine participant status. This feedback loop enables the system to adaptively adjust audio processing in real-time, maintaining automatic operation while achieving selective cancellation based on participant identification.
Solution Approach 2:
The system integrates multiple functions into a unified audio processing pipeline: voice activity detection, speaker localization using cameras and sensors, participant identification, and selective noise cancellation. This multi-functional approach allows the system to handle diverse audio scenarios within a single framework.
3Measurement precision
If directional sound processing is used to maintain participant communication, then clear audio is preserved for participants, but non-participant voices may still cause distractions
Solution Approach 1:
The system applies different audio processing qualities to different spatial locations and participant groups. Participant voices receive high-fidelity processing with preservation, while non-participant voices in peripheral or unwanted directions receive noise cancellation processing, creating localized audio zones with different characteristics.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances immersive XR experiences by reducing distractions from non-participant voices while maintaining clear communication among participants, using active noise cancellation and voice recognition technologies.
Implementation Method 1
generating an antinoise signal to cancel the first audio activity; and, by a loudspeaker, producing an acoustic signal that is based on the antinoise signal
Data Source
AI summary
Methods, systems, computer-readable media, and apparatuses for audio signal processing are presented. Some configurations include determining that first audio activity in at least one microphone signal is voice activity; determining whether the voice activity is voice activity of a participant in an application session active on a device; based at least on a result of the determining whether the voice activity is voice activity of a participant in the application session, generating an antinoise signal to cancel the first audio activity; and by a loudspeaker, producing an acoustic signal that is based on the antinoise signal. Applications relating to shared virtual spaces are described.


