Pre-Generated Inverse Audio Tracks for Selective Voice Cancellation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio systems fail to effectively cancel specific audio components, such as undesirable voices or commentary, when multiple audio sources are present, leading to an undesirable user experience, especially in environments with extended reality headsets.
Innovation Solution
The use of pre-generated inverse audio tracks, encoded during multimedia content creation, which are synchronized with the original audio to cancel out specific audio components, allowing selective audio silencing or replacement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If active noise canceling is used to cancel ambient audio, then audio cancellation is achieved, but the ability to selectively cancel specific audio components (e.g., particular voices or commentary) is lost
Solution Approach 1:
The audio stream is segmented into multiple independent components (e.g., commentary track, soundtrack, sound effects) during encoding. Each component can be independently identified and canceled using pre-generated inverse audio tracks, allowing selective cancellation of specific elements like a particular commentator's voice while preserving other audio elements.
Solution Approach 2:
Inverse audio tracks are pre-generated during the multimedia content creation process rather than being generated in real-time. This preliminary action enables the system to have ready-to-use cancellation signals for each audio component, improving the speed and accuracy of selective cancellation while reducing computational complexity during playback.
2Adaptability or versatility
If real-time audio processing is used to identify and cancel specific audio components, then selective cancellation is achieved, but computational complexity and processing time increase
Solution Approach 1:
The computationally intensive task of generating inverse audio tracks is performed in advance during content creation, not in real-time during playback. This shifts the computational burden to a preprocessing stage where complex algorithms can be used without time constraints, while playback devices only need to play back pre-computed cancellation signals.
Solution Approach 2:
Instead of processing and generating cancellation signals from scratch during playback, the system uses pre-generated copies of inverse audio tracks that have already been computed. This copying approach eliminates the need for real-time computational processing of complex algorithms while maintaining cancellation effectiveness.
3Adaptability or versatility
If multiple audio tracks are encoded and delivered within video or audio streams, then selective audio delivery is enabled, but data transmission volume and processing overhead increase
Solution Approach 1:
Specific audio components (e.g., commentary, unwanted voices) are extracted as separate identifiable tracks within the audio stream. This extraction allows the system to target and cancel only the problematic components without needing to process or transmit separate cancellation data for the entire audio stream, reducing overall data requirements.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables users to selectively silence or replace specific audio components, enhancing user experience by reducing ambient audio interference and allowing personalized audio delivery.
Implementation Method 1
a derived inverse audio track which is 180 degrees out of phase (inverted) with respect to the 'heard' audio may be played back, which when combined with the ambient ('heard') sound causes the 'heard' sound to be 'canceled' out
Data Source
AI summary
System and method are provided for pre-generated inverse audio canceling. A system may identify source audio content that a first device is playing via a first speaker, retrieve pre-generated inverse audio content associated with the identified audio content, modify at least a portion of the retrieved inverse audio content, and cause the modified inverse audio content to be played in synchronization with the identified source audio content to attenuate at least a portion of the source audio content that is playing via the first speaker.


