Dynamic Spatial Separation of Sound Objects
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for enhancing audio intelligibility in video content are limited, particularly in complex audio environments, and do not effectively utilize modern video formats or advanced listening technologies like spatial audio and inertial measurement units, leading to imperfect voice separation and localization.
Innovation Solution
The use of orientation-responsive audio enhancement, dynamic spatial separation, and frequency spreading techniques that leverage spatial audio and inertial measurement units to enhance voice intelligibility by identifying and amplifying sound objects based on user head orientation and gaze, and adjusting sound positions and frequencies to improve localization and separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If differential equalization is applied to amplify voice frequencies, then voice intelligibility is improved, but non-voice sounds in the same frequency range are also amplified, causing confusion
Solution Approach 1:
The patent segments the audio spectrum into multiple bands and applies different processing to each band. Instead of applying uniform differential equalization across all voice frequencies, the system processes different frequency ranges separately, allowing selective amplification of voice components while preserving or attenuating non-voice sounds in specific bands.
Solution Approach 2:
The patent applies local quality by treating different frequency regions differently. The system identifies which frequency bands contain voice and which contain non-voice sounds, then applies enhanced processing only to the voice-containing bands while maintaining original characteristics for non-voice bands, thus avoiding the confusion caused by uniform amplification.
2Measurement precision
If voice tracks are isolated and redirected to center-channel speaker, then voice presentation is strengthened, but sounds from original positions are lost and spatial accuracy is degraded
Solution Approach 1:
The patent implements dynamic processing where the system continuously monitors audio spatial characteristics and adjusts the redirection strength in real-time. When a sound source is clearly identified as voice, the system applies strong center-channel redirection; when the source position is ambiguous or the sound is non-voice, the system maintains the original spatial positioning, thus preserving spatial accuracy while enhancing voice presentation.
Solution Approach 2:
The patent changes the spatial distribution parameters of audio signals dynamically. Instead of fixed redirection to center-channel, the system adjusts the spatial parameters based on identified sound characteristics, applying different panning and volume distributions to voice versus non-voice sounds, thereby maintaining spatial information where appropriate while strengthening voice presentation where needed.
3Ease of manufacture
If simple frequency filtering is used to identify voice sounds, then implementation is straightforward, but other non-voice sounds in the voice frequency band are incorrectly identified
Solution Approach 1:
The patent segments the voice identification process into multiple stages: initial frequency band filtering, temporal pattern analysis, spectral characteristic analysis, and contextual verification. This multi-stage segmentation allows the system to maintain implementation simplicity through modular processing while achieving high identification accuracy by combining multiple verification criteria that distinguish voice from non-voice sounds.
Solution Approach 2:
The patent applies preliminary action by performing initial frequency filtering to narrow down candidate regions, then applying additional preliminary analyses such as temporal envelope detection and spectral flux calculation before final voice identification. This staged approach maintains ease of implementation through sequential processing while significantly improving identification accuracy by eliminating false positives through multiple preliminary checks.
4Adaptability or versatility
If audio content is presented in simple stereo mix, then compatibility is maintained, but selective attention to specific sound sources is difficult
Solution Approach 1:
The patent implements universality by designing a processing system that works across multiple audio formats and device types. The system can process Dolby Atmos, spatial audio, and traditional stereo mixes, adapting its processing accordingly. This multi-functionality maintains compatibility with existing formats while enabling selective attention capabilities through spatial audio processing and head-tracking integration.
Solution Approach 2:
The patent applies dynamics by making the audio presentation adaptive to user head orientation and attention state. The system dynamically adjusts spatial audio rendering based on real-time head tracking data, automatically directing sound sources toward the user's point of interest. This dynamic adaptation enables selective audio attention without requiring users to manually adjust settings, while maintaining compatibility with various audio formats through intelligent processing.
Data Source
AI summary
Sound objects are identified within a content item and location metadata is extracted from the content item for each sound object. A reference layout is generated, relative to a user position, for the sound objects based on the location metadata. If a first sound object is within a threshold angle, relative to the user position, from a second sound object, a virtual position of either the first sound object or the second sound object is adjusted by an adjustment angle.


