Dynamic Spatial Separation of Sound Objects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for enhancing audio intelligibility in video content are limited, particularly in complex audio environments, and do not effectively utilize modern video formats or advanced listening technologies like spatial audio and inertial measurement units, leading to imperfect voice separation and localization.

Innovation Solution

The use of orientation-responsive audio enhancement, dynamic spatial separation, and frequency spreading techniques that leverage spatial audio and inertial measurement units to enhance voice intelligibility by identifying and amplifying sound objects based on user head orientation and gaze, and adjusting sound positions and frequencies to improve localization and separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If differential equalization is applied to amplify voice frequencies, then voice intelligibility is improved, but non-voice sounds in the same frequency range are also amplified, causing confusion

Engineering Contradiction:
Improvevoice intelligibilityVSAvoidconfusion from amplified non-voice sounds
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent segments the audio spectrum into multiple bands and applies different processing to each band. Instead of applying uniform differential equalization across all voice frequencies, the system processes different frequency ranges separately, allowing selective amplification of voice components while preserving or attenuating non-voice sounds in specific bands.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by treating different frequency regions differently. The system identifies which frequency bands contain voice and which contain non-voice sounds, then applies enhanced processing only to the voice-containing bands while maintaining original characteristics for non-voice bands, thus avoiding the confusion caused by uniform amplification.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If voice tracks are isolated and redirected to center-channel speaker, then voice presentation is strengthened, but sounds from original positions are lost and spatial accuracy is degraded

Engineering Contradiction:
Improvevoice presentation strengthVSAvoidspatial position information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent implements dynamic processing where the system continuously monitors audio spatial characteristics and adjusts the redirection strength in real-time. When a sound source is clearly identified as voice, the system applies strong center-channel redirection; when the source position is ambiguous or the sound is non-voice, the system maintains the original spatial positioning, thus preserving spatial accuracy while enhancing voice presentation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the spatial distribution parameters of audio signals dynamically. Instead of fixed redirection to center-channel, the system adjusts the spatial parameters based on identified sound characteristics, applying different panning and volume distributions to voice versus non-voice sounds, thereby maintaining spatial information where appropriate while strengthening voice presentation where needed.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If simple frequency filtering is used to identify voice sounds, then implementation is straightforward, but other non-voice sounds in the voice frequency band are incorrectly identified

Engineering Contradiction:
Improveimplementation simplicityVSAvoidvoice sound identification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent segments the voice identification process into multiple stages: initial frequency band filtering, temporal pattern analysis, spectral characteristic analysis, and contextual verification. This multi-stage segmentation allows the system to maintain implementation simplicity through modular processing while achieving high identification accuracy by combining multiple verification criteria that distinguish voice from non-voice sounds.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing initial frequency filtering to narrow down candidate regions, then applying additional preliminary analyses such as temporal envelope detection and spectral flux calculation before final voice identification. This staged approach maintains ease of implementation through sequential processing while significantly improving identification accuracy by eliminating false positives through multiple preliminary checks.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If audio content is presented in simple stereo mix, then compatibility is maintained, but selective attention to specific sound sources is difficult

Engineering Contradiction:
Improveaudio format compatibilityVSAvoidselective audio attention
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements universality by designing a processing system that works across multiple audio formats and device types. The system can process Dolby Atmos, spatial audio, and traditional stereo mixes, adapting its processing accordingly. This multi-functionality maintains compatibility with existing formats while enabling selective attention capabilities through spatial audio processing and head-tracking integration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies dynamics by making the audio presentation adaptive to user head orientation and attention state. The system dynamically adjusts spatial audio rendering based on real-time head tracking data, automatically directing sound sources toward the user's point of interest. This dynamic adaptation enables selective audio attention without requiring users to manually adjust settings, while maintaining compatibility with various audio formats through intelligent processing.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12167224B2Systems and methods for dynamic spatial separation of sound objects
Publication Date: 2024.12.10 ADEIA GUIDES INC
  • US12167224B2 patent drawing
  • US12167224B2 patent drawing
  • US12167224B2 patent drawing

AI summary

Sound objects are identified within a content item and location metadata is extracted from the content item for each sound object. A reference layout is generated, relative to a user position, for the sound objects based on the location metadata. If a first sound object is within a threshold angle, relative to the user position, from a second sound object, a virtual position of either the first sound object or the second sound object is adjusted by an adjustment angle.