Audio Signal Separation for Stable Ambience Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to effectively separate and render voice and ambience signals in a way that enhances speech intelligibility while maintaining an immersive audio experience, especially when the capture device moves, causing ambience sounds to appear to change direction.

Innovation Solution

A method involving microphone arrays that capture audio signals, process them into frequency domain signals, extract primary speech and ambience signals, generate spatial parameters for ambience, and encode these signals for playback, which includes spatializing ambience sounds to remain stationary while allowing primary speech to adapt to device movement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If ambience sounds are captured and played back with the capture device, then an immersive audio environment is provided, but the ambience sounds distract from and detract from the primary speaker's speech intelligibility

Engineering Contradiction:
Improveimmersive audio environmentVSAvoidspeech intelligibility
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The audio signal is segmented into distinct components: primary speaker speech and ambience sounds. The system separates these components through signal processing, allowing independent handling of each. The primary speaker is identified and extracted as a separate audio stream from the ambience, enabling selective spatial rendering where speech remains clear and focused while ambience provides immersive background context without distracting from the primary speaker

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different spatial rendering qualities are applied to different audio components. The primary speaker's speech is rendered with high clarity and directness (local quality optimized for intelligibility), while ambience sounds are rendered with spatial characteristics that create immersion (local quality optimized for environmental context). This local differentiation allows the system to optimize each component for its specific purpose rather than applying a uniform rendering approach

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If the capture device moves, then the user can interact with the environment dynamically, but the ambience sounds appear to change direction which is distracting and disorienting to the listener

Engineering Contradiction:
Improvedynamic interaction capabilityVSAvoidspatial stability of ambience sounds
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The system applies preliminary anti-action by detecting device movement and pre-compensating for its effect on ambience sound spatial positioning. When the capture device moves, the system calculates the movement and applies an offsetting transformation to the ambience sounds to counteract the apparent directional changes. This preliminary compensation prevents the disorienting effect before it occurs, maintaining stable spatial perception of ambience sounds even during device movement

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

Instead of allowing ambience sounds to naturally follow device movement (which causes disorientation), the system inverts the expected behavior by deliberately offsetting the ambience sound positioning in the opposite direction of device movement. This inversion creates a stable reference frame for ambience sounds that remains consistent relative to the listener, counterintuitively improving spatial stability by moving against the natural physical expectation

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS12283289B2Separating and rendering voice and ambience signals by offsetting impact of device movements
Publication Date: 2025.04.22 APPLE INC
  • US12283289B2 patent drawing
  • US12283289B2 patent drawing
  • US12283289B2 patent drawing

AI summary

Processing of ambience and speech can include extracting from audio signals, ambience and speech signals. One or more spatial parameters can be generated that define spatial characteristics of ambience sound in the one or more ambience audio signals. The primary speech signal, the one or more ambience audio signals, and the spatial parameters can be encoded into one or more encoded data streams. Other aspects are described and claimed.