Spatial Audio Depth Encoding for Nearfield Source Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio capture systems fail to encode depth information with spatial audio signals, leading to nearfield components being relegated to farfield, limiting the ability of decoders to differentiate and render nearfield sources accurately.

Innovation Solution

Concurrently using a depth sensor with an audio sensor to capture acoustic and visual information, enabling the encoding of spatial audio signals with depth characteristics, such as through the integration of a three-dimensional depth camera and microphone array to correlate audio objects with their respective depths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth sensor is integrated with audio sensor to capture depth information, then spatial audio rendering accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvedepth measurement precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines a depth sensor (such as a time-of-flight camera or structured light sensor) with an audio sensor array to create an integrated capture system. This merging allows the system to simultaneously capture both visual depth information and audio spatial information, enabling accurate association of audio sources with their corresponding spatial positions. The depth sensor provides Z-axis distance measurements that complement the audio sensors' directional information, resolving the technical contradiction by improving measurement precision through sensor fusion while managing system complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a processing system that acts as an intermediary between the depth sensor and audio sensor data. This intermediary component correlates depth information from the visual sensor with audio information from the microphones, creating a unified spatial audio representation. The processing system matches audio objects with their corresponding depth values by analyzing temporal and spatial correlations, effectively mediating between the two sensor types to achieve accurate spatial audio rendering without requiring direct physical integration of all components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If multiple sensor types are integrated to capture spatial audio with depth, then audio rendering accuracy is improved, but ease of manufacture deteriorates

Engineering Contradiction:
Improvespatial audio encoding precisionVSAvoidsystem integration difficulty
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent divides the spatial audio capture system into distinct functional modules: a depth sensor module for capturing Z-axis distance information, an audio sensor module for capturing acoustic spatial information, and a processing module for correlating the data. This segmentation allows each module to be optimized and manufactured independently using standard technologies, then integrated through well-defined interfaces. The segmented approach maintains high spatial audio encoding precision while improving ease of manufacture by avoiding the need for custom-integrated multi-sensor systems.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a processing system that can handle multiple sensor types and produce universal spatial audio output compatible with various playback systems. The correlation algorithm is designed to work with different depth sensor technologies (time-of-flight, structured light, stereo vision) and different audio sensor configurations, making the system universally applicable. This multi-functionality approach maintains manufacturing precision by providing accurate spatial encoding while improving ease of manufacture by allowing standard components to be used across different implementations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If depth information is encoded with spatial audio signals, then nearfield source differentiation is improved, but loss of information is reduced

Engineering Contradiction:
Improvespatial information lossVSAvoidencoding complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent adds the depth dimension (Z-axis distance information) to the traditional spatial audio encoding that primarily captures directional information (azimuth and elevation). By incorporating depth values from the depth sensor correlated with audio objects, the system creates a four-dimensional spatial representation (X, Y, Z, time) instead of the conventional three-dimensional representation. This dimensional enhancement reduces information loss by preserving nearfield source characteristics that would otherwise be indistinguishable from farfield sources, while the encoding complexity is managed through efficient data structures that store depth values alongside existing spatial parameters.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12501209B2Spatial audio capture and analysis with depth
Publication Date: 2025.12.16 DTS INC(US)
  • US12501209B2 patent drawing
  • US12501209B2 patent drawing
  • US12501209B2 patent drawing

AI summary

Spatial audio signals can include audio objects that can be respectively encoded and rendered at each of multiple different depths. In an example, a method for encoding a spatial audio signal can include receiving audio scene information from an audio capture source in an environment, and receiving a depth characteristic of a first object in the environment. The depth characteristic can be determined using information from a depth sensor. A correlation can be identified between at least a portion of the audio scene information and the first object. The spatial audio signal can be encoded using the portion of the audio scene and the depth characteristic of the first object.