Stereo Microphones for 3D Audio Positioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As the number of channels and speaker configurations in audio playback systems increases from 2D to 3D, accurately positioning and rendering sounds becomes increasingly complex, especially with the introduction of height speakers, requiring innovative methods to capture and process audio data effectively.

Innovation Solution

The use of a pair of coincident, vertically-stacked directional microphones to capture azimuthal and elevation angles, allowing for the determination of sound source location and generation of audio objects with associated metadata, which can be rendered in complex playback environments with multiple speaker zones.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the number of channels and speaker configurations increases from 2D to 3D, then spatial audio rendering capability is improved, but system complexity increases

Engineering Contradiction:
Improvespatial audio rendering capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transitions from traditional 2D audio rendering to 3D spatial audio rendering by introducing elevation angle determination. The vertically-stacked microphone configuration enables capture of sound sources in three-dimensional space, adding the vertical dimension to the conventional horizontal plane audio field, thereby achieving immersive spatial audio reproduction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent decomposes the complex task of 3D sound source localization into distinct angular components: azimuth angle determination for horizontal positioning and elevation angle determination for vertical positioning. This segmentation allows independent processing of horizontal and vertical spatial information, simplifying the overall system architecture.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If vertically-stacked directional microphones are used to capture azimuthal and elevation angles, then sound source localization accuracy is improved, but microphone system complexity increases

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidmicrophone system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent employs a vertically-stacked microphone configuration where microphones are arranged along the vertical axis. This vertical stacking enables the system to capture both azimuthal information (through the directional pattern of individual microphones) and elevation information (through the vertical spatial distribution of microphone signals), achieving 3D sound source localization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The vertically-stacked directional microphone system serves multiple functions simultaneously: it captures azimuthal angles for horizontal sound source positioning and elevation angles for vertical positioning. This multi-functional design eliminates the need for separate microphone arrays for horizontal and vertical localization, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If audio objects with positional metadata are generated for complex playback environments, then sound positioning accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvesound positioning accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent separates the determination of sound source position into independent angular parameters: azimuth angle for horizontal positioning and elevation angle for vertical positioning. By processing these angular parameters independently and then combining them to form complete positional metadata, the system achieves accurate 3D sound positioning while managing processing complexity through modular computation.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables accurate spatial audio rendering in multi-channel environments, improving sound localization and immersion by providing detailed positional metadata for audio objects, even in 3D configurations, simplifying the process of sound positioning and reproduction.

Implementation Method 1

a pair of coincident, vertically-stacked directional microphones to capture azimuthal and elevation angles

Methodology Applied
Scientific EffectSound wave propagation: Sound

Implementation Method 2

determining, based at least in part on a temporal difference between the first microphone audio signals and the second microphone audio signals, an elevation angle

Methodology Applied
Scientific EffectTime difference of arrival: Time of Flight

Data Source

PatentEP3318070B1Determining azimuth and elevation angles from stereo recordings
Publication Date: 2024.05.22 DOLBY LABORATORIES LICENSING CORP
  • EP3318070B1 patent drawingFigure 1
  • EP3318070B1 patent drawingFigure 2
  • EP3318070B1 patent drawingFigure 3A~3B

AI summary

Input audio data, including first microphone audio signals and second microphone audio signals output by a pair of coincident, vertically-stacked directional microphones, may be received. An azimuthal angle corresponding to a sound source location may be determined, based at least in part on an intensity difference between the first microphone audio signals and the second microphone audio signals. An elevation angle corresponding to a sound source location may be determined, based at least in part on a temporal difference between the first microphone audio signals and the second microphone audio signals. Output audio data, including at least one audio object corresponding to a sound source, may be generated. The audio object may include audio object signals and associated audio object metadata. The audio object metadata may include at least audio object location data corresponding to the sound source location.