Spatial Audio Direction Identification Using Coherence Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio processing technologies face challenges in accurately identifying multiple directions of sound arrival from audio signals captured by microphones, leading to artifacts in the rendered audio output, especially in devices with restricted microphone configurations like mobile phones.

Innovation Solution

An apparatus and method using processing circuitry and memory circuitry to identify multiple directions of sound arrival by analyzing delay parameters and energy ratios between audio signals captured by a microphone array, employing coherence analysis in the time-frequency domain to determine directional information and energy parameters for each frequency band, and providing metadata for synthesizing spatial audio signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing spatial audio processing technologies are used to identify directions of sound arrival, then the processing can be performed with simple microphone configurations, but the accuracy of identifying multiple directions is poor leading to artifacts in rendered audio

Engineering Contradiction:
Improveaccuracy of identifying multiple directions of sound arrivalVSAvoidartifacts in rendered audio output
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent segments the audio signal processing into multiple frequency bands, performing coherence analysis separately for each band. This allows accurate identification of multiple directions of sound arrival by analyzing different frequency components independently, resolving the contradiction between measurement precision and artifact generation in restricted microphone configurations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of analysis by using coherence analysis in the time-frequency domain with varying time delays. By adjusting delay parameters and analyzing coherence across different frequency bands, the system accurately identifies multiple sound directions even with limited microphones, thereby improving measurement precision while reducing artifacts

Inventive Principle:
Principle #35Parameter changes

2Reliability

If multiple directions of sound arrival are accurately identified using coherence analysis in time-frequency domain, then spatial audio quality is improved, but the processing complexity increases

Engineering Contradiction:
Improvespatial audio qualityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the complex processing task into segmented frequency bands, where coherence analysis is performed independently for each band. This segmentation approach improves spatial audio quality through accurate multi-direction identification while managing processing complexity by breaking down the overall computation into smaller, more manageable frequency-specific operations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies coherence analysis selectively across different frequency bands rather than uniformly across the entire spectrum. By focusing computational resources on specific frequency ranges where directional information is most useful, the system achieves high spatial audio quality while avoiding unnecessary processing complexity in less critical frequency regions

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11950063B2Apparatus, method and computer program for audio signal processing
Publication Date: 2024.04.02 NOKIA TECHNOLOGIES OY
  • US11950063B2 patent drawing
  • US11950063B2 patent drawing
  • US11950063B2 patent drawing

AI summary

A device may be configured to: obtain at least a first audio signal and a second audio signal, wherein the first audio signal and the second audio signal are captured with a microphone array comprising at least two microphones; identify a first direction for a plurality of frequency bands; and identify a second direction for the plurality of frequency bands; wherein the first direction and the second direction are identified based, at least partially, on spatial analysis of the first audio signal and the second audio signal.