Directional Audio Component Extraction for Source Categorization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio analysis methods for multi-channel signals, particularly in video game applications, face challenges in accurately categorizing and characterizing individual audio sources in scenarios where spatial audio information is limited, such as with single or stereo setups, leading to reduced user experience and information availability.

Innovation Solution

An apparatus and method that includes a receiver for multi-channel audio signals, an extractor for directional audio components using spatial filtering, a feature processor for determining features, a categorizer for assigning audio source categories, and an assigner for linking properties to these categories, allowing for improved audio source characterization and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multi-channel audio rendering is used to provide spatial audio information, then audio localization accuracy is improved, but device complexity increases

Engineering Contradiction:
Improveaudio localization accuracyVSAvoidspatial audio rendering system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts directional audio components from the multi-channel audio signal by applying spatial filtering. This extraction process isolates specific directional information from the mixed audio signal, allowing the system to obtain accurate localization data without requiring complex multi-channel rendering hardware. The extracted directional components can be processed independently to provide localization accuracy while reducing the complexity burden on the audio rendering system.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If spatial filtering is applied to extract directional audio components, then audio source characterization accuracy is improved, but processing time increases

Engineering Contradiction:
Improveaudio source characterization accuracyVSAvoidaudio processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the audio signal processing into distinct stages: first applying spatial filtering to extract directional components, then separately processing the extracted components for characterization. This segmentation allows the spatial filtering to be optimized for accuracy while the subsequent processing can be optimized for speed. The directional audio components are extracted as discrete elements that can be processed independently, reducing the overall processing time compared to analyzing the complete mixed signal.

Inventive Principle:
Principle #1Segmentation

3Loss of information

If directional audio components are extracted and analyzed, then information availability about audio sources is improved, but device complexity increases

Engineering Contradiction:
Improveaudio source information availabilityVSAvoidaudio analysis system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces directional audio components as intermediary representations between the raw multi-channel audio signal and the final audio source characterization. These extracted directional components serve as mediators that contain the essential localization information while being easier to process and analyze. By using these intermediary components, the system can provide rich audio source information without directly processing the complex multi-channel signal throughout the entire analysis pipeline.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11956616B2Apparatus and method for audio analysis
Publication Date: 2024.04.09 GN STORE NORD AS
  • US11956616B2 patent drawing
  • US11956616B2 patent drawing
  • US11956616B2 patent drawing

AI summary

An apparatus comprises a receiver (201) receiving a multi-channel audio signal representing audio for a scene. An extractor (203) extracts at least one directional audio component by applying a spatial filtering to the multi-channel signal where the spatial filtering is dependent on the multi-channel audio signal. A feature processor (205) determines a set of features for the first directional audio component and a categorizer (207) determines a first audio source category out of a plurality of audio source categories for the directional audio signal in response to the set of features. An assigner (209) assigns a first audio source property to the first directional audio component from a set of audio source properties for the first audio source category. The apparatus may provide very advantageous categorization and characterization of individual audio sources/components present in a multi-channel signal. This may be advantageous e.g. for visualization of audio events.