Spatial-Region Neural Feature Extraction for Flexible Sound Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural network-based audio processing systems struggle with efficiently separating multiple localized sound sources due to the need for predefined spatial angles and computational inefficiencies, lacking versatility and flexibility in source selection.

Innovation Solution

A neural network architecture that incorporates spatial region-dependent feature extraction, allowing user-defined target regions at runtime, reduces computational redundancy by focusing on a specified spatial range, and integrates DOA-dependent layers to enhance separation efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the neural network is trained to extract signal components from all directions simultaneously, then the system can handle any target direction, but computational resources are wasted due to redundancy in output streams

Engineering Contradiction:
Improveability to handle any target directionVSAvoidcomputational efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically adapts the neural network's spatial processing based on the target direction parameter provided at runtime. Instead of a fixed architecture that processes all directions, the system configures the spatial processing dynamically to focus only on the relevant direction, reducing computational redundancy while maintaining versatility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by making the neural network's spatial processing specialized for a specific target direction rather than uniformly processing all directions. The spatial processing is localized to the region of interest, eliminating redundant computations for irrelevant directions while preserving the ability to handle any direction by changing the target parameter.

Inventive Principle:
Principle #3Local quality

2Productivity

If the neural network is trained for a fixed spatial angle, then computational resources are used efficiently, but the system lacks flexibility in selecting different angles at runtime

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidflexibility in angle selection
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system achieves dynamic adaptability by accepting a target direction as a runtime parameter and configuring the spatial processing accordingly. This allows the system to maintain computational efficiency for a specific angle while being flexible in selecting different angles when needed, resolving the contradiction between efficiency and adaptability.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multiple output streams are generated for all directions, then complete coverage of all possible targets is achieved, but the majority of output streams are unused and create redundancy

Engineering Contradiction:
Improvecoverage of all possible targetsVSAvoidredundancy in output data
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent extracts only the relevant spatial information corresponding to the target direction from the full set of possible directions. By taking out only the necessary component (target direction) and discarding the rest (other directions), the system eliminates redundancy while maintaining the ability to handle any target by changing the extraction parameter.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4171064B1Spatial dependent feature extraction in neural network based audio processing
Publication Date: 2025.09.17 GOODIX TECH HK CO LTD
  • EP4171064B1 patent drawingFigure 1~2
  • EP4171064B1 patent drawingFigure 3~4A
  • EP4171064B1 patent drawingFigure 4B

AI summary

A method for estimating target sound sources located in at least one target spatial region among a plurality of spatial regions by receiving a plurality of signals associated with one of a plurality of microphone signals comprising sound events generated by the plurality of sound sources, extracting via a neural network a plurality of features obtained by training the neural network for a different spatial region among the plurality of spatial regions, generating, by the processor, another plurality of features based on the extracted plurality of features wherein the another plurality of features corresponds to the at least one target spatial region, detecting or estimating, by the processor, at least one sound source among the target sound sources in the target spatial region based on the another plurality of features corresponding to the at least one target spatial region.