Acoustic Processing Device Spatial Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound source separation technologies face challenges in efficiently processing dynamic changes in the number and positions of sound sources in real acoustic environments, leading to high spatial complexity and reduced quality of target sound source acquisition.

Innovation Solution

An acoustic processing device and method that includes a spatial normalization unit to normalize the orientation component of a microphone array for a target direction, a mask function estimating unit using machine learning, and a mask processing unit to extract the target sound source component, reducing spatial complexity by employing steering vectors and space filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all spatial patterns of sound sources are considered in advance for sound source separation, then the quality of target sound source acquisition is improved, but the device complexity and processing effort increase enormously

Engineering Contradiction:
Improvequality of target sound source acquisitionVSAvoidspatial complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the spatial coordinates of sound sources from arbitrary directions to a standardized coordinate system where all sound sources are mapped to a standard direction (e.g., 0 degrees). This parameter transformation allows the machine learning model to process all spatial patterns using a unified framework, improving generalization while reducing the complexity of handling each spatial pattern separately

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal processing framework that handles all spatial patterns of sound sources through a single machine learning model. By normalizing spatial orientations to a standard direction, the system achieves multi-functionality where one model can process any spatial configuration of sound sources, eliminating the need for separate models for each spatial pattern

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If the number and positions of sound sources are set in advance, then the processing effort is reduced, but the reliability of target sound source separation deteriorates when sound sources dynamically change

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidquality of target sound source acquisition
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a dynamic processing framework where the system continuously adapts to changing sound source configurations. The spatial normalization process dynamically transforms any incoming sound source pattern to the standard coordinate system, allowing the machine learning model to reliably process dynamic changes in number and positions of sound sources without requiring pre-set configurations

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses parameter transformation to map dynamic spatial configurations to a standardized form. By changing the coordinate system parameters and normalizing orientations, the system maintains processing efficiency while reliably handling dynamic sound source scenarios, as the machine learning model receives consistently formatted input regardless of the actual spatial configuration

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11818557B2Acoustic processing device including spatial normalization, mask function estimation, and mask processing, and associated acoustic processing method and storage medium
Publication Date: 2023.11.14 HONDA MOTOR CO LTD
  • US11818557B2 patent drawing
  • US11818557B2 patent drawing
  • US11818557B2 patent drawing

AI summary

A spatial normalization unit generates a normalized spectrum by normalizing an orientation component of a microphone array for a target direction included in a spectrum of an acoustic signal acquired from each of a plurality of microphones forming the microphone array into an orientation component for a predetermined standard direction. A mask function estimating unit determines a mask function used for extracting a component of a target sound source arriving in the target direction on the basis of the normalized spectrum using a machine learning model. A mask processing unit estimates the component of the target sound source installed in the target direction by applying the mask function to the acoustic signal.