Acoustic Processing Device Spatial Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sound source separation technologies face challenges in efficiently processing dynamic changes in the number and positions of sound sources in real acoustic environments, leading to high spatial complexity and reduced quality of target sound source acquisition.
Innovation Solution
An acoustic processing device and method that includes a spatial normalization unit to normalize the orientation component of a microphone array for a target direction, a mask function estimating unit using machine learning, and a mask processing unit to extract the target sound source component, reducing spatial complexity by employing steering vectors and space filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all spatial patterns of sound sources are considered in advance for sound source separation, then the quality of target sound source acquisition is improved, but the device complexity and processing effort increase enormously
Solution Approach 1:
The patent transforms the spatial coordinates of sound sources from arbitrary directions to a standardized coordinate system where all sound sources are mapped to a standard direction (e.g., 0 degrees). This parameter transformation allows the machine learning model to process all spatial patterns using a unified framework, improving generalization while reducing the complexity of handling each spatial pattern separately
Solution Approach 2:
The patent creates a universal processing framework that handles all spatial patterns of sound sources through a single machine learning model. By normalizing spatial orientations to a standard direction, the system achieves multi-functionality where one model can process any spatial configuration of sound sources, eliminating the need for separate models for each spatial pattern
2Productivity
If the number and positions of sound sources are set in advance, then the processing effort is reduced, but the reliability of target sound source separation deteriorates when sound sources dynamically change
Solution Approach 1:
The patent implements a dynamic processing framework where the system continuously adapts to changing sound source configurations. The spatial normalization process dynamically transforms any incoming sound source pattern to the standard coordinate system, allowing the machine learning model to reliably process dynamic changes in number and positions of sound sources without requiring pre-set configurations
Solution Approach 2:
The patent uses parameter transformation to map dynamic spatial configurations to a standardized form. By changing the coordinate system parameters and normalizing orientations, the system maintains processing efficiency while reliably handling dynamic sound source scenarios, as the machine learning model receives consistently formatted input regardless of the actual spatial configuration
Data Source
AI summary
A spatial normalization unit generates a normalized spectrum by normalizing an orientation component of a microphone array for a target direction included in a spectrum of an acoustic signal acquired from each of a plurality of microphones forming the microphone array into an orientation component for a predetermined standard direction. A mask function estimating unit determines a mask function used for extracting a component of a target sound source arriving in the target direction on the basis of the normalized spectrum using a machine learning model. A mask processing unit estimates the component of the target sound source installed in the target direction by applying the mask function to the acoustic signal.


