Signal Processing Apparatus for Occupancy Mask Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing signal processing techniques using neural networks struggle to accurately distinguish target sound from background noise when the sound source position or direction changes, requiring relearning of the network and significant processing resources.
Innovation Solution
A signal processing apparatus that acquires and integrates occupancy masks from multiple microphone groups, allowing for accurate determination of target and obstructing sounds without relearning the neural network, even when the number of channels or microphone positions change.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a neural network is used to determine target sound occupancy and noise occupancy for each time frequency point, then the extraction accuracy of target sound is improved, but the processing load and time required for relearning the network when sound source position or direction changes increases
Solution Approach 1:
The neural network is learned in advance with features from multiple microphone groups and sound source positions incorporated into the training data. This preliminary action enables the network to handle various sound source positions without requiring relearning when conditions change, thus reducing the time loss associated with retraining while maintaining high extraction accuracy.
2Adaptability or versatility
If the neural network is relearned when the number of channels or microphone positions changes, then the adaptability to different sound source positions is improved, but the processing load and time required for relearning increases
Solution Approach 1:
The neural network is designed to universally handle multiple sound source positions and microphone configurations by incorporating features from multiple microphone groups during training. This multi-functionality enables the same network to adapt to different sound source positions and channel configurations without requiring relearning, thus maintaining high adaptability while reducing processing load.
Solution Approach 2:
Multiple microphone group features and various sound source position data are incorporated into the training phase in advance. This preliminary action equips the neural network with the capability to handle diverse configurations, eliminating the need for relearning when the number of channels or microphone positions changes, thereby reducing processing load while maintaining versatility.
3Measurement precision
If features from multiple microphone groups are integrated, then the accuracy of determining target and obstructing sounds is improved, but the complexity of the signal processing system increases
Solution Approach 1:
Features from multiple microphone groups are merged and integrated into a unified neural network processing framework. This combining approach maintains high accuracy in determining target and obstructing sounds by leveraging information from all microphone groups, while the integrated architecture avoids the need for separate processing systems, thus managing complexity effectively.
Data Source
AI summary
A signal processing apparatus includes one or more processors. The processors acquire a plurality of observed signals acquired from a plurality of microphone groups each including at least one microphone selected from a plurality of microphones. The microphone groups include respective microphone combinations each including at least one microphone, the combinations are different from each other, and at least one of the microphone groups includes a plurality of microphones. The processors estimate a mask indicating occupancy for each of time frequency points of a sound signal of a space corresponding to the observed signal in a plurality of spaces, for each of the observed signals. The processors integrate masks estimated for the observed signals to generate an integrated mask indicating occupancy for each of time frequency points of a sound signal in a space determined based on the spaces.


