Signal processing device, signal processing method, signal processing program, learning device, learning method, and learning program
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal extraction technologies are limited to speeches of persons and face increased calculation loads with multiple audio classes, lacking the ability to efficiently extract desired audio signals from mixed audio signals of various classes without proportional calculation increases.
Innovation Solution
A signal processing device using a neural network to extract desired audio signals from mixed audio signals, employing an auxiliary and main neural network to embed and transform audio features based on user-defined extraction targets, maintaining a constant calculation load regardless of the number of classes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional audio signal extraction technologies are used to extract multiple audio classes from mixed audio signals, then the extraction capability is improved, but the calculation amount increases proportionally with the number of audio classes
Solution Approach 1:
The patent segments the audio signal processing into distinct neural network components: an auxiliary neural network for feature extraction and a main neural network for signal reconstruction. This segmentation allows the system to process multiple audio classes efficiently by dividing the computational workload across specialized network modules rather than using a single monolithic processing approach.
Solution Approach 2:
The auxiliary neural network performs preliminary action by extracting and embedding audio features before the main neural network processes the signal reconstruction. This preliminary feature extraction stage prepares the data in advance, enabling the main network to focus on reconstruction tasks and reducing the overall calculation amount required for multi-class extraction.
2Adaptability or versatility
If conventional technologies are extended to support audio signals other than speeches of persons, then the applicability is improved, but the calculation complexity increases
Solution Approach 1:
The patent implements universality by designing a neural network architecture that can handle multiple audio classes including but not limited to speech signals. The auxiliary and main neural networks are configured to process various audio types (speech, environmental sounds, etc.) through unified feature extraction and reconstruction mechanisms, enabling one system to serve multiple functions without proportionally increasing complexity.
3Adaptability or versatility
If the number of audio classes to be extracted is increased, then the extraction comprehensiveness is improved, but the calculation amount increases proportionally
Solution Approach 1:
The patent merges the processing of multiple audio classes into a unified neural network framework where the auxiliary network extracts features for all classes simultaneously and the main network reconstructs all target signals in parallel. This combining approach allows comprehensive extraction of multiple audio classes while maintaining constant calculation efficiency by avoiding sequential processing of each class.
Data Source
AI summary
A signal processing device includes processing circuitry configured to receive an input of extraction target information indicating which audio class of an audio signal is to be extracted from a mixture audio signal constituted by a mixture of audio signals of a plurality of audio classes, and output a result of extracting the audio signal of the audio class indicated by the extraction target information from the mixture audio signal, with a neural network by using a feature value of the mixture audio signal and the extraction target information.


