Audio Signal Processing Apparatus Voice Component Exclusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing methods degrade in separation performance when the voice signal is mixed with the audio signal, leading to incorrect production of the non-voice signal basis matrix.
Innovation Solution
An audio signal processing apparatus that includes an audio acquisition unit, likelihood calculation unit, spectral feature extraction unit, first basis matrix producing unit, and second basis matrix producing unit, which calculates and excludes components associated with the voice signal from the first basis matrix to produce a second basis matrix for nonnegative matrix factorization, improving separation performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If nonnegative matrix factorization is used to separate voice signal from audio signal, then sound source separation can be achieved, but the basis matrix of non-voice signal cannot be correctly produced when voice signal is mixed in the audio signal, resulting in degraded separation performance
Solution Approach 1:
The patent segments the basis matrix production process into two distinct stages: first producing a basis matrix from intervals with high non-voice probability, then using likelihood calculation to identify and exclude voice-containing components. This segmentation allows the system to handle mixed signals by processing different signal types separately and combining results.
Solution Approach 2:
The patent introduces likelihood calculation as an intermediary step between basis matrix production and final separation. The likelihood values serve as a mediator to identify which components of the basis matrix correspond to voice signals, enabling the system to exclude these components and produce a corrected basis matrix for accurate non-voice signal representation.
2Quantity of substance
If the basis matrix is produced from audio signal intervals with high non-voice probability, then non-voice signal components can be captured, but voice signal components are incorrectly included in the basis matrix, degrading separation accuracy
Solution Approach 1:
The patent extracts and removes voice signal components from the basis matrix by calculating likelihood values for each component and excluding those with high voice probability. This extraction process purifies the basis matrix, ensuring it contains only non-voice signal components while maintaining the complete set of non-voice signal characteristics.
Data Source
AI summary
Likelihood calculation means extracts audio features expressing features of a voice signal and a non-voice signal from an acquired audio signal, and calculates likelihood expressing probability that the voice signal is included in the audio signal using the audio features. Spectral feature extraction means performs a frequency analysis to the audio signal to extract a spectral feature. Using the spectral feature, first basis matrix producing means produces a first basis matrix expressing the feature of the non-voice signal. Second basis matrix producing means specifies a component having a high association with the voice signal in the first basis matrix using the likelihood, and excludes the component to produce a second basis matrix. Spectral feature estimation means estimates a spectral feature of the voice signal or a spectral feature of the non-voice signal by performing nonnegative matrix factorization to the spectral feature using the second basis matrix.


