Audio Signal Processing Apparatus Voice Component Exclusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing methods degrade in separation performance when the voice signal is mixed with the audio signal, leading to incorrect production of the non-voice signal basis matrix.

Innovation Solution

An audio signal processing apparatus that includes an audio acquisition unit, likelihood calculation unit, spectral feature extraction unit, first basis matrix producing unit, and second basis matrix producing unit, which calculates and excludes components associated with the voice signal from the first basis matrix to produce a second basis matrix for nonnegative matrix factorization, improving separation performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If nonnegative matrix factorization is used to separate voice signal from audio signal, then sound source separation can be achieved, but the basis matrix of non-voice signal cannot be correctly produced when voice signal is mixed in the audio signal, resulting in degraded separation performance

Engineering Contradiction:
Improveseparation performanceVSAvoidaccuracy of basis matrix production
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the basis matrix production process into two distinct stages: first producing a basis matrix from intervals with high non-voice probability, then using likelihood calculation to identify and exclude voice-containing components. This segmentation allows the system to handle mixed signals by processing different signal types separately and combining results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces likelihood calculation as an intermediary step between basis matrix production and final separation. The likelihood values serve as a mediator to identify which components of the basis matrix correspond to voice signals, enabling the system to exclude these components and produce a corrected basis matrix for accurate non-voice signal representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If the basis matrix is produced from audio signal intervals with high non-voice probability, then non-voice signal components can be captured, but voice signal components are incorrectly included in the basis matrix, degrading separation accuracy

Engineering Contradiction:
Improvenon-voice signal componentsVSAvoidpurity of basis matrix
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts and removes voice signal components from the basis matrix by calculating likelihood values for each component and excluding those with high voice probability. This extraction process purifies the basis matrix, ensuring it contains only non-voice signal components while maintaining the complete set of non-voice signal characteristics.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9224392B2Audio signal processing apparatus and audio signal processing method
Publication Date: 2015.12.29 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9224392B2 patent drawing
  • US9224392B2 patent drawing
  • US9224392B2 patent drawing

AI summary

Likelihood calculation means extracts audio features expressing features of a voice signal and a non-voice signal from an acquired audio signal, and calculates likelihood expressing probability that the voice signal is included in the audio signal using the audio features. Spectral feature extraction means performs a frequency analysis to the audio signal to extract a spectral feature. Using the spectral feature, first basis matrix producing means produces a first basis matrix expressing the feature of the non-voice signal. Second basis matrix producing means specifies a component having a high association with the voice signal in the first basis matrix using the likelihood, and excludes the component to produce a second basis matrix. Spectral feature estimation means estimates a spectral feature of the voice signal or a spectral feature of the non-voice signal by performing nonnegative matrix factorization to the spectral feature using the second basis matrix.