Audio Signal Processing for Single-Microphone Noise Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital cameras with a single microphone face challenges in separating desired sound components from undesired noise, such as wind noise, due to the inability to apply techniques that require multiple microphone signals for sound source separation and noise reduction.
Innovation Solution
An audio signal processing apparatus that transforms mixed audio signals into time-frequency signals, divides them into bands, and uses nonnegative matrix factorization to separate sound components without prelearning, employing a teacher activity matrix to distinguish between desired and undesired sound components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple microphone signals are used for sound source separation and noise reduction, then the ability to separate desired sound from noise is improved, but the device complexity and cost increase
Solution Approach 1:
The patent segments the audio signal into multiple frequency bands and applies nonnegative matrix factorization separately to each band. This allows effective noise reduction in a single-channel system by processing different frequency components independently, achieving multi-microphone-like separation performance without requiring multiple physical microphones.
Solution Approach 2:
The patent transforms the audio signal from time domain to time-frequency domain representation, and processes different frequency bands with different parameters. By changing the representation domain and processing parameters according to frequency characteristics, the system achieves effective sound separation without increasing hardware complexity.
2Measurement precision
If nonnegative matrix factorization is applied to separate sound components, then noise reduction capability is improved, but the need for prelearned dictionaries increases device complexity
Solution Approach 1:
The patent employs blind nonnegative matrix factorization that automatically learns the basis matrices and activity matrices directly from the input signal without requiring prelearned dictionaries. The system serves itself by adapting to the specific characteristics of each input signal, eliminating the need for external prelearning resources and reducing device complexity.
Solution Approach 2:
The patent performs band division and preliminary signal processing before applying nonnegative matrix factorization. By preparing the signal in advance through frequency band segmentation and time-frequency transformation, the system enables effective noise reduction without requiring complex prelearned dictionaries, as the preliminary processing extracts sufficient structural information.
3Measurement precision
If clustering is performed to distinguish noise from desired sound, then noise reduction accuracy is improved, but the processing time and complexity increase
Solution Approach 1:
The patent extracts noise components and desired sound components separately through nonnegative matrix factorization, representing the mixed signal as a sum of nonnegative basis functions with time-varying activities. This extraction approach enables direct separation without time-consuming clustering operations, as the mathematical decomposition inherently isolates different sound sources based on their spectral characteristics.
Solution Approach 2:
The patent replaces the mechanical clustering process with a mathematical optimization approach using nonnegative matrix factorization. Instead of iteratively clustering separated signals to identify noise versus desired sound, the system uses constrained optimization to directly decompose the signal, significantly reducing processing time while maintaining or improving separation accuracy.
Data Source
AI summary
An audio signal in which a first audio component and a second audio component are mixed is inputted. The audio signal is transformed into a time-frequency signal representing relation between time and signal frequency. The time-frequency signal is divided into a plurality of bands. A time-frequency signal in a band, in which influence from the second audio component is small among the plurality of bands, is factorized into activity vectors composing a first activity matrix. The time-frequency signal transformed from the audio signal is factorized into activity vectors composing a second activity matrix using the first activity matrix as a teacher activity.


