Split-Domain Speech Signal Enhancement via Frequency Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile devices equipped with laser microphones struggle to capture high-quality speech signals due to noise interference from background noise, surface materials, and vibrations, leading to low signal-to-noise ratios.
Innovation Solution
The method involves receiving input signals with both noise and speech components, performing linear predictive filtering to generate coefficients and residual signals, converting these signals into frequency domains to estimate magnitude and phase spectra, and synthesizing output signals based on these spectra to enhance speech quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional noise suppression methods (spectral subtraction, Wiener filtering) are used, then computation complexity is reduced, but speech signal quality and intelligibility improvement is limited
Solution Approach 1:
The speech signal is segmented into multiple frequency bins through Fourier transformation, allowing independent processing of different frequency components. This segmentation enables selective enhancement of speech frequencies while suppressing noise frequencies, resolving the contradiction by providing targeted processing that improves quality without requiring globally complex algorithms.
Solution Approach 2:
The patent transforms the signal from time domain to frequency domain, changing the parameter space for noise suppression. By working in the frequency domain, the system can selectively modify magnitude and phase parameters of different frequency components, achieving better speech quality with more efficient computation compared to time-domain methods.
2Reliability
If deep neural network or non-negative matrix factorization algorithms are used, then speech signal quality and intelligibility are improved, but computation complexity increases significantly
Solution Approach 1:
The patent extracts the magnitude and phase components of the speech signal separately through Fourier transformation. By taking out these components independently, the system can apply simple magnitude scaling and phase preservation operations, achieving effective noise suppression without requiring complex neural network computations.
Solution Approach 2:
The patent replaces complex machine learning algorithms (deep neural networks, non-negative matrix factorization) with simpler signal processing operations in the frequency domain. By substituting mechanical/mathematical transformations (Fourier transform, magnitude scaling) for computational intensive learning-based methods, the system achieves comparable quality improvement with reduced computation complexity suitable for mobile devices.
3Ease of operation
If laser microphone captures audio based on surface vibrations, then portable and contactless audio capture is achieved, but noise from background, surface material, and vibrations reduces signal quality
Solution Approach 1:
The patent applies different processing strategies to different frequency bins based on their local characteristics. Speech frequencies receive enhancement through magnitude scaling, while noise frequencies are suppressed. This local quality approach resolves the contradiction by improving signal quality in specific frequency regions without compromising the portable contactless capture capability.
Solution Approach 2:
The patent converts the harmful noise interference into a separable component in the frequency domain. By transforming to frequency domain, the noise and speech components become distinguishable, allowing the system to selectively enhance speech frequencies while suppressing noise frequencies, thus converting the harmful mixed signal into a beneficial separable structure.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively improves the intelligibility and quality of speech signals by separating noise from speech components, resulting in enhanced speech signal processing even in noisy environments.
Implementation Method 1
the microphone may direct the light beam to a surface that is proximate to a sound source, and vibrations of the surface, caused by sound waves from the sound source, may change properties of the reflected light beam. For example, the vibrations of the surface may change a frequency of the light beam and a phase of the light beam.
Data Source
AI summary
A method and an apparatus for estimating speech signal in split-domain is disclosed. The method includes performing LP analysis on a noisy speech signal to generate a first plurality of LPC and a first residual signal. The method also includes estimating speech LPC spectrum to generate cleaned LPC. The method further includes estimating speech residual spectrum to generate cleaned residual signal. The method also includes synthesizing output signals based on the cleaned LPC and the cleaned residual signal.


