Speech Synthesis Noise Reduction for Natural Audio Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise reduction techniques, such as Spectral Subtraction and the Ephraim and Malah algorithm, suffer from musical artifacts and distort the speech when the signal-to-noise ratio decreases, making enhanced speech sound unnatural and preferring noisy speech over distorted 'enhanced' speech.
Innovation Solution
A system using speech recognition and synthesis that converts speech audio signals to text, generates synthetic speech based on a user's speech data corpus, and selectively transmits either the original or synthetic speech, whichever has higher predicted quality, determined by objective quality metrics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If spectral subtraction or Ephraim and Malah algorithm is used for noise reduction, then noise is reduced from speech signal, but the enhanced speech becomes distorted and unnatural when signal-to-noise ratio decreases
Solution Approach 1:
The patent creates a clean copy of the speech signal by synthesizing speech from text using a speech synthesis model. Instead of trying to clean the noisy signal directly, the system converts noisy speech to text, then generates a clean synthetic speech copy from that text, effectively replacing the distorted enhanced speech with a pristine synthetic version.
Solution Approach 2:
The patent replaces traditional signal processing mechanical systems (spectral subtraction, spectral amplitude estimation) with an information-based system using speech recognition and synthesis. This substitution transforms the problem from signal manipulation to linguistic processing, avoiding the artifacts inherent in spectral manipulation methods.
2Loss of energy
If traditional noise reduction techniques are applied, then noise power is reduced, but the speech becomes more distorted as signal-to-noise ratio decreases
Solution Approach 1:
The system generates a clean copy of the speech through synthesis from text representation. This synthetic copy has no noise contamination and maintains natural speech characteristics, replacing the need to manipulate the original noisy signal and its associated distortion risks.
Solution Approach 2:
The patent fundamentally changes the processing parameters from operating in the speech signal domain to operating in the text domain. By converting speech to text and back, the system changes the state of the information, allowing noise reduction without the parameter-related artifacts that plague direct signal processing methods.
3Object-affected harmful factors
If spectral subtraction is used to estimate and subtract noise power spectrum, then noise reduction is achieved, but musical artifacts are introduced
Solution Approach 1:
The patent replaces the mechanical spectral subtraction process with an information-theoretic approach using speech recognition and synthesis. This substitution eliminates the musical artifacts that arise from spectral manipulation because the system works with linguistic representations rather than frequency spectra.
Solution Approach 2:
Instead of subtracting noise from the spectrum (which creates artifacts), the system creates a clean synthetic copy from text. This copying approach bypasses the spectral manipulation entirely, avoiding the introduction of musical artifacts while still achieving noise reduction.
Data Source
AI summary
The present disclosure describes a system (100) for reducing background noise from a speech audio signal generated by a user. The system (100) includes a user device (102) receiving the speech audio signal, a noise reduction device (118) in communication with a stored data repository (208), where the noise reduction device is configured to convert the speech audio signal to text; generate synthetic speech based on the converted text; optionally determine the user as an actual subscriber based on a comparison between the speech audio signal with the synthetic speech; and selectively transmit the speech audio signal or the synthetic speech based on comparison between the predicted subjective quality of the recorded speech and the synthetic speech.


