Speech Signal Reconstruction via Distortion Recovery Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately separating speech signals from noise due to the limited number of microphones in a microphone array, leading to impaired speech signals and reduced recognition accuracy.
Innovation Solution
A method and terminal for reconstructing speech signals using a distortion recovery model trained with generative adversarial networks, which separates and reconstructs speech signals to minimize distortion, improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If frequency-domain Wiener filtering is applied to separate speech signals from noise signals, then noise elimination is improved, but speech signal quality deteriorates
Solution Approach 1:
The patent replaces traditional frequency-domain Wiener filtering (a mathematical/mechanical processing system) with a deep neural network-based speech separation system. The DNN model learns optimal separation strategies during training and performs speech-enhanced signal separation in the time domain, substituting the conventional filtering approach with an intelligent, adaptive system that preserves speech quality while eliminating noise.
Solution Approach 2:
The patent transforms the speech separation problem from frequency-domain filtering to time-domain processing using DNN. By changing the domain and the parameters used for separation (from fixed filter coefficients to dynamic neural network outputs), the system achieves better speech quality preservation while maintaining noise elimination capability.
2Measurement precision
If the number of microphones in the microphone array is increased, then speech signal separation capability is improved, but device complexity increases
Solution Approach 1:
The patent substitutes the mechanical solution of increasing microphone array size with an intelligent software-based DNN speech separation system. Instead of adding more physical sensors to improve separation capability, the system uses a trained neural network to achieve high-quality speech separation with the existing limited number of microphones, thereby avoiding increased device complexity.
3Object-affected harmful factors
If traditional speech separation filtering is applied, then noise removal is improved, but speech recognition accuracy deteriorates
Solution Approach 1:
The patent replaces traditional speech separation filtering with a DNN-based speech enhancement system that outputs speech-enhanced signals optimized for speech recognition. The neural network is trained to preserve speech characteristics that are important for recognition while removing noise, thereby improving both noise removal and maintaining speech recognition accuracy.
Solution Approach 2:
The patent introduces a DNN-based speech enhancement module as an intermediary between the microphone array and the speech recognition system. This intermediary processes the raw speech signals to remove noise while preserving recognition-critical features, acting as a bridge that reconciles the conflicting requirements of noise removal and recognition accuracy.
Data Source
AI summary
The present disclosure discloses a method performed at a terminal for reconstructing a speech signal, and a computer storage medium, and relates to the field of speech recognition. The method includes: collecting, by the terminal, a plurality of sound signals through a plurality of sensors of a microphone array; determining, by the terminal, a first speech signal in the plurality of sound signals; performing, by the terminal, signal separation on the first speech signal to obtain a second speech signal; and performing, by the terminal, reconstruction on the second speech signal through a distortion recovery model to obtain a reconstructed speech signal; the distortion recovery model being obtained by training based on a clean speech signal and a distorted speech signal. The embodiments of the present disclosure improve accuracy of speech recognition results.


