Voice Correction Apparatus for Low SNR Band Distortion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems face challenges in increasing recognition accuracy for voice signals that have undergone noise suppression processing, particularly in environments with low signal-to-noise ratios, as current methods often introduce distortion and fail to adequately correct sound quality in bands with low signal noise ratios.
Innovation Solution
A voice correction method that involves obtaining voice information from both noisy and ideal environments, emphasizing components in bands with low signal noise ratios, and using machine learning to generate a model that corrects noise-suppressed signals, prioritizing weighting for bands with low signal noise ratios to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If noise suppression processing is performed on voice signals, then noise reduction is achieved, but distortion occurs in low SNR bands resulting in unnatural voice and decreased voice recognition rate
Solution Approach 1:
The patent applies different processing strategies to different frequency bands based on their SNR characteristics. Low SNR bands undergo noise suppression with correction, while high SNR bands are preserved with minimal processing. This localized approach allows noise reduction in problematic bands without degrading the quality of already-clean bands.
Solution Approach 2:
The system dynamically adjusts the noise suppression strength and correction parameters based on the measured SNR of each frequency band. By changing processing parameters adaptively rather than applying fixed suppression levels, the system achieves effective noise reduction while maintaining voice naturalness in varying acoustic conditions.
2Object-affected harmful factors
If noise suppression processing is performed, then noise is reduced, but voice recognition rate decreases due to distortion in low SNR bands
Solution Approach 1:
The patent applies different processing strategies to different frequency bands based on their SNR characteristics. Low SNR bands undergo noise suppression with correction, while high SNR bands are preserved with minimal processing. This localized approach allows noise reduction in problematic bands without degrading the quality of already-clean bands.
Solution Approach 2:
The system uses the measured SNR of each frequency band as feedback to adjust the noise suppression and correction parameters. This closed-loop approach allows the system to optimize the balance between noise reduction and voice quality preservation based on real-time acoustic conditions, thereby maintaining higher voice recognition rates.
3Measurement precision
If machine learning is used to correct sound quality, then recognition accuracy improves, but processing complexity increases
Solution Approach 1:
The patent divides the voice signal processing into distinct stages: noise suppression, SNR estimation, distortion correction, and voice recognition. By segmenting the processing pipeline, the system can apply machine learning selectively at critical stages rather than throughout the entire process, reducing overall complexity while maintaining accuracy.
Solution Approach 2:
The system performs noise suppression and SNR estimation before voice recognition, preparing the signal in advance. This preliminary processing reduces the complexity of the subsequent recognition task by pre-correcting obvious distortions and organizing the signal characteristics, allowing the recognition algorithm to focus on pattern matching rather than basic signal cleaning.
Data Source
AI summary
A voice correction method implemented by a computer, the method includes: obtaining first voice information which is voice information recorded when noise is generated and on which noise suppression processing is performed and second voice information indicating voice information recorded in an environment in which no noise is generated, and generating emphasized information by emphasizing a component of a band corresponding to a band having a low signal noise ratio (SNR) of the first voice information, among bands of the second voice information; performing machine learning on a model based on the first voice information and the emphasized information; and generating corrected voice information by correcting third voice information on which noise suppression processing is performed, based on the machine-learned model.


