Multiway Speech Recognition in Low Signal-to-Noise Ratio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Far-field speech recognition accuracy is significantly reduced in noisy environments due to low signal-to-noise ratios, as spatial filtering depends on sound source localization, which is also impaired in such conditions.
Innovation Solution
A method that separates input audio signals into multiple signals, generates a denoised signal based on primary and secondary signals, and performs preliminary recognition using a combination of these signals, integrating array signal processing with speech recognition to enhance recognition rates even in low signal-to-noise ratios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If spatial filtering is performed through a microphone array based on sound source localization, then speech recognition can be performed in noisy environments, but when the signal-to-noise ratio is low, the accuracy of sound source localization is significantly reduced, leading to reduced recognition rate
Solution Approach 1:
The patent segments the speech recognition process into multiple independent recognition channels (original microphone array signal, separated speech signal, and denoised signal), each processed through different signal processing paths. This segmentation allows the system to maintain multiple recognition streams that can be independently evaluated and combined, preventing the degradation of one channel from completely compromising the overall recognition reliability in low signal-to-noise ratio conditions.
Solution Approach 2:
The patent applies parameter changes by transforming the input signal through different processing operations to create multiple versions with different characteristics. Specifically, it separates the mixed audio signal into speech and noise components, generates denoised signals with adjusted noise levels, and processes these through different recognition models. This parameter transformation allows the system to adapt to varying noise conditions and maintain recognition accuracy when traditional localization methods fail.
2Device complexity
If a single-channel speech signal is output after spatial filtering, then the system complexity is reduced, but the recognition rate is greatly reduced in noisy environments with low signal-to-noise ratio
Solution Approach 1:
The patent merges multiple recognition results from different signal processing channels (original spatially filtered signal, separated speech signal, and denoised signal) into a unified recognition output. By combining the strengths of different processing approaches, the system achieves improved recognition reliability in noisy environments without requiring an excessively complex individual processing chain for each channel.
Solution Approach 2:
The patent creates a multi-functional speech recognition system that can handle both clean and noisy speech conditions through a single integrated framework. The system processes the same input signal through multiple pathways (spatial filtering, blind source separation, denoising) and uses a unified recognition model that can adaptively utilize the most appropriate signal representation, making the system universally applicable across different noise conditions without requiring separate specialized systems.
Data Source
AI summary
Disclosed is a method and an apparatus for recognizing speech, and the method comprises: separating an input audio signal into at least two separated signals; generating a denoised signal at a current frame; performing a preliminary recognition on each interesting signal at the current frame; and performing a recognition decision according to a recognition score of each interesting signal at the current frame. The method and apparatus of the present disclosure deeply integrate an array signal processing and a speech recognition and use multiway recognitions such that a good recognition rate may be obtained even in a case of a low signal-to-noise ratio.


