Multiway Speech Recognition in Low Signal-to-Noise Ratio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Far-field speech recognition accuracy is significantly reduced in noisy environments due to low signal-to-noise ratios, as spatial filtering depends on sound source localization, which is also impaired in such conditions.

Innovation Solution

A method that separates input audio signals into multiple signals, generates a denoised signal based on primary and secondary signals, and performs preliminary recognition using a combination of these signals, integrating array signal processing with speech recognition to enhance recognition rates even in low signal-to-noise ratios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If spatial filtering is performed through a microphone array based on sound source localization, then speech recognition can be performed in noisy environments, but when the signal-to-noise ratio is low, the accuracy of sound source localization is significantly reduced, leading to reduced recognition rate

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidsound source localization accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech recognition process into multiple independent recognition channels (original microphone array signal, separated speech signal, and denoised signal), each processed through different signal processing paths. This segmentation allows the system to maintain multiple recognition streams that can be independently evaluated and combined, preventing the degradation of one channel from completely compromising the overall recognition reliability in low signal-to-noise ratio conditions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies parameter changes by transforming the input signal through different processing operations to create multiple versions with different characteristics. Specifically, it separates the mixed audio signal into speech and noise components, generates denoised signals with adjusted noise levels, and processes these through different recognition models. This parameter transformation allows the system to adapt to varying noise conditions and maintain recognition accuracy when traditional localization methods fail.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If a single-channel speech signal is output after spatial filtering, then the system complexity is reduced, but the recognition rate is greatly reduced in noisy environments with low signal-to-noise ratio

Engineering Contradiction:
Improvesystem complexityVSAvoidrecognition rate
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent merges multiple recognition results from different signal processing channels (original spatially filtered signal, separated speech signal, and denoised signal) into a unified recognition output. By combining the strengths of different processing approaches, the system achieves improved recognition reliability in noisy environments without requiring an excessively complex individual processing chain for each channel.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a multi-functional speech recognition system that can handle both clean and noisy speech conditions through a single integrated framework. The system processes the same input signal through multiple pathways (spatial filtering, blind source separation, denoising) and uses a unified recognition model that can adaptively utilize the most appropriate signal representation, making the system universally applicable across different noise conditions without requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11183179B2Method and apparatus for multiway speech recognition in noise
Publication Date: 2021.11.23 NANJING HORIZON ROBOTICS TECH CO LTD
  • US11183179B2 patent drawing
  • US11183179B2 patent drawing
  • US11183179B2 patent drawing

AI summary

Disclosed is a method and an apparatus for recognizing speech, and the method comprises: separating an input audio signal into at least two separated signals; generating a denoised signal at a current frame; performing a preliminary recognition on each interesting signal at the current frame; and performing a recognition decision according to a recognition score of each interesting signal at the current frame. The method and apparatus of the present disclosure deeply integrate an array signal processing and a speech recognition and use multiway recognitions such that a good recognition rate may be obtained even in a case of a low signal-to-noise ratio.