Dual Noise Suppression for Speech Recognition Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound processing devices face high calculation costs and reduced speech recognition accuracy due to the complexity of estimating sound source directions and separating sound signals from multiple channels, especially in noisy environments.

Innovation Solution

A sound processing device with dual noise suppression units and a speech section detection unit to identify and isolate speech sections within noise-removed signals, allowing for reduced distortion and improved speech recognition rates by selectively performing speech recognition on the most intense channel or separated sound sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If sound source directions are estimated and sound signals are separated from multiple channels, then speech recognition accuracy is improved, but calculation cost increases significantly

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcalculation cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the sound signal processing into distinct stages: noise component extraction, speech section detection, and speech recognition. By dividing the processing pipeline and applying different operations to different sections, the system reduces overall calculation cost while maintaining recognition accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies noise suppression selectively only to detected speech sections rather than processing the entire sound signal continuously. This partial action approach reduces calculation cost by avoiding unnecessary processing in non-speech periods while maintaining speech recognition accuracy.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If sound source directions are estimated regardless of the number of sound sources, then speech separation is achieved, but processing time increases

Engineering Contradiction:
Improvespeech separation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent dynamically adapts the processing approach based on the detected number of sound sources. When multiple sound sources are detected, the system applies appropriate separation techniques; when a single source is detected, simpler processing is used. This dynamic adaptation reduces processing time while maintaining separation accuracy when needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes processing parameters based on the detected speech sections and sound source characteristics. By adjusting the processing intensity and methods according to the specific situation (number of speakers, speech activity), the system optimizes processing time without sacrificing separation accuracy.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If noise suppression is applied with high suppression amount, then noise removal is improved, but speech distortion increases

Engineering Contradiction:
Improvenoise removal effectivenessVSAvoidspeech signal quality
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The patent applies different noise suppression amounts to different time sections based on speech activity detection. During speech sections, moderate suppression is applied to maintain quality; during non-speech sections, stronger suppression removes noise more effectively. This local differentiation resolves the contradiction between noise removal and speech quality.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent applies strong noise suppression only partially during non-speech sections and moderate suppression during speech sections. This selective application of different suppression levels achieves effective noise removal without excessively distorting speech signals.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9384760B2Sound processing device and sound processing method
Publication Date: 2016.07.05 HONDA MOTOR CO LTD
  • US9384760B2 patent drawing
  • US9384760B2 patent drawing
  • US9384760B2 patent drawing

AI summary

A sound processing device includes a first noise suppression unit configured to suppress a noise component included in an input sound signal using a first suppression amount, a second noise suppression unit configured to suppress the noise component included in the input sound signal using a second suppression amount greater than the first suppression amount, a speech section detection unit configured to detect whether the sound signal whose noise component has been suppressed by the second noise suppression unit includes a speech section having a speech for every predetermined time, and a speech recognition unit configured to perform a speech recognizing process on a section, which is detected to be a speech section by the speech section detection unit, in the sound signal whose noise component has been suppressed by the first noise suppression unit.