Impulse Noise Correction in Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

First responders face delays and increased processing demands due to impulse noise interfering with speech recognition systems, causing incomplete utterance recognition and repeated requests in noisy environments, which can be critical in emergency situations.

Innovation Solution

The system detects impulse noise in utterances and generates clarifying voice prompts to confirm or repeat affected words, reducing the need for users to repeat entire utterances, thereby improving interaction and resource efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech recognition systems process utterances in noisy environments, then the system can provide information and services to first responders, but impulse noise causes incomplete utterance recognition and requires repeated requests, increasing processing demands and delays

Engineering Contradiction:
Improveutterance recognition accuracyVSAvoidprocessing delays
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by detecting impulse noise and generating clarifying voice prompts before the user needs to repeat their entire utterance. The electronic processor identifies portions of utterances affected by impulse noise and proactively requests clarification on specific words or phrases, preventing incomplete recognition from propagating through the system and reducing subsequent reprocessing needs.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the system requests users to repeat entire utterances due to impulse noise, then recognition accuracy may be maintained, but communication efficiency decreases and battery resources are consumed

Engineering Contradiction:
Improverecognition accuracyVSAvoidbattery consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system applies segmentation by dividing the utterance into discrete portions and identifying specific segments affected by impulse noise. Instead of requiring repetition of the entire utterance, the electronic processor generates clarifying voice prompts that target only the noisy segments or specific words/phrases that were missed, reducing the amount of audio data that needs to be reprocessed and thereby conserving battery resources.

Inventive Principle:
Principle #1Segmentation

3Productivity

If the system processes every utterance completely despite impulse noise, then no clarification requests are needed, but critical information may be missed and processing resources are wasted

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidmissed critical information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system implements feedback by monitoring the output of speech recognition and detecting when impulse noise has degraded utterance quality. The electronic processor analyzes recognition confidence and audio quality metrics, then provides feedback in the form of targeted clarifying voice prompts for specific portions of the utterance. This closed-loop approach ensures critical information is captured while avoiding unnecessary reprocessing of clear segments.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10811011B2Correcting for impulse noise in speech recognition systems
Publication Date: 2020.10.20 MOTOROLA SOLUTIONS INC
  • US10811011B2 patent drawing
  • US10811011B2 patent drawing
  • US10811011B2 patent drawing

AI summary

System and method for correcting for impulse noise in speech recognition systems. One example system includes a microphone, a speaker, and an electronic processor. The electronic processor is configured to receive an audio signal representing an utterance. The electronic processor is configured to detect, within the utterance, the impulse noise, and, in response, generate an annotated utterance including a timing of the impulse noise. The electronic processor is configured to segment the annotated utterance into silence, voice content, and other content, and, when a length of the other content is greater than or equal to an average word length for the annotated utterance, determine, based on the voice content, an intent portion and an entity portion. The electronic processor is configured to generate a voice prompt based on the timing of the impulse noise and the intent portion and/or the entity portion, and to play the voice prompt.