Microphone Initialization for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional microphone initialization techniques in speaker phones and mobile phone car installations suppress the microphone during audio playback, leading to loss of low-volume speech sounds at the beginning of an utterance, resulting in truncated recognition and potential misinterpretation of spoken phrases.

Innovation Solution

A method and system that detect the transition of the microphone switching on and perform speech recognition with reduced contribution from the initial portion of the signal, allowing for optional deletion of initial sounds or inclusion of a redundant word to ensure accurate recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-generated harmful factors

If the microphone is suppressed during audio playback to eliminate echo and feedback, then echo and feedback are eliminated, but low-volume speech sounds at the beginning of an utterance are lost

Engineering Contradiction:
Improveecho and feedbackVSAvoidlow-volume speech sounds
Core Design Contradiction:
Object-generated harmful factorsVSLoss of information

Solution Approach 1:

The system dynamically adjusts the microphone suppression state by detecting transitions in the audio signal. When a high-energy signal is detected indicating the user has started speaking, the system switches the microphone from suppressed to active state, allowing low-volume speech sounds to be captured while maintaining echo suppression during playback

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary detection of high-energy audio signals before fully activating the microphone. This preliminary action allows the system to prepare for speech capture in advance, ensuring that low-volume speech sounds at the beginning of an utterance are not lost while maintaining echo suppression during the transition period

Inventive Principle:
Principle #10Preliminary action

2Object-generated harmful factors

If the microphone remains off until high energy audio signal is received, then echo and feedback are suppressed, but the spoken phrase is truncated and recognition fails

Engineering Contradiction:
Improveecho and feedback suppressionVSAvoidspeech recognition accuracy
Core Design Contradiction:
Object-generated harmful factorsVSReliability

Solution Approach 1:

The system extracts and processes only the relevant portion of the speech signal for recognition. By detecting the microphone switch-on transition, the system identifies where speech actually begins and extracts this portion for recognition, effectively removing the problematic initial truncated portion while maintaining recognition reliability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameter of microphone activation from a static high-energy threshold to a dynamic transition-detection mechanism. This parameter change allows the microphone to be activated at the precise moment speech begins, ensuring complete speech capture while maintaining echo suppression, thereby improving recognition reliability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7636661B2Microphone initialization enhancement for speech recognition
Publication Date: 2009.12.22 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7636661B2 patent drawing
  • US7636661B2 patent drawing
  • US7636661B2 patent drawing

AI summary

A method and arrangement for improved speech recognition in a telephonically challenging speakerphone in-car environment. The method includes receiving a signal from a microphone representative of speech to be recognized, performing detection of a transition in the signal indicative of switch on of the microphone, and, in response to the detection, performing speech recognition on the signal with reduced contribution from an initial portion thereof. The initial portion may be treated as optional speech, the speech recognition may be performed with a predetermined redundant sound, and a user may be requested to speak the predetermined redundant sound when speech recognition has fallen below a predetermined threshold. Thus, recognition may be made possible when otherwise it would not be possible, recognition match scoring will be increased as the low weighting given by deleted initial sounds will be eliminated and therefore confusion of the recognized phrase will be reduced.