Microphone Initialization for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional microphone initialization techniques in speaker phones and mobile phone car installations suppress the microphone during audio playback, leading to loss of low-volume speech sounds at the beginning of an utterance, resulting in truncated recognition and potential misinterpretation of spoken phrases.
Innovation Solution
A method and system that detect the transition of the microphone switching on and perform speech recognition with reduced contribution from the initial portion of the signal, allowing for optional deletion of initial sounds or inclusion of a redundant word to ensure accurate recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-generated harmful factors
If the microphone is suppressed during audio playback to eliminate echo and feedback, then echo and feedback are eliminated, but low-volume speech sounds at the beginning of an utterance are lost
Solution Approach 1:
The system dynamically adjusts the microphone suppression state by detecting transitions in the audio signal. When a high-energy signal is detected indicating the user has started speaking, the system switches the microphone from suppressed to active state, allowing low-volume speech sounds to be captured while maintaining echo suppression during playback
Solution Approach 2:
The system performs preliminary detection of high-energy audio signals before fully activating the microphone. This preliminary action allows the system to prepare for speech capture in advance, ensuring that low-volume speech sounds at the beginning of an utterance are not lost while maintaining echo suppression during the transition period
2Object-generated harmful factors
If the microphone remains off until high energy audio signal is received, then echo and feedback are suppressed, but the spoken phrase is truncated and recognition fails
Solution Approach 1:
The system extracts and processes only the relevant portion of the speech signal for recognition. By detecting the microphone switch-on transition, the system identifies where speech actually begins and extracts this portion for recognition, effectively removing the problematic initial truncated portion while maintaining recognition reliability
Solution Approach 2:
The system changes the parameter of microphone activation from a static high-energy threshold to a dynamic transition-detection mechanism. This parameter change allows the microphone to be activated at the precise moment speech begins, ensuring complete speech capture while maintaining echo suppression, thereby improving recognition reliability
Data Source
AI summary
A method and arrangement for improved speech recognition in a telephonically challenging speakerphone in-car environment. The method includes receiving a signal from a microphone representative of speech to be recognized, performing detection of a transition in the signal indicative of switch on of the microphone, and, in response to the detection, performing speech recognition on the signal with reduced contribution from an initial portion thereof. The initial portion may be treated as optional speech, the speech recognition may be performed with a predetermined redundant sound, and a user may be requested to speak the predetermined redundant sound when speech recognition has fallen below a predetermined threshold. Thus, recognition may be made possible when otherwise it would not be possible, recognition match scoring will be increased as the low weighting given by deleted initial sounds will be eliminated and therefore confusion of the recognized phrase will be reduced.


