Utterance Detection Apparatus Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In speech translation systems, especially in medical settings where operators' hands are occupied, existing techniques struggle to accurately detect the end of an utterance due to noise interference from other speakers, leading to delays in speech recognition and translation output.
Innovation Solution
The system employs two directional microphones to calculate sound pressure differences, suppresses the utterance start direction sound pressure when it falls below the non-utterance start direction sound pressure, and uses this suppression to improve the detection of utterance ends by adjusting the sound pressure thresholds, thereby reducing noise interference and enhancing detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If noise suppression is applied to improve utterance detection accuracy, then detection precision improves, but the system complexity increases
Solution Approach 1:
The patent segments the audio signal processing into distinct phases: pre-utterance period processing, utterance detection period processing, and post-utterance period processing. During the pre-utterance period, the system establishes baseline noise characteristics. During the utterance detection period, the system dynamically adjusts suppression parameters based on real-time audio analysis. This segmentation allows the complex noise suppression to be managed through structured temporal phases rather than a monolithic approach.
Solution Approach 2:
The patent implements dynamic noise suppression by adjusting suppression parameters in real-time based on the detection state. The suppression amount is not fixed but varies according to the utterance detection status, the pre-utterance noise characteristics, and the current audio signal properties. This dynamic adaptation allows the system to maintain high detection accuracy while reducing unnecessary processing complexity during stable states.
2Loss of time
If the detection threshold is lowered to detect utterance ends faster, then response time improves, but false detection rate increases
Solution Approach 1:
The patent performs preliminary analysis during the pre-utterance period to establish baseline noise characteristics and suppression parameters before actual utterance detection begins. This preliminary action allows the system to be prepared with appropriate threshold settings and noise models, enabling faster and more accurate utterance end detection without increasing false detection rates, as the thresholds are already optimized based on the specific acoustic environment.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors detection results and adjusts thresholds accordingly. When utterance ends are detected, the system uses this feedback to refine the noise suppression parameters for subsequent detections. This feedback loop allows the system to maintain high reliability while achieving fast response times, as the thresholds are continuously optimized based on actual performance data.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively reduces delays in detecting utterance ends, improving the accuracy and speed of speech recognition and translation output, ensuring timely and accurate language translation in multi-speaker environments.
Implementation Method 1
detect an utterance start based on a first sound pressure based on first audio data acquired from a first microphone and a second sound pressure based on second audio data acquired from a second microphone
Data Source
AI summary
An utterance detection apparatus includes a processor configured to: detect an utterance start based on a first sound pressure based on first audio data acquired from a first microphone and a second sound pressure based on second audio data acquired from a second microphone; suppress an utterance start direction sound pressure when the utterance start direction sound pressure, which is one of the first sound pressure and the second sound pressure being larger at a time point of detecting the utterance start, falls below a non-utterance start direction sound pressure, which is the other one of the first sound pressure and the second sound pressure being smaller at the time point of detecting the utterance start; and detect an utterance end based on the suppressed utterance start direction sound pressure.


