Utterance Detection Apparatus Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In speech translation systems, especially in medical settings where operators' hands are occupied, existing techniques struggle to accurately detect the end of an utterance due to noise interference from other speakers, leading to delays in speech recognition and translation output.

Innovation Solution

The system employs two directional microphones to calculate sound pressure differences, suppresses the utterance start direction sound pressure when it falls below the non-utterance start direction sound pressure, and uses this suppression to improve the detection of utterance ends by adjusting the sound pressure thresholds, thereby reducing noise interference and enhancing detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If noise suppression is applied to improve utterance detection accuracy, then detection precision improves, but the system complexity increases

Engineering Contradiction:
Improveutterance end detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio signal processing into distinct phases: pre-utterance period processing, utterance detection period processing, and post-utterance period processing. During the pre-utterance period, the system establishes baseline noise characteristics. During the utterance detection period, the system dynamically adjusts suppression parameters based on real-time audio analysis. This segmentation allows the complex noise suppression to be managed through structured temporal phases rather than a monolithic approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic noise suppression by adjusting suppression parameters in real-time based on the detection state. The suppression amount is not fixed but varies according to the utterance detection status, the pre-utterance noise characteristics, and the current audio signal properties. This dynamic adaptation allows the system to maintain high detection accuracy while reducing unnecessary processing complexity during stable states.

Inventive Principle:
Principle #15Dynamics

2Loss of time

If the detection threshold is lowered to detect utterance ends faster, then response time improves, but false detection rate increases

Engineering Contradiction:
Improveutterance end detection delayVSAvoiddetection accuracy
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent performs preliminary analysis during the pre-utterance period to establish baseline noise characteristics and suppression parameters before actual utterance detection begins. This preliminary action allows the system to be prepared with appropriate threshold settings and noise models, enabling faster and more accurate utterance end detection without increasing false detection rates, as the thresholds are already optimized based on the specific acoustic environment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors detection results and adjusts thresholds accordingly. When utterance ends are detected, the system uses this feedback to refine the noise suppression parameters for subsequent detections. This feedback loop allows the system to maintain high reliability while achieving fast response times, as the thresholds are continuously optimized based on actual performance data.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach effectively reduces delays in detecting utterance ends, improving the accuracy and speed of speech recognition and translation output, ensuring timely and accurate language translation in multi-speaker environments.

Implementation Method 1

detect an utterance start based on a first sound pressure based on first audio data acquired from a first microphone and a second sound pressure based on second audio data acquired from a second microphone

Methodology Applied
Scientific EffectSound pressure: Sound

Data Source

PatentUS11205416B2Non-transitory computer-read able storage medium for storing utterance detection program, utterance detection method, and utterance detection apparatus
Publication Date: 2021.12.21 FUJITSU LTD
  • US11205416B2 patent drawing
  • US11205416B2 patent drawing
  • US11205416B2 patent drawing

AI summary

An utterance detection apparatus includes a processor configured to: detect an utterance start based on a first sound pressure based on first audio data acquired from a first microphone and a second sound pressure based on second audio data acquired from a second microphone; suppress an utterance start direction sound pressure when the utterance start direction sound pressure, which is one of the first sound pressure and the second sound pressure being larger at a time point of detecting the utterance start, falls below a non-utterance start direction sound pressure, which is the other one of the first sound pressure and the second sound pressure being smaller at the time point of detecting the utterance start; and detect an utterance end based on the suppressed utterance start direction sound pressure.