Speech Recognition Device Using Dynamic Volume Threshold Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition devices face difficulties in distinguishing background noise from speech, especially in environments with multiple speakers or reverberation, and struggle to set an appropriate volume threshold for effective speech recognition.

Innovation Solution

A speech recognition device with an interactive adjustment mechanism, utilizing a microphone, adjustment processor, and recognition processor, which allows users to set and adjust a threshold based on user input, discarding audio signals below the threshold and performing recognition on signals above it, while incorporating features like trigger words and noise power ratios to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a high volume threshold is set for speech recognition, then noise can be filtered out more effectively, but speech with low volume may be discarded incorrectly

Engineering Contradiction:
Improvenoise filtering accuracyVSAvoidspeech recognition accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The threshold is made dynamic rather than fixed. The adjustment processor allows the threshold to be changed based on user feedback and environmental conditions. The system transitions from a static threshold to an adaptive threshold that learns from user corrections, resolving the contradiction between filtering noise and preserving low-volume speech.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates user feedback through correction instructions. When users correct recognition errors, the system adjusts the threshold accordingly. This feedback loop enables the threshold to adapt to actual speech patterns and environmental noise levels, improving both noise filtering and speech recognition accuracy.

Inventive Principle:
Principle #23Feedback

2Productivity

If a low volume threshold is set to capture all speech, then more speech can be recognized, but noise differentiation becomes more difficult

Engineering Contradiction:
Improvespeech capture rateVSAvoidnoise filtering accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

User feedback through correction instructions enables the system to learn which low-volume signals are actual speech versus noise. The adjustment processor uses this feedback to refine the threshold, maintaining high speech capture rates while improving noise differentiation over time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-adjustment of the threshold based on user corrections. The adjustment processor automatically modifies the threshold parameter without requiring manual reconfiguration, enabling the system to adapt to changing environmental conditions and speech patterns autonomously.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If the threshold is adjusted frequently to adapt to environmental changes, then recognition accuracy improves, but system complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidadjustment mechanism complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system adjusts a single critical parameter (the volume threshold) rather than multiple complex parameters. This simple parameter change approach enables adaptation to environmental changes while maintaining relatively low system complexity. The adjustment processor focuses on modifying only the threshold value based on user feedback.

Inventive Principle:
Principle #35Parameter changes

4Object-affected harmful factors

If the threshold is set too high to filter noise, then noise rejection improves, but speech recognition is lost in noisy environments

Engineering Contradiction:
Improvenoise rejectionVSAvoidspeech recognition reliability
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The threshold dynamically adapts to environmental noise levels through user feedback. In noisy environments, the threshold adjusts downward to capture speech while still filtering noise. This dynamic adjustment maintains both noise rejection and speech recognition reliability across varying environmental conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

User corrections provide feedback about actual speech patterns in the environment. The adjustment processor uses this feedback to calibrate the threshold appropriately for the current noise level, ensuring that speech is not lost while maintaining noise rejection capabilities.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables accurate speech recognition by filtering out noise and focusing on the target speaker's voice, improving recognition accuracy and reducing errors in noisy environments.

Implementation Method 1

a microphone 101, an adjustment processor 103 and a recognition processor 104. The microphone 101 detects sound

Methodology Applied
Scientific EffectElectromagnetic induction: Electromagnetic Induction

Data Source

PatentUS10579327B2Speech recognition device, speech recognition method and storage medium using recognition results to adjust volume level threshold
Publication Date: 2020.03.03 KK TOSHIBA
  • US10579327B2 patent drawing
  • US10579327B2 patent drawing
  • US10579327B2 patent drawing

AI summary

In a speech recognition device according to one embodiment, a microphone detects sound and generates an audio signal corresponding to the sound, an adjustment processor adjusts a threshold to be a value less than a first volume level of first input audio signal generated by the microphone, and registers the adjusted threshold, a recognition processor reads the registered threshold, compares the registered threshold with a second input audio signal, discards the second input audio signal when a second volume level of the second input audio signal is less than the registered threshold, and performs a recognition process as the audio signal of a user to be recognized when the second volume level of the second input audio signal is greater than or equal to the registered threshold.