Adaptive Speech Response Volume for Noisy Voice Dialogue
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech sound dialogue devices struggle to accurately recognize speech sounds due to difficulties in adjusting the volume of response speech sounds based on input speech sound volumes and environmental noise levels.
Innovation Solution
A speech sound response device equipped with a processor that detects voice input, generates responses, and determines the response volume based on both the voice volume and environmental sound volume, using pre-defined volume functions to optimize recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the speech sound dialogue device uses a preset volume or user-defined volume for response speech sound, then the device structure is simple, but the device cannot flexibly change the volume of the response speech sound in response to input volume variations
Solution Approach 1:
The patent implements dynamic volume control by making the response volume variable based on detected input speech volume levels. The system transitions from static preset volumes to dynamic adjustment where the response volume changes according to the input speech characteristics, allowing flexible adaptation without complex manual configuration
Solution Approach 2:
The system employs feedback mechanisms by detecting the volume of input speech and using this information to automatically adjust the response speech volume. The microphone captures input speech, the processor analyzes its volume characteristics, and the speaker outputs responses at appropriately adjusted volumes, creating a closed-loop adaptive system
2Measurement precision
If the speech sound dialogue device uses a microphone to collect all sounds, then the device can capture the talker's voice, but the device cannot distinguish the talker's voice from other environmental sounds
Solution Approach 1:
The patent applies local quality by focusing the microphone's attention on specific sound characteristics associated with human speech rather than treating all sounds uniformly. The system enhances detection of speech-specific features (frequency ranges, temporal patterns) while filtering out environmental noise, improving voice detection accuracy without requiring multiple microphones or complex arrays
Solution Approach 2:
The system changes detection parameters dynamically by adjusting sensitivity thresholds and filtering characteristics based on the detected sound environment. The processor modifies detection parameters to optimize speech recognition in varying acoustic conditions, enabling accurate voice detection while maintaining simple hardware configuration
3Reliability
If the volume of the input speech sound is too loud or too low, then the speech sound dialogue device can still collect the sound, but the device cannot obtain a correct recognition result
Solution Approach 1:
The system performs preliminary volume assessment by detecting and analyzing the volume characteristics of input speech before processing the recognition. This preliminary action allows the system to identify potentially problematic volume levels and apply appropriate preprocessing or alert the user, ensuring recognition accuracy across a wide volume range
Solution Approach 2:
The patent implements dynamic volume adaptation by adjusting the response speech volume based on the detected input volume characteristics. When input speech is detected at extreme volumes, the system dynamically modifies its response strategy to maintain effective communication, expanding its operational volume range while maintaining reliable recognition
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
A speech sound response device includes a microphone, a processor, and a speaker. The microphone is configured to acquire sound. The processor is configure to detect a voice in the sound, generate a response to the voice, determine a response volume for outputting the response based on (i) a voice volume of the voice in the sound and (ii) an environmental volume of an environmental sound other than the voice, and control the speaker to output the response at the response volume.