Voice Assistant Audio Adaptation for Environmental Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-assistant devices face challenges in ensuring the intelligibility of synthesized speech across varying listening environments, as factors like noise levels, reverberation, and user speech characteristics can significantly impact the clarity and comfort of the audio output.
Innovation Solution
The implementation of an audio appliance with an audio acquisition module, speech classifier, decision component, and speech synthesizer that adapts synthesized speech by classifying user utterances and identifying environmental cues to select appropriate speech synthesis parameters or models, modifying speech to match or differ from the user's speech mode, and adjusting playback volume to enhance intelligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synthesized speech is output at high volume to overcome background noise, then intelligibility in noisy environments improves, but user comfort deteriorates in quiet environments
Solution Approach 1:
The system dynamically adjusts speech output characteristics by detecting environmental noise levels and user speech modes in real-time. The decision component selects from multiple speech output modes (normal, whisper, Lombard-effect) based on current conditions, making the system adaptive rather than static. This resolves the contradiction by allowing high volume/Lombard-effect mode only when noise is detected, while using quiet whisper mode when the environment is calm.
Solution Approach 2:
The system changes multiple speech parameters simultaneously including volume level, spectral tilt, pitch contour, and speech rate based on detected conditions. The speech synthesizer modifies these parameters to match the selected speech output mode, enabling the same synthesized speech to be adapted for different acoustic environments and user preferences without requiring separate systems.
2Reliability
If speech synthesis parameters are heavily modified to match Lombard effect in noisy environments, then intelligibility improves, but naturalness of speech deteriorates
Solution Approach 1:
The system applies Lombard-effect modifications dynamically only when noisy environments are detected through acoustic scene analysis. In quiet environments, the system uses normal or whisper mode with minimal modifications, preserving naturalness. The transition between modes is smooth and context-appropriate, preventing the permanent artificialization of speech output.
Solution Approach 2:
The system applies selective parameter modifications based on the selected speech output mode. Lombard-effect mode modifies spectral tilt, pitch contour, and speech rate to enhance intelligibility in noise, while normal mode preserves natural speech characteristics. The speech synthesizer applies these parameter changes only when appropriate, maintaining a balance between intelligibility enhancement and naturalness preservation.
3Adaptability or versatility
If multiple speech synthesis models are maintained for different environments, then adaptability improves, but device complexity increases
Solution Approach 1:
The system uses a single speech synthesis model but applies different parameter settings and modification techniques tailored to specific environmental conditions. Rather than maintaining multiple complete models, the system locally adapts the output of one model through parameter adjustment and post-processing, achieving environmental adaptability with reduced complexity.
Solution Approach 2:
The system achieves multiple speech output modes (normal, whisper, Lombard-effect) by modifying parameters of a single base model rather than maintaining separate models for each mode. This parameter-based approach reduces memory requirements and computational overhead while preserving the ability to adapt speech characteristics to different environments and user preferences.
Data Source
AI summary
An appliance can include a microphone transducer, a processor, and a memory storing instructions. The appliance is configured to receive an audio signal at the microphone transducer and to detect an utterance in the audio signal. The appliance is further configured to classify a speech mode based on the utterance. The appliance is further configured to determine conditions of an environment of the appliance. The appliance is further configured to select at least one of a playback volume or a speech output mode from a plurality of speech output modes based on the classification, and the conditions of the environment of the appliance. The appliance is further configured to adapt the playback volume and/or mode of played-back speech according to the speech output mode. The appliance may be configured to synthesize speech according to the speech output mode, or to modify synthesized speech according to the speech output mode.


