Lombard Effect Speech Synthesis for Noisy Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis systems fail to deliver clear notification voices in noisy environments, leading to auditory stress and increased maximum output sound volume, which is costly and inefficient.
Innovation Solution
A method involving a speech synthesis system that extracts utterance features, generates a feature vector, and applies it to a pre-trained model with a parameter adjusting the Lombard effect, measuring external noise to determine the appropriate weight for synthesizing speech, thereby optimizing sound output in noisy conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech synthesis outputs at the highest volume in noisy environments, then notification voice clarity is improved, but auditory stress increases and device cost increases
Solution Approach 1:
The speech synthesis system dynamically adjusts the Lombard effect parameter based on measured external noise levels. When noise exceeds a threshold, the parameter weight is increased to apply Lombard effect (higher pitch, stronger intensity); when noise is low, the weight is reduced or set to zero. This dynamic adaptation resolves the contradiction by providing high volume only when necessary for clarity, reducing auditory stress during normal conditions.
Solution Approach 2:
The system changes the Lombard effect parameter (controlling pitch and intensity characteristics) based on environmental noise conditions. By adjusting this parameter weight from 0 to higher values depending on noise levels, the system achieves clear notification voice in noisy environments without consistently operating at maximum volume, thereby reducing auditory stress and device cost.
2Reliability
If speech synthesis outputs at the highest volume in noisy environments, then notification voice clarity is improved, but device cost increases
Solution Approach 1:
The system dynamically controls output volume based on external noise measurement. By adjusting the Lombard effect parameter weight according to noise levels, the device avoids consistently operating at maximum volume, reducing the need for high-power output components and lowering overall device cost while maintaining clarity when needed.
Solution Approach 2:
By changing the Lombard effect parameter to adjust pitch and intensity characteristics rather than simply increasing volume, the system achieves clear notification voice without requiring maximum power output capability, thereby reducing device manufacturing cost.
3Reliability
If Lombard effect parameter weight is increased to enhance speech clarity in noise, then notification voice clarity is improved, but energy consumption increases
Solution Approach 1:
The system dynamically adjusts the Lombard effect parameter weight based on real-time noise level measurement. The weight is increased only when external noise exceeds a threshold, and reduced or set to zero when noise is low. This dynamic control achieves speech clarity in noisy environments without continuously consuming additional energy, resolving the contradiction between clarity and energy consumption.
Solution Approach 2:
By changing the Lombard effect parameter to adjust pitch and intensity characteristics rather than simply increasing volume, the system achieves clearer speech in noise with more efficient energy consumption compared to maximum volume output.
Data Source
AI summary
Disclosed is speech synthesis in a noisy environment. According to an embodiment of the disclosure, a method of speech synthesis may generate a Lombard effect-applied synthesized speech using a feature vector generated from an utterance feature. According to the disclosure, the speech synthesis method and device may be related to artificial intelligence (AI) modules, unmanned aerial vehicles (UAVs), robots, augmented reality (AR) devices, virtual reality (VR) devices, and 5G service-related devices.


