Voice Enhancement System with Variable Gain Normalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices face challenges in maintaining consistent audio loudness levels, leading to difficulties in understanding voice commands due to variations in speech loudness, especially when users are at different distances from the microphone.

Innovation Solution

A system that performs voice enhancement by identifying active intervals in audio data, applying variable gains based on desired output loudness and flatness values to normalize power levels, and selectively amplifies voice data to achieve uniform loudness across words and sentences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If variable gain amplification is applied to all audio data, then loudness consistency is improved, but speech intelligibility deteriorates due to amplification of background noise

Engineering Contradiction:
Improveloudness consistencyVSAvoidspeech intelligibility
Core Design Contradiction:
Stability of the object's compositionVSLoss of information

Solution Approach 1:

The patent applies different gain values to different portions of the audio signal based on voice activity detection. During voice activity intervals, higher gain is applied to amplify speech; during non-voice intervals, lower gain is applied to minimize background noise amplification. This local differentiation resolves the contradiction by making the amplification property spatially and temporally selective rather than uniform.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the gain parameter based on real-time voice activity detection results. The gain value changes over time according to whether voice activity is present, transforming from a static amplification approach to a dynamic one that adapts to the audio content, thereby maintaining speech intelligibility while improving loudness consistency.

Inventive Principle:
Principle #15Dynamics

2Loss of information

If voice activity detection is used to selectively amplify speech, then speech intelligibility is improved, but system complexity increases due to additional processing requirements

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The audio signal is segmented into discrete intervals, and voice activity detection is performed on each interval independently. This segmentation approach simplifies the overall processing by breaking down the continuous signal analysis into manageable discrete units, reducing computational complexity while maintaining speech intelligibility through selective amplification of detected voice intervals.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10600432B1Methods for voice enhancement
Publication Date: 2020.03.24 AMAZON TECH INC
  • US10600432B1 patent drawing
  • US10600432B1 patent drawing
  • US10600432B1 patent drawing

AI summary

A system configured to perform power normalization for voice enhancement. The system may identify active intervals corresponding to voice activity and may selectively amplify the active intervals in order to generate output audio data at a near uniform loudness. The system may determine a variable gain for each of the active intervals based on a desired output loudness and a flatness value, which indicates how much a signal envelope is to be modified. For example, a low flatness value corresponds to no modification, with peak active interval values corresponding to the desired output loudness and lower active intervals being lower than the desired output loudness. In contrast, a high flatness value corresponds to extensive modification, with peak active interval values and lower active interval values both corresponding to the desired output loudness. Thus, individual words may share the same peak power level.