Speech Model Audio Enhancement for Low-Bandwidth Delay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Low-bandwidth network connections in electronic communication systems lead to lag and delays in audio transmission, which can negatively impact the quality of communication, and conventional compression methods often reduce the overall quality of the transmitted signal.

Innovation Solution

The use of bandwidth-efficient representations of audio data combined with trained speech models that enhance compressed audio in real-time, utilizing both compressed and uncompressed audio signals to improve quality without significant delays, where the uncompressed audio serves as a ground truth for training the models during communication sessions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If audio data is transmitted in raw uncompressed format, then audio quality is maintained, but transmission delays and bandwidth consumption increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidtransmission delay
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary compression of audio data before transmission to reduce bandwidth requirements and transmission time. The compressed audio is sent ahead of time, while the original uncompressed audio is retained for later use as ground truth data for model training and enhancement.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A trained speech enhancement model acts as an intermediary between the compressed audio signal and the final output. The model enhances the compressed audio by predicting and restoring high-frequency content and other audio features that were lost during compression, using the uncompressed audio as training data to learn the transformation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If audio data is compressed to reduce bandwidth consumption, then transmission delays are reduced, but audio quality deteriorates

Engineering Contradiction:
Improvetransmission delayVSAvoidaudio quality
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The trained speech enhancement model serves as an intermediary processing stage that receives compressed audio and transforms it into enhanced audio output. The model learns to predict missing audio features and restore quality by training on pairs of compressed and uncompressed audio, effectively mediating between the bandwidth-efficient compressed format and the high-quality uncompressed target.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameters of the audio signal by applying learned transformations through the speech enhancement model. The model adjusts spectral characteristics, frequency content, and temporal features of the compressed audio to match the quality characteristics of uncompressed audio, effectively changing the audio parameters to improve quality without increasing bandwidth.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If uncompressed audio is transmitted to maintain quality, then bandwidth consumption increases, but network efficiency decreases

Engineering Contradiction:
Improveaudio qualityVSAvoidnetwork efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs preliminary compression of audio data before transmission to reduce bandwidth requirements and transmission time. The compressed audio is sent ahead of time, while the original uncompressed audio is retained for later use as ground truth data for model training and enhancement.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A trained speech enhancement model acts as an intermediary between the compressed audio signal and the final output. The model enhances the compressed audio by predicting and restoring high-frequency content and other audio features that were lost during compression, using the uncompressed audio as training data to learn the transformation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10839809B1Online training with delayed feedback
Publication Date: 2020.11.17 AMAZON TECH INC
  • US10839809B1 patent drawing
  • US10839809B1 patent drawing
  • US10839809B1 patent drawing

AI summary

Bandwidth-efficient (i.e., compressed) representations of audio data can be utilized for near real-time presentation of the audio on one or more receiving devices. Persons identified as having speech represented in the audio data can have trained speech models provided to the devices. These trained models can be used to classify the compressed audio in order to improve the quality to correspond more closely to the uncompressed version, without experiencing lag that might otherwise be associated with transmission of the uncompressed audio. The uncompressed audio is also received, with potential lag, and is used to further train the speech models in near real time. The ability to utilize the uncompressed audio as it is received prevents a need to store or further transmit the audio data for offline processing, and enables the further trained model to be used during the communication session.