Neural Network Audio Upsampling for Bandwidth Quality Trade-off

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio transmission technologies often compromise audio quality by reducing sample rates to minimize bandwidth, resulting in lower quality audio that is less clear and less enjoyable for users, particularly in applications like video conferencing and media streaming.

Innovation Solution

An audio upsampling system utilizing a deep learning network running on a GPU, which increases the audio resolution of received audio data from lower sample rates to higher sample rates, such as from 8 kHz to 48 kHz, providing high-quality audio with low latency and minimal GPU consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If audio sample rate is reduced to minimize bandwidth, then bandwidth requirements are decreased, but audio quality deteriorates

Engineering Contradiction:
ImprovebandwidthVSAvoidaudio quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical/audio signal processing with neural network-based deep learning models to perform audio upsampling. The neural network learns complex patterns from low-sample-rate audio and synthesizes high-sample-rate audio signals, achieving quality restoration without conventional interpolation methods. This substitution enables high-quality audio reconstruction from compressed data, resolving the contradiction between bandwidth efficiency and audio quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of audio representation by transforming audio from time-domain samples to frequency-domain spectral representations through neural networks. The model processes audio at low sample rates (e.g., 8kHz) and outputs enhanced audio at high sample rates (e.g., 48kHz or 96kHz), effectively changing the sampling parameter while maintaining or improving quality. This parameter transformation allows bandwidth reduction without sacrificing audio fidelity.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If audio sample rate is increased to improve quality, then audio clarity and richness are improved, but bandwidth requirements increase

Engineering Contradiction:
Improveaudio qualityVSAvoidbandwidth
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training neural network models offline using large datasets of high-quality audio. The trained models are then deployed to perform real-time upsampling of low-sample-rate audio streams. This preliminary training phase captures complex audio patterns and relationships, enabling the system to generate high-quality high-sample-rate audio from compressed low-sample-rate input without requiring continuous high bandwidth during actual audio transmission.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating synthetic high-sample-rate audio signals from low-sample-rate inputs through neural network inference. The model copies and reconstructs the essential audio information, adding missing frequency components and temporal details that were lost during initial compression. This copying process generates virtual high-quality audio data that enhances the original signal without requiring transmission of the complete high-sample-rate original signal.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If neural network-based upsampling is applied to restore high sample rate, then audio quality is improved, but computational resources and latency are increased

Engineering Contradiction:
Improveaudio qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the audio processing task into distinct stages: (1) audio compression at low sample rate for transmission, (2) neural network-based upsampling to restore high sample rate, and (3) audio playback. The neural network model itself is segmented into efficient computational layers (convolutional layers, recurrent layers, attention mechanisms) that process audio in manageable frames. This segmentation allows the system to achieve high-quality upsampling while managing computational complexity through optimized architecture and processing stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamics by making the upsampling system adaptive to different audio conditions and requirements. The neural network can dynamically adjust processing based on input characteristics, audio content type, and computational resources available. The system can operate with varying degrees of computational intensity depending on the specific application needs, balancing quality restoration with resource consumption in real-time.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230076431A1Audio upsampling using one or more neural networks
Publication Date: 2023.03.09 NVIDIA CORP
  • US20230076431A1 patent drawing
  • US20230076431A1 patent drawing
  • US20230076431A1 patent drawing

AI summary

Apparatuses, systems, and techniques are presented to upsample audio. In at least one embodiment, one or more neural networks are used to determine one or more second frequencies of one or more audio signals based, at least in part, on only one or more first frequencies of the one or more audio signals