Neural Network Video Synthesis for Low Bandwidth Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing technologies face challenges with poor network connections and low bandwidth, leading to dropped frames or poor video quality, despite using compression techniques like H.264, which may still result in suboptimal performance.

Innovation Solution

The use of neural networks to generate video by processing audio data and a reference image, allowing for the transmission of only compressed audio, which reduces data transmission requirements and maintains presentation quality, even in low-bandwidth conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If video compression techniques like H.264 are used to reduce data transmission, then the amount of data to be transmitted is reduced, but video quality deteriorates and may still result in poor quality video presentation

Engineering Contradiction:
Improvedata transmission volumeVSAvoidvideo quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent replaces conventional mechanical video compression algorithms (H.264) with a neural network-based generative model. This substitution enables the system to generate high-quality video frames from audio data and reference images, achieving both reduced data transmission (only audio needs to be sent) and maintained video quality through intelligent synthesis rather than lossy compression

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates synthetic video copies by generating video frames that replicate the appearance and motion of the original video source using only audio input and a reference image. The neural network generates accurate visual copies of video content without requiring transmission of the actual video data, thus reducing bandwidth while maintaining quality

Inventive Principle:
Principle #26Copying

2Quantity of substance

If conventional video compression is used to handle low bandwidth, then data transmission is reduced, but reliability of video presentation deteriorates due to dropped frames

Engineering Contradiction:
Improvedata transmission volumeVSAvoidvideo presentation reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent replaces fragile mechanical compression-decompression pipelines with a robust neural network synthesis system. By generating video frames intelligently from audio and reference images, the system avoids the frame dropping and quality degradation inherent in conventional compression, ensuring reliable video presentation even under low bandwidth conditions

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary actions by capturing a reference image in advance and using it as the basis for generating subsequent video frames. This pre-established visual template allows the neural network to reliably reconstruct video content from audio alone, ensuring consistent and reliable video presentation without requiring continuous high-bandwidth video transmission

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220374637A1Synthesizing video from audio using one or more neural networks
Publication Date: 2022.11.24 NVIDIA CORP
  • US20220374637A1 patent drawing
  • US20220374637A1 patent drawing
  • US20220374637A1 patent drawing

AI summary

Apparatuses, systems, and techniques are presented to reduce an amount of data to be transmitted for media content. In at least one embodiment, one or more neural networks are used to generate video and audio information corresponding to one or more people based, at least in part, on at least one image and voice information corresponding to the one or more people.