Voice-Driven Video Synthesis for Low-Bandwidth Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is a need for technologies that can generate continuous video streams during video conferences without requiring live video capture, especially when connection speeds are insufficient or participants feel anxious about being continuously viewed.

Innovation Solution

A system that combines audio data with pre-recorded or captured video data altered by a machine learning model to predict facial speaking movements, creating a synchronized video stream that can be used as a virtual camera, allowing participants to contribute to video conferences without their actual video being on.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If live video capture is used during video conferences, then video quality and realism are improved, but processing requirements and network bandwidth consumption increase

Engineering Contradiction:
Improvevideo qualityVSAvoidprocessing requirements
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system pre-captures video data before the video conference and stores it for later use. This preliminary action allows the actual video conference to use pre-rendered video segments instead of requiring real-time video processing and transmission, significantly reducing processing requirements and network bandwidth consumption while maintaining video quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy of the pre-captured video data and uses this copy during the video conference instead of the original live video feed. This copying approach allows multiple participants to receive identical video data without requiring continuous real-time video streaming, reducing network bandwidth consumption and processing requirements

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If live video capture is enabled, then video presence and engagement are improved, but network bandwidth consumption and connection requirements increase

Engineering Contradiction:
Improvevideo presenceVSAvoidnetwork bandwidth
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

Video data is captured and prepared in advance before the video conference occurs. This pre-capture approach allows the system to store video segments locally or on a server, enabling participants to receive pre-rendered video instead of continuous live streams, thereby reducing network bandwidth consumption while maintaining video presence

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of continuous live video streaming, the system uses periodic or segmented video transmission where pre-captured video segments are transmitted at intervals or on-demand. This periodic approach reduces the overall network bandwidth requirement compared to continuous real-time video streaming while maintaining adequate video presence

Inventive Principle:
Principle #19Periodic action

3Speed

If continuous video streaming is implemented, then real-time visual feedback is improved, but participant anxiety and self-consciousness increase

Engineering Contradiction:
Improvereal-time feedbackVSAvoidparticipant anxiety
Core Design Contradiction:
SpeedVSObject-affected harmful factors

Solution Approach 1:

The system captures and prepares video content in advance before the participant needs to present. This preliminary preparation allows the participant to review and approve the pre-captured video, reducing anxiety about real-time performance while still providing visual feedback to other participants during the conference

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The pre-captured video acts as an intermediary between the participant and the video conference. Instead of directly streaming live video from the participant's camera, the system uses the pre-recorded video as a mediator, giving participants control over their video content and reducing the pressure of real-time self-awareness while maintaining visual presence

Inventive Principle:
Principle #24Intermediary (Mediator)

4Use of energy by moving object

If pre-recorded video is used instead of live capture, then processing requirements and network bandwidth are reduced, but video continuity and real-time synchronization may be compromised

Engineering Contradiction:
Improveprocessing requirementsVSAvoidvideo continuity
Core Design Contradiction:
Use of energy by moving objectVSStability of the object's composition

Solution Approach 1:

The system captures video data in advance and prepares it for synchronized playback during the video conference. This pre-capture approach allows the video to be pre-rendered and buffered, ensuring smooth continuous playback without real-time processing requirements while maintaining synchronization with audio and other conference elements

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system ensures continuous video playback by pre-capturing sufficient video segments and buffering them for uninterrupted transmission. This approach maintains video continuity by having ready-to-play segments available, eliminating gaps or interruptions that might occur with real-time processing, while still reducing overall processing requirements through offline video preparation

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20230027741A1Continuous video generation from voice data
Publication Date: 2023.01.26 EMC IP HLDG CO LLC
  • US20230027741A1 patent drawing
  • US20230027741A1 patent drawing
  • US20230027741A1 patent drawing

AI summary

One example method includes capturing audio data at a client engine while outputting an output video, the output video being based upon an original video stored at the client engine, delivering the captured audio data to a prediction engine upon the captured audio data being captured for a pre-determined time, receiving from the prediction engine substitute frame data used by the client engine to stitch one or more frames into the original video stored at the client engine, and following stitching the one or more frames into the output video to generate an altered output video, outputting the captured audio data and the altered video from the client engine.