Text-Driven Video Synthesis for Low-Bandwidth Messaging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently transmitting video data over low-bandwidth channels like SMS due to high data requirements, making it difficult to send realistic video messages without incurring excessive costs or infrastructure changes.

Innovation Solution

A method and system that generate a video sequence from a text sequence, simulating visual and audible emotional expressions of a person, using a priori knowledge and minimal data transmission, allowing for realistic video representation without actual video data transfer over low-bandwidth channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video data is transmitted over low-bandwidth channels like SMS, then video communication is enabled, but the data transmission bandwidth requirement cannot be met due to limited channel capacity

Engineering Contradiction:
Improvevideo communication capabilityVSAvoiddata transmission bandwidth
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system pre-captures and stores a priori visual and audio information about the person before the actual communication event. This preliminary data collection enables the generation of realistic video sequences during low-bandwidth transmission without requiring real-time video data flow, thus resolving the bandwidth limitation while maintaining video communication capability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of transmitting actual video data, the system creates a synthetic copy of the video sequence by combining stored a priori information with text sequence data. This copying approach allows video communication to occur over low-bandwidth channels by generating video content locally at the receiving end rather than transmitting it

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If MMS is used to send multimedia messages, then video transmission capability is improved, but transmission cost increases and existing SMS infrastructure cannot be utilized

Engineering Contradiction:
Improvemultimedia messaging capabilityVSAvoidtransmission cost and infrastructure compatibility
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system makes the SMS infrastructure multi-functional by enabling it to carry not only traditional text messages but also to trigger the generation of video sequences. The text sequence sent via SMS serves dual purposes: as the message content and as the control signal for synthesizing the corresponding video, thus allowing existing SMS infrastructure to support multimedia messaging without additional cost or infrastructure changes

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The text sequence acts as an intermediary between the low-bandwidth SMS channel and the high-fidelity video output. Instead of directly transmitting video data over SMS, the text sequence mediates the process by controlling the synthesis of video sequences from a priori information, enabling cost-effective multimedia messaging through existing infrastructure

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a realistic video of a person is produced, then visual and audible emotional expressions are simulated, but large amounts of data need to be transmitted

Engineering Contradiction:
Improverealism of emotional expressionVSAvoidvideo data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system segments the video generation process into separate components: stored a priori visual information, stored a priori audio information, and transmitted text sequence data. By segmenting the data requirements and generating the video locally from these segments, the system achieves realistic emotional expressions without requiring transmission of complete video data

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transmits only the essential text sequence data rather than complete video data. This partial action approach transmits minimal necessary information (the text sequence that controls synthesis) while the actual video generation happens locally using pre-captured a priori information, significantly reducing data transmission volume while maintaining realism

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9082400B2Video generation based on text
Publication Date: 2015.07.14 REZVANI BEHROOZ
  • US9082400B2 patent drawing
  • US9082400B2 patent drawing
  • US9082400B2 patent drawing

AI summary

Techniques for generating a video sequence of a person based on a text sequence, are disclosed herein. Based on the received text sequence, a processing device generates the video sequence of a person to simulate visual and audible emotional expressions of the person, including using an audio model of the person's voice to generate an audio portion of the video sequence. The emotional expressions in the visual portion of the video sequence are simulated based a priori knowledge about the person. For instance, the a priori knowledge can include photos or videos of the person captured in real life.