Real-Time AI Speech Video Synthesis with Standby Motion Templates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI technologies face challenges in generating real-time speech videos due to the time-consuming synthesis process and high data requirements.

Innovation Solution

A computing device generates a standby state video and a speech state video, synthesizing them in real time by using a reference frame and back motion images to create a synthesized speech video, reducing data requirements and synthesis time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional speech video synthesis is performed using existing AI technologies, then speech video can be generated, but the synthesis process takes too much time and requires large amounts of data

Engineering Contradiction:
Improvespeech video generation speedVSAvoidsynthesis time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-generating standby state videos containing mouth shape data, head motion data, and body motion data before actual speech synthesis is needed. This pre-computation stores motion information that can be quickly combined with speech content during real-time synthesis, eliminating the need to generate motion data from scratch during the synthesis process itself.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the speech video synthesis process into distinct components: standby state video generation, speech audio generation, and video synthesis. By dividing the complex synthesis task into separate modules that can be independently processed and cached, the system reduces overall synthesis time while maintaining data efficiency.

Inventive Principle:
Principle #1Segmentation

2Productivity

If traditional speech video synthesis is performed using existing AI technologies, then speech video can be generated, but large amounts of data are required

Engineering Contradiction:
Improvespeech video generation efficiencyVSAvoiddata volume
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent pre-generates and caches standby state videos containing mouth shape sequences, head motion data, and body motion data. This preliminary action stores motion information that can be quickly combined with speech content during real-time synthesis, eliminating the need to generate motion data from scratch during the synthesis process itself.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates standby state videos as reusable templates that can be copied and applied to different speech content. Instead of generating complete speech videos from scratch each time, the system copies pre-computed motion data and mouth shape sequences from standby state videos, significantly reducing the computational data requirements for each new synthesis operation.

Inventive Principle:
Principle #26Copying

3Loss of time

If real-time speech video generation is implemented, then synthesis time is reduced, but the complexity of the system increases

Engineering Contradiction:
Improvesynthesis timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent segments the speech video synthesis process into distinct components: standby state video generation, speech audio generation, and video synthesis. By dividing the complex synthesis task into separate modules that can be independently processed and cached, the system reduces overall synthesis time while maintaining data efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces standby state videos as an intermediary element between speech content and final video output. These standby state videos act as pre-computed templates containing mouth shape data and motion information, serving as a mediator that simplifies the real-time synthesis process by providing ready-to-use motion data that can be quickly combined with new speech content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12367892B2Method for providing speech video and computing device for executing the method
Publication Date: 2025.07.22 DEEPBRAIN AI INC
  • US12367892B2 patent drawing
  • US12367892B2 patent drawing
  • US12367892B2 patent drawing

AI summary

In a method of providing a speech video according to an embodiment, a standby state video in which a person in a video is in a standby state is reproduced, a speech state video in which a person in a video is in a speech state based on a source of speech content is generated, the standby state video being reproduced to a reference frame of the standby state video being reproduced based on a back motion image is returned, and a synthesized speech video by synthesizing the returned reference frame and the speech state video is generated.