Real-Time Speech Video Synthesis Using Standby State Segments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in generating speech videos in real time due to the time-consuming process of synthesizing speech images and voices, which requires significant data and processing time.

Innovation Solution

The method involves sequentially playing back first sections of standby state videos, generating speech state images and voices based on speech content, and then synthesizing these images and voices with the standby state videos to create a real-time speech video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech images are synthesized using existing AI technologies, then the quality of speech video is improved, but the processing time increases significantly and real-time generation becomes difficult

Engineering Contradiction:
Improvespeech image qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system pre-generates multiple standby state videos showing the same person in different poses and expressions before actual speech video generation. These pre-prepared video segments are stored and can be quickly selected and combined with synthesized speech, eliminating the need to generate entire speech videos from scratch in real-time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech video generation process is divided into separate components: standby state video segments, synthesized speech audio, and face image generation. These segments are processed independently and then combined, allowing parallel processing and reducing overall generation time while maintaining quality.

Inventive Principle:
Principle #1Segmentation

2Reliability

If comprehensive data is used for synthesizing speech images, then the realism of speech video is improved, but the data requirements and processing complexity increase

Engineering Contradiction:
Improvespeech video realismVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system extracts and uses only the essential facial region from standby state videos to generate speech-related face images, rather than processing entire video frames. This extraction of the critical face portion reduces data volume and processing complexity while maintaining the realism of speech expressions.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of generating entirely new speech video data, the system copies and modifies existing standby state video segments by replacing facial regions with AI-generated speech face images. This copying approach maintains realism while reducing the computational complexity of creating completely new video data.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250157112A1Apparatus and method for providing speech video
Publication Date: 2025.05.15 DEEPBRAIN AI INC
  • US20250157112A1 patent drawing
  • US20250157112A1 patent drawing
  • US20250157112A1 patent drawing

AI summary

In a method for providing a speech video performed by a computing device according to one embodiment, first sections of a plurality of standby state videos are sequentially played back, wherein each standby state video includes the first section in which a person in the video is in a standby state and a second section for image interpolation between a last frame of the first section and a reference frame, a plurality of speech state images in which the person in the video is in a speech state and a speech voice based on a source of speech contents are generated and played back, when the generating of the plurality of speech state images and the speech voice is completed, the second section of the standby state video being played back at the time of completion, and a synthesized speech video is generated by synthesizing the plurality of speech state images and the speech voice with at least some of the plurality of standby state videos.