Real-Time AI Speech Image Synthesis Using Pre-Generated Frames

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI technologies face challenges in generating natural speech images in real time, requiring significant data and time to synthesize images of conversations, leading to unnatural AI speech images from the perspective of the other person.

Innovation Solution

A computing device with modules for generating standby and speech state images, using a back-motion image set for interpolation between frames, allowing real-time generation of synthetic speech images by combining standby and speech state images during playback, reducing data requirements and generating natural speech shapes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If AI synthesizes speech images in real time for conversation, then responsiveness and naturalness improve, but it requires significant data and time for synthesis which is difficult to achieve in real time

Engineering Contradiction:
Improvespeech image generation speedVSAvoiddata requirement
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The system pre-generates multiple standby state images showing different mouth shapes and facial expressions that can be used during conversation. These images are prepared in advance and stored for quick retrieval during real-time speech generation, eliminating the need to generate all images on-demand

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech image generation process is divided into separate components: standby state images (pre-generated), speech state images (generated based on audio), and back-motion images (for interpolation). This segmentation allows each component to be optimized independently and combined efficiently during real-time operation

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If AI generates speech images with high quality and natural speech shapes, then image quality improves, but generation time increases making real-time conversation difficult

Engineering Contradiction:
Improvespeech image qualityVSAvoidgeneration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

Back-motion images are introduced as intermediary elements between standby state images and speech state images. These back-motion images serve as transition frames that smoothly connect different speech states, enabling natural interpolation without requiring complex real-time generation of intermediate frames

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

Back-motion images and standby state images are generated in advance and stored in a library. During real-time conversation, the system retrieves pre-generated images based on the current speech state and interpolates between them, avoiding the need for complex real-time synthesis while maintaining high image quality

Inventive Principle:
Principle #10Preliminary action

3Productivity

If the system uses pre-generated standby state images for real-time playback, then real-time performance improves, but maintaining natural speech shape transitions becomes challenging

Engineering Contradiction:
Improvereal-time speech image provisionVSAvoidspeech shape continuity
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The system dynamically selects and interpolates between different standby state images and back-motion images based on the current speech state. This dynamic approach allows the system to maintain natural speech shape transitions while using pre-generated images for real-time playback

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Back-motion images serve as intermediary transition frames that bridge gaps between different standby state images. These intermediate images ensure smooth and natural transitions in speech shapes during real-time playback, maintaining visual continuity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11830120B2Speech image providing method and computing device for performing the same
Publication Date: 2023.11.28 DEEPBRAIN AI INC
  • US11830120B2 patent drawing
  • US11830120B2 patent drawing
  • US11830120B2 patent drawing

AI summary

A computing device according to an embodiment includes one or more processors, a memory storing one or more programs executed by the one or more processors, a standby state image generating module configured to generate a standby state image in which a person is in a standby state, and generate a back-motion image set including a plurality of back-motion images at a preset frame interval from the standby state image for image interpolation between a preset reference frame of the standby state image, a speech state image generating module configured to generate a speech state image in which a person is in a speech state based on a source of speech content, and an image playback module configured to generate a synthetic speech image by combining the standby state image and the speech state image while playing the standby state image.