Speech Image Interpolation for Real-Time Conversational Avatars

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing artificial intelligence technologies struggle to generate speech images in real time, particularly for conversations, due to the time-consuming process of synthesizing images and the large amount of data required.

Innovation Solution

A computing device with modules for generating a standby state image, interpolation images, and connection images using interpolation image sets to seamlessly transition between a standby and speech state, enabling real-time generation of speech images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech images are generated using traditional AI synthesis methods, then the quality and realism of the speech image is improved, but the generation time increases and real-time performance deteriorates

Engineering Contradiction:
Improvespeech image qualityVSAvoidgeneration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent pre-generates standby state images and speech preparation images before actual speech occurs. These pre-generated images are stored and ready for rapid retrieval and combination, eliminating the need for real-time synthesis of these components and enabling real-time speech image generation while maintaining high quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides the speech image generation process into separate components: standby state images, speech preparation images, and speech state images. Each component is generated and stored independently, then quickly combined during actual speech events, reducing overall generation time while preserving quality

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If traditional speech image synthesis is performed, then accurate speech representation is achieved, but the system complexity and data requirements increase

Engineering Contradiction:
Improvespeech representation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech image into distinct state components (standby, preparation, speech states) that can be independently generated and managed. This segmentation simplifies the overall system architecture by breaking down the complex synthesis process into manageable, pre-processable units with clearer data requirements

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

By pre-generating and storing speech preparation images and standby state images, the system reduces the computational complexity during actual speech events. The heavy synthesis work is performed in advance when system resources are more readily available, simplifying real-time processing

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12536727B2Speech image providing method and computing device for performing the same
Publication Date: 2026.01.27 DEEPBRAIN AI INC
  • US12536727B2 patent drawing
  • US12536727B2 patent drawing
  • US12536727B2 patent drawing

AI summary

A computing device according to an embodiment disclosed includes one or more processors and a memory storing one or more programs executed by the one or more processors, and a standby state image generating module configured to generate a standby state image in which a person is in a standby state, an interpolation image generating module configured to generate an interpolation image set for interpolation between the standby state image and a pre-stored speech preparation image, and an image playback module configured to generate a connection image for connecting the standby state image and a speech state image based on the interpolation image set when the speech state image is generated.