Speech Image Interpolation for Real-Time Conversational Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence technologies struggle to generate speech images in real time, particularly for conversations, due to the time-consuming process of synthesizing images and the large amount of data required.
Innovation Solution
A computing device with modules for generating a standby state image, interpolation images, and connection images using interpolation image sets to seamlessly transition between a standby and speech state, enabling real-time generation of speech images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If speech images are generated using traditional AI synthesis methods, then the quality and realism of the speech image is improved, but the generation time increases and real-time performance deteriorates
Solution Approach 1:
The patent pre-generates standby state images and speech preparation images before actual speech occurs. These pre-generated images are stored and ready for rapid retrieval and combination, eliminating the need for real-time synthesis of these components and enabling real-time speech image generation while maintaining high quality
Solution Approach 2:
The patent divides the speech image generation process into separate components: standby state images, speech preparation images, and speech state images. Each component is generated and stored independently, then quickly combined during actual speech events, reducing overall generation time while preserving quality
2Manufacturing precision
If traditional speech image synthesis is performed, then accurate speech representation is achieved, but the system complexity and data requirements increase
Solution Approach 1:
The patent segments the speech image into distinct state components (standby, preparation, speech states) that can be independently generated and managed. This segmentation simplifies the overall system architecture by breaking down the complex synthesis process into manageable, pre-processable units with clearer data requirements
Solution Approach 2:
By pre-generating and storing speech preparation images and standby state images, the system reduces the computational complexity during actual speech events. The heavy synthesis work is performed in advance when system resources are more readily available, simplifying real-time processing
Data Source
AI summary
A computing device according to an embodiment disclosed includes one or more processors and a memory storing one or more programs executed by the one or more processors, and a standby state image generating module configured to generate a standby state image in which a person is in a standby state, an interpolation image generating module configured to generate an interpolation image set for interpolation between the standby state image and a pre-stored speech preparation image, and an image playback module configured to generate a connection image for connecting the standby state image and a speech state image based on the interpolation image set when the speech state image is generated.


