Real-Time AI Speech Image Synthesis Using Pre-Generated Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI technologies face challenges in generating natural speech images in real time, requiring significant data and time to synthesize images of conversations, leading to unnatural AI speech images from the perspective of the other person.
Innovation Solution
A computing device with modules for generating standby and speech state images, using a back-motion image set for interpolation between frames, allowing real-time generation of synthetic speech images by combining standby and speech state images during playback, reducing data requirements and generating natural speech shapes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If AI synthesizes speech images in real time for conversation, then responsiveness and naturalness improve, but it requires significant data and time for synthesis which is difficult to achieve in real time
Solution Approach 1:
The system pre-generates multiple standby state images showing different mouth shapes and facial expressions that can be used during conversation. These images are prepared in advance and stored for quick retrieval during real-time speech generation, eliminating the need to generate all images on-demand
Solution Approach 2:
The speech image generation process is divided into separate components: standby state images (pre-generated), speech state images (generated based on audio), and back-motion images (for interpolation). This segmentation allows each component to be optimized independently and combined efficiently during real-time operation
2Manufacturing precision
If AI generates speech images with high quality and natural speech shapes, then image quality improves, but generation time increases making real-time conversation difficult
Solution Approach 1:
Back-motion images are introduced as intermediary elements between standby state images and speech state images. These back-motion images serve as transition frames that smoothly connect different speech states, enabling natural interpolation without requiring complex real-time generation of intermediate frames
Solution Approach 2:
Back-motion images and standby state images are generated in advance and stored in a library. During real-time conversation, the system retrieves pre-generated images based on the current speech state and interpolates between them, avoiding the need for complex real-time synthesis while maintaining high image quality
3Productivity
If the system uses pre-generated standby state images for real-time playback, then real-time performance improves, but maintaining natural speech shape transitions becomes challenging
Solution Approach 1:
The system dynamically selects and interpolates between different standby state images and back-motion images based on the current speech state. This dynamic approach allows the system to maintain natural speech shape transitions while using pre-generated images for real-time playback
Solution Approach 2:
Back-motion images serve as intermediary transition frames that bridge gaps between different standby state images. These intermediate images ensure smooth and natural transitions in speech shapes during real-time playback, maintaining visual continuity
Data Source
AI summary
A computing device according to an embodiment includes one or more processors, a memory storing one or more programs executed by the one or more processors, a standby state image generating module configured to generate a standby state image in which a person is in a standby state, and generate a back-motion image set including a plurality of back-motion images at a preset frame interval from the standby state image for image interpolation between a preset reference frame of the standby state image, a speech state image generating module configured to generate a speech state image in which a person is in a speech state based on a source of speech content, and an image playback module configured to generate a synthetic speech image by combining the standby state image and the speech state image while playing the standby state image.


