Avatar Visual Continuity During Speech Output
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for acoustic and visual output of content by an avatar result in significant time delays and unnatural transitions, disrupting direct communication between the user and the avatar.
Innovation Solution
A method utilizing a predetermined movement sequence with a playback length of less than 1 second, allowing for immediate and synchronized acoustic and visual output of content by the avatar, eliminating the need for interrupting video sequences and ensuring smooth transitions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If a video sequence with avatar movements is played for dynamic display, then the visual representation is smooth and natural, but the output of spoken content must wait until the video sequence finishes, causing time delay of several minutes
Solution Approach 1:
The patent segments the avatar representation into two distinct components: a long video sequence for background visual continuity and short movement sequences (less than 1 second) for active speech output. This segmentation allows the system to switch between different visual representations based on whether the avatar is listening or speaking, thereby reducing time delay while maintaining visual stability.
2Loss of time
If the video sequence is interrupted as soon as spoken content can be output, then the time delay is reduced, but a sudden visual transition and unnatural avatar movement occurs
Solution Approach 1:
The patent applies preliminary action by preparing short movement sequences in advance that can be seamlessly integrated into the long video sequence. These pre-prepared movement sequences contain the avatar's speech movements and are designed to match the timing and visual style of the main video sequence, allowing smooth transitions without abrupt changes when the avatar begins speaking.
3Device complexity
If the avatar is displayed statically using a still image, then the processing is simple, but the communication is less engaging and natural
Solution Approach 1:
The patent implements dynamics by switching the avatar display from static images or long video sequences to active short movement sequences when speech output is required. This dynamic adjustment allows the system to maintain simplicity during listening phases while providing engaging, natural-looking speech movements during communication phases, thereby improving user experience without permanently increasing system complexity.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A method for the acoustic and visual output of spoken content by an avatar is proposed. A user inputs the content, and in response to the input, the content is generated and output by the avatar. During the user input until the output of the spoken content, the avatar is visually represented using a predetermined movement sequence with a predetermined playback length of less than 1 second. Furthermore, a device, a computer program, and a computer-readable data carrier for implementing the proposed method are proposed.