Avatar Visual Continuity During Speech Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for acoustic and visual output of content by an avatar result in significant time delays and unnatural transitions, disrupting direct communication between the user and the avatar.

Innovation Solution

A method utilizing a predetermined movement sequence with a playback length of less than 1 second, allowing for immediate and synchronized acoustic and visual output of content by the avatar, eliminating the need for interrupting video sequences and ensuring smooth transitions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If a video sequence with avatar movements is played for dynamic display, then the visual representation is smooth and natural, but the output of spoken content must wait until the video sequence finishes, causing time delay of several minutes

Engineering Contradiction:
Improvevisual continuityVSAvoidtime delay
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent segments the avatar representation into two distinct components: a long video sequence for background visual continuity and short movement sequences (less than 1 second) for active speech output. This segmentation allows the system to switch between different visual representations based on whether the avatar is listening or speaking, thereby reducing time delay while maintaining visual stability.

Inventive Principle:
Principle #1Segmentation

2Loss of time

If the video sequence is interrupted as soon as spoken content can be output, then the time delay is reduced, but a sudden visual transition and unnatural avatar movement occurs

Engineering Contradiction:
Improvetime delayVSAvoidvisual smoothness
Core Design Contradiction:
Loss of timeVSStability of the object's composition

Solution Approach 1:

The patent applies preliminary action by preparing short movement sequences in advance that can be seamlessly integrated into the long video sequence. These pre-prepared movement sequences contain the avatar's speech movements and are designed to match the timing and visual style of the main video sequence, allowing smooth transitions without abrupt changes when the avatar begins speaking.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If the avatar is displayed statically using a still image, then the processing is simple, but the communication is less engaging and natural

Engineering Contradiction:
Improveprocessing complexityVSAvoiduser experience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The patent implements dynamics by switching the avatar display from static images or long video sequences to active short movement sequences when speech output is required. This dynamic adjustment allows the system to maintain simplicity during listening phases while providing engaging, natural-looking speech movements during communication phases, thereby improving user experience without permanently increasing system complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4567583A1Method for the acoustic and visual output of a content to be transmitted by an avatar
Publication Date: 2025.06.11 GOAVA GMBH
  • EP4567583A1 patent drawingFigure 1
  • EP4567583A1 patent drawingFigure 2
  • EP4567583A1 patent drawingFigure 3~4

AI summary

A method for the acoustic and visual output of spoken content by an avatar is proposed. A user inputs the content, and in response to the input, the content is generated and output by the avatar. During the user input until the output of the spoken content, the avatar is visually represented using a predetermined movement sequence with a predetermined playback length of less than 1 second. Furthermore, a device, a computer program, and a computer-readable data carrier for implementing the proposed method are proposed.