Interactive Avatar Response Using Spatial Context and Event Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing interactive avatars lack the ability to dynamically adjust their responses based on user context and environment, limiting their effectiveness in providing human-like interactions.
Innovation Solution
An electronic device and server system that processes user-uttered voice and spatial information to generate adaptive avatar animations and responses, incorporating facial expressions and voice synthesis, and adjusts animations in real-time based on detected events and environmental factors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If avatar responses are generated based on basic voice recognition, then the system complexity is low, but the interaction naturalness and emotional intimacy are insufficient
Solution Approach 1:
The system segments the avatar service into multiple independent modules: voice input processing, spatial information acquisition, response mode determination, voice answer generation, facial expression sequence generation, and animation reproduction. Each module handles a specific aspect of the interaction, allowing the system to achieve high naturalness while maintaining manageable complexity through modular architecture.
Solution Approach 2:
The system transitions from traditional single-dimensional text-based chatbot interactions to multi-dimensional interactions by incorporating spatial information (direction, distance), facial expressions, voice tone variations, and contextual awareness. This dimensional expansion enables more natural and emotionally intimate interactions without proportionally increasing system complexity.
2Adaptability or versatility
If the avatar uses fixed facial expressions and voice patterns, then the system is simple to implement, but the emotional expressiveness and user engagement are limited
Solution Approach 1:
The system implements dynamic facial expressions and voice patterns that adapt in real-time based on the determined response mode and generated facial expression sequences. The avatar's emotional expressiveness changes dynamically according to the conversation context, spatial information, and identified events, rather than using fixed patterns, thereby enhancing user engagement while managing complexity through algorithmic adaptation.
Solution Approach 2:
The system changes multiple parameters simultaneously to achieve emotional expressiveness: facial expression parameters (eye shape, mouth curvature, eyebrow position), voice parameters (pitch, tone, volume), and animation parameters (head orientation, body posture). These parameter changes are coordinated based on the determined response mode and generated sequences, allowing rich emotional expression without requiring completely separate systems for each parameter.
3Productivity
If the avatar responds immediately to all user inputs, then the interaction speed is high, but the contextual understanding and appropriate response timing are reduced
Solution Approach 1:
The system performs preliminary processing of spatial information and contextual analysis before generating the final response. By acquiring and analyzing spatial information upfront and determining the response mode in advance, the system prepares the groundwork for more accurate contextual understanding while maintaining efficient response generation, thus balancing interaction speed with contextual comprehension.
4Adaptability or versatility
If the system processes only basic voice data, then the processing load is low, but the spatial awareness and environmental adaptability are insufficient
Solution Approach 1:
The system uses a multi-functional processing architecture where the same computational framework handles multiple types of data: voice recognition, spatial information processing, contextual analysis, and response generation. This universal processing approach enables the system to acquire spatial awareness and environmental adaptability without requiring completely separate specialized systems for each function, thereby managing processing complexity while enhancing versatility.
Data Source
AI summary
A method of providing an avatar service includes obtaining a user-uttered voice and a spatial information of a user-utterance space, transmitting the user-uttered voice and the spatial information to a server, receiving, from the server, a first avatar voice answer and an avatar facial expression sequence corresponding to the first avatar voice, which are determined based on the user-uttered voice and the spatial information, determining first avatar facial expression data, based on the first avatar voice answer and the avatar facial expression sequence, identifying a certain event during reproduction of a first avatar animation created based on the first avatar voice answer and the first avatar facial expression data, determining second avatar facial expression data or a second avatar voice answer, based on the certain event, and reproducing a second avatar animation created based on the second avatar facial expression data or the second avatar voice answer.


