Multimodal Empathetic Conversation System for Virtual Human Emotion Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI conversation systems are limited in expressing complex emotions due to their unimodal methods, which restrict their ability to deliver empathetic responses effectively.
Innovation Solution
A method for real-time generation of empathy expressions in virtual humans using multimodal emotion recognition, which combines voice-based conversation and facial expression analysis to synchronize the emotional state of the virtual human with the user's emotions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If unimodal methods are used in AI conversation systems, then the system structure is simple, but the ability to express complex emotions is limited
Solution Approach 1:
The patent merges multiple modalities (voice-based conversation and facial expression analysis) into a unified empathetic conversation system. The voice-based conversation module processes audio inputs while the facial expression analysis module processes visual inputs, and both are integrated through the empathetic conversation generation module to produce comprehensive emotional responses, thereby resolving the contradiction between emotional expression capability and system complexity
Solution Approach 2:
The AI system is designed with multi-functionality to handle multiple types of input modalities (voice and facial expressions) and generate appropriate empathetic responses. The universal empathetic conversation generation module can process different input types and adapt its response generation accordingly, enabling the system to express complex emotions effectively while maintaining a unified architecture
2Adaptability or versatility
If unimodal methods are used in AI conversation systems, then the system is easier to implement, but the empathetic response capability is limited
Solution Approach 1:
The system is segmented into distinct functional modules: voice-based conversation module, facial expression analysis module, and empathetic conversation generation module. Each module can be independently developed, tested, and optimized, which facilitates easier implementation while achieving high empathetic response capability through the coordinated interaction of these specialized modules
Solution Approach 2:
The empathetic conversation generation module acts as an intermediary that integrates information from both the voice-based conversation module and the facial expression analysis module. This intermediary component synthesizes multi-modal data and generates appropriate empathetic responses, enabling complex empathetic capability while maintaining modular implementation ease
3Measurement precision
If multimodal emotion recognition is applied, then the accuracy of emotion assessment is improved, but the processing time increases
Solution Approach 1:
The system performs preliminary processing of voice and facial expression data independently before integrating them in the empathetic conversation generation module. The voice-based conversation module and facial expression analysis module process their respective inputs in parallel, preparing processed information in advance for integration, which reduces overall processing time while maintaining high emotion assessment accuracy through comprehensive multi-modal analysis
Data Source
AI summary
Provided are a conversational artificial intelligence (AI) system and method based on real-time multimodal emotion recognition. The system includes a model server configured to provide a machine learning-based conversational model, a terminal configured to perform a conversation with the machine learning-based conversational model through the model server, display a virtual human responding to a user during a conversation with the user, and capture a facial image of the user during the conversation, and a multimodal empathetic conversation-generation system configured to access the model server and receive a response to a question of the user from the terminal, and assess an emotion of the user from the facial image of the user and control, based on the assessed emotion, an expression of the virtual human displayed on the terminal.


