Voice Response Personalization via Emotion Parameter Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice-recognizable electronic devices fail to provide responses that adequately reflect a user's context, style, and emotion, resulting in mechanical and unpersonalized interactions.
Innovation Solution
An electronic device equipped with a microphone and processor that transmits user utterances to an external server for partial automatic speech recognition and natural language understanding, modifying neutral responses based on identified conversation style and emotion parameters to generate contextually and emotionally relevant outputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If neutral text response is provided to user utterance, then response generation is simple and fast, but the response cannot reflect user's emotion, style, and context making it appear mechanical
Solution Approach 1:
The response generation process is segmented into multiple stages: first generating a neutral baseline response, then separately analyzing user emotion and style parameters, and finally combining these analyses to personalize the response. This segmentation allows the system to maintain simplicity in the base response generation while adding personalization layers without completely redesigning the core system.
Solution Approach 2:
The system performs preliminary analysis of user emotion and conversation style parameters before generating the final response. By pre-processing the user input to extract emotional and stylistic characteristics, the system prepares personalized parameters in advance that can be applied to the neutral response template, enabling efficient personalization without slowing down the overall response time.
2Adaptability or versatility
If user's emotion and style parameters are analyzed and applied to modify response text, then response quality and personalization are improved, but system complexity increases
Solution Approach 1:
The external server is designed to perform multiple functions: it conducts automatic speech recognition, analyzes user emotion, determines conversation style parameters, and generates personalized response text. By consolidating these diverse functions into a single multi-functional server, the system achieves high response personalization without proportionally increasing overall system complexity, as the server handles all these tasks through an integrated architecture.
Solution Approach 2:
An external server acts as an intermediary between the user's utterance and the electronic device's response. This intermediary handles the complex tasks of emotion analysis, style parameter determination, and response customization, allowing the local electronic device to remain relatively simple while still delivering personalized responses. The intermediary server absorbs the computational complexity of personalization algorithms.
3Speed
If conventional voice-recognizable devices provide mechanical responses, then processing speed is fast, but user engagement and interaction quality are poor
Solution Approach 1:
The system applies partial personalization to the response by selectively incorporating user emotion and style parameters into the neutral response template. Rather than completely regenerating the response from scratch with full personalization, the system modifies specific aspects of the neutral response based on analyzed parameters, achieving good user engagement while maintaining relatively fast processing speeds through this partial customization approach.
Data Source
AI summary
An electronic device includes a microphone, a communication circuit, and a processor configured to obtain a user's utterance through the microphone, transmit first information about the utterance through the communication circuit to an external server for at least partially automatic speech recognition (ASR) or natural language understanding (NLU), obtain a second text from the external server through the communication circuit, the second text being a text resulting from modifying at least part of a first text included in a neutral response to the utterance based on parameters corresponding to the user's conversation style and emotion identified based on the first information, and provide a voice corresponding to the second text or a message including the second text in response to the utterance.


