Sympathetic Back-Channel Signal Timing for Natural Voice Avatars
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital humans, intelligence robots, and voice avatar chatbots lack the ability to deliver natural back-channel signals, leading to unnatural conversations.
Innovation Solution
A system and method for generating sympathetic back-channel signals based on user input voice and image information, determining appropriate timing for signal output, and outputting both image and voice signals to enhance conversation naturality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If digital humans and intelligence robots use fixed appearance and simple voice synthesis for conversations, then device complexity is reduced, but conversation naturality deteriorates
Solution Approach 1:
The system dynamically adjusts the digital human's appearance expressions and back-channel signals based on real-time voice signal analysis. Instead of fixed appearance, the system generates appropriate facial expressions, eye movements, and head gestures that correspond to the conversation context, making the conversation more natural while managing complexity through algorithmic control.
Solution Approach 2:
The system analyzes voice signals in real-time and generates appropriate back-channel signals (facial expressions, eye gestures, head movements) as feedback to the user. This feedback mechanism creates a more natural conversation flow by simulating human-like responsive behavior, improving conversation naturality without requiring overly complex hardware.
2Reliability
If digital humans deliver back-channel signals with emotional sympathy, then conversation naturality is improved, but device complexity increases
Solution Approach 1:
The system changes multiple parameters simultaneously including facial expressions, eye gestures, head movements, and voice tone to convey emotional sympathy. By coordinating these parameter changes based on voice signal analysis, the system achieves natural conversation with emotional depth without requiring proportionally complex hardware for each individual parameter.
Solution Approach 2:
The voice signal analysis system serves multiple functions: it detects conversation timing, determines emotional tone, identifies speaker intent, and triggers appropriate back-channel responses. This multi-functionality reduces overall system complexity by using a single analysis engine to control multiple output parameters for natural conversation.
3Productivity
If back-channel signals are generated at inappropriate timing, then conversation responsiveness is improved, but conversation naturality deteriorates
Solution Approach 1:
The system performs preliminary analysis of voice signals to determine optimal timing for back-channel signal generation. By analyzing conversation flow, pause duration, and speaker intent in advance, the system generates back-channel signals at appropriate moments that simulate natural human conversation timing, balancing responsiveness with naturality.
Data Source
AI summary
A method of generating a sympathetic back-channel signal is provided. The method includes receiving a voice signal from a user, determining whether predetermined timing is timing at which a back-channel signal is output in response to the input of the voice signal at the predetermined timing, storing the voice signal that has been input so far if the predetermined timing is the timing at which the back-channel signal is output as a result of the determination, determining back-channel signal information based on the stored voice signal, and outputting the determined back-channel signal information.


