Intelligent Robot Speech Pace Adaptation for User Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing intelligent question-answering robot systems respond with a preset speech pace, failing to adjust intelligently to match individual users' speech pace habits, leading to poor or uncomfortable communication.
Innovation Solution
A multimodal human-machine collaborative interaction method that collects monitoring videos from intelligent question-answering robots, extracts human voices and speech paces, adjusts response speech paces based on semantic relationships and key coefficients, and synthesizes a response voice to match individual users' speech habits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the robot uses a preset speech pace for responses, then the system complexity is reduced, but the communication effectiveness deteriorates because it cannot match individual users' speech pace habits
Solution Approach 1:
The robot's speech pace is transformed from a static preset value to a dynamic parameter that automatically adjusts based on real-time analysis of user speech patterns. The system continuously monitors user speech speed and adapts the robot's response pace accordingly, enabling flexible adaptation to different users without manual configuration.
Solution Approach 2:
The system implements feedback by monitoring user speech pace in real-time and using this information to adjust the robot's response pace. The feedback loop involves: (1) capturing user speech, (2) analyzing speech pace, (3) adjusting robot response pace based on the analysis, and (4) continuously optimizing the interaction based on the adjusted performance.
2Productivity
If the robot adjusts speech pace to match each user, then communication effectiveness improves, but the processing time and computational resources increase
Solution Approach 1:
The system performs preliminary action by analyzing user speech pace early in the interaction and using this information to pre-adjust the robot's response parameters before the actual conversation begins. This allows the robot to be optimized for each user's preferences from the start, reducing the need for continuous real-time adjustments throughout the interaction.
Solution Approach 2:
The system changes the speech pace parameter dynamically based on user characteristics. By detecting user speech pace and transforming it into a guide for robot response, the system optimizes the temporal parameters of communication to match individual user preferences, thereby improving interaction efficiency without requiring excessive processing resources.
3Ease of operation
If the robot uses fixed speech pace, then the system is easier to implement, but the communication quality deteriorates due to mismatch with user preferences
Solution Approach 1:
The system implements self-service by automatically detecting user speech pace and adjusting its own response parameters without requiring external intervention or manual programming for each user. The robot serves itself by autonomously adapting to different users' preferences, maintaining ease of implementation while significantly improving communication quality through automatic personalization.
Data Source
AI summary
The present invention discloses a multimodal human-machine collaborative interaction system and method, and relates to the technical field of human-machine collaborative interaction. The method includes: collecting a monitoring video captured by an intelligent question-answering robot and having a moving object in a picture, and extracting a target moving object and a human voice corresponding to the target moving object; obtaining a human speech pace corresponding to the human voice, obtaining a response text corresponding to the human voice, and obtaining a key coefficient corresponding to each keyword in the response text; obtaining a response speech pace of each keyword; and adjusting a response speech pace between any two adjacent keywords according to the two adjacent keywords, and then obtaining a response voice corresponding to the response text.

