Guide Robot User Intent Recognition via Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing guide robots require a wake-up word or specific gesture for voice interaction, which can be inconvenient for new users or those without prior knowledge of these triggers, and often lead to misrecognition issues.
Innovation Solution
A guide robot capable of recognizing a user's intention to speak without a wake-up word, using sensors and cameras to detect user proximity and face angle, and triggering a voice conversation mode accordingly, allowing for customizable wake-up words and noise filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a wake-up word recognition system is used to activate voice conversation, then voice interaction can be triggered intentionally, but new users cannot interact without prior knowledge of the wake-up word
Solution Approach 1:
The system performs preliminary detection of user approach and intent before requiring wake-up word interaction. The sensor detects when a user approaches the robot, and the camera captures facial expressions and gestures to determine interaction intent in advance, eliminating the need for users to know wake-up words beforehand
Solution Approach 2:
The system introduces sensors and camera as intermediary detection means between the user and the voice conversation activation. These intermediaries detect user presence and intent, serving as a bridge that translates physical presence into system activation without requiring verbal wake-up words
2Speed
If wake-up word detection is always active to enable quick interaction, then voice conversation can be activated quickly, but noise and unintentional voice cause misrecognition
Solution Approach 1:
The system performs preliminary detection of user presence through sensors and preliminary analysis of facial expressions and gestures through the camera before activating voice conversation. This preliminary action filters out unintentional voices and noise by confirming genuine user intent before activation
Solution Approach 2:
The system uses feedback from sensor detection and camera analysis to determine whether to activate voice conversation. The feedback loop analyzes user approach, facial expressions, and gestures to confirm intentional interaction, reducing misrecognition from noise and unintentional voices
3Ease of operation
If multiple sensors and cameras are used to detect user intent without wake-up words, then user interaction becomes more intuitive, but device complexity increases
Solution Approach 1:
The existing sensors and camera in the robot are made multi-functional. The sensors that originally detected obstacles are now also used to detect user approach, and the camera that captured environment information is now used to analyze facial expressions and gestures. This eliminates the need for additional dedicated components
Data Source
AI summary
A guide robot can include a travel part to move the guide robot, a touch screen and a camera, a sensor to detect an approach of a user, and a voice reception part to receive a voice. The guide robot further includes a controller to display at least one digital signage while the guide robot is traveling, in response to detecting the approach of the user, stop the traveling of the guide robot and transition the camera from a deactivated state to an activated state, and detect a face and a face angle of the user. Also, in response to determining that the user intends to use the guide robot, the controller can trigger a voice conversation mode by activating the voice reception part, stopping the display of the at least one digital signage and outputting usage guide information for the voice conversation mode.


