Communication Robot Voice Reception Direction Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing communication robots face challenges in accurately performing speech recognition when there are multiple utterers nearby, as they often misidentify the direction of continuous word utterers after detecting a wake-up word, leading to degraded speech recognition performance.
Innovation Solution
The communication robot employs a method that analyzes both spoken speech and photographed images to determine the voice reception enhancement direction, using techniques like TDOA for wake-up word detection and gaze tracking and lipreading for continuous word identification, thereby distinguishing between wake-up and continuous word utterers and adjusting sensitivity accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the robot enhances sensitivity in the direction of the wake-up word utterer, then wake-up word detection is improved, but continuous word recognition by other utterers deteriorates
Solution Approach 1:
The robot dynamically adjusts the voice reception enhancement direction based on real-time identification of which utterer is speaking the continuous word. The system transitions from a static enhancement direction (fixed after wake-up word detection) to a dynamic one that can be reconfigured during operation, allowing the robot to switch between enhancing the wake-up word utterer's direction and other directions as needed for accurate continuous word recognition
Solution Approach 2:
The system uses image recognition to provide feedback about the current speaker's identity and position. This feedback loop allows the robot to continuously monitor which utterer is speaking and adjust the voice reception enhancement direction accordingly, preventing the degradation of continuous word recognition accuracy that would occur with a fixed enhancement direction
2Device complexity
If the robot uses speech recognition alone, then processing is simple, but accuracy deteriorates when multiple utterers are present
Solution Approach 1:
The system merges image recognition technology with speech recognition technology to create a multi-modal processing system. The image recognition component identifies which utterer is speaking and provides spatial information, while the speech recognition component processes the audio input. By combining these two technologies, the system achieves accurate speech recognition in multi-utterer scenarios without requiring overly complex processing
3Speed
If the robot constantly activates speech recognition, then response speed is improved, but power consumption increases
Solution Approach 1:
The system performs preliminary action by constantly monitoring for wake-up words in the background without fully activating speech recognition. When a wake-up word is detected, the system then activates full speech recognition processing. This preliminary monitoring approach allows the robot to maintain readiness for user commands while consuming less power than constant full speech recognition activation
Data Source
AI summary
Disclosed are a communication robot and a method for operating the same capable of smoothly processing speech recognition by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm in 5G environment connected for Internet of things. A method for operating a communication robot according to an embodiment of the present disclosure may include collecting speech uttered by two or more utterers approaching within a predetermined distance from the communication robot, collecting photographed images of the two or more utterers, determining whether a case where utterers of a wake-up word and a continuous word included in the uttered speech are the same is a first case, or whether a case where the utterers of the wake-up and the continuous word included in the uttered speech are different is a second case, and determining a voice reception enhancement direction according to the first case or the second case.


