Communication Robot Voice Reception Direction Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing communication robots face challenges in accurately performing speech recognition when there are multiple utterers nearby, as they often misidentify the direction of continuous word utterers after detecting a wake-up word, leading to degraded speech recognition performance.

Innovation Solution

The communication robot employs a method that analyzes both spoken speech and photographed images to determine the voice reception enhancement direction, using techniques like TDOA for wake-up word detection and gaze tracking and lipreading for continuous word identification, thereby distinguishing between wake-up and continuous word utterers and adjusting sensitivity accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the robot enhances sensitivity in the direction of the wake-up word utterer, then wake-up word detection is improved, but continuous word recognition by other utterers deteriorates

Engineering Contradiction:
Improvewake-up word detection accuracyVSAvoidcontinuous word recognition accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The robot dynamically adjusts the voice reception enhancement direction based on real-time identification of which utterer is speaking the continuous word. The system transitions from a static enhancement direction (fixed after wake-up word detection) to a dynamic one that can be reconfigured during operation, allowing the robot to switch between enhancing the wake-up word utterer's direction and other directions as needed for accurate continuous word recognition

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses image recognition to provide feedback about the current speaker's identity and position. This feedback loop allows the robot to continuously monitor which utterer is speaking and adjust the voice reception enhancement direction accordingly, preventing the degradation of continuous word recognition accuracy that would occur with a fixed enhancement direction

Inventive Principle:
Principle #23Feedback

2Device complexity

If the robot uses speech recognition alone, then processing is simple, but accuracy deteriorates when multiple utterers are present

Engineering Contradiction:
Improveprocessing system simplicityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system merges image recognition technology with speech recognition technology to create a multi-modal processing system. The image recognition component identifies which utterer is speaking and provides spatial information, while the speech recognition component processes the audio input. By combining these two technologies, the system achieves accurate speech recognition in multi-utterer scenarios without requiring overly complex processing

Inventive Principle:
Principle #5Merging (Combining)

3Speed

If the robot constantly activates speech recognition, then response speed is improved, but power consumption increases

Engineering Contradiction:
Improvespeech response speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary action by constantly monitoring for wake-up words in the background without fully activating speech recognition. When a wake-up word is detected, the system then activates full speech recognition processing. This preliminary monitoring approach allows the robot to maintain readiness for user commands while consuming less power than constant full speech recognition activation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11217246B2Communication robot and method for operating the same
Publication Date: 2022.01.04 LG ELECTRONICS INC
  • US11217246B2 patent drawing
  • US11217246B2 patent drawing
  • US11217246B2 patent drawing

AI summary

Disclosed are a communication robot and a method for operating the same capable of smoothly processing speech recognition by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm in 5G environment connected for Internet of things. A method for operating a communication robot according to an embodiment of the present disclosure may include collecting speech uttered by two or more utterers approaching within a predetermined distance from the communication robot, collecting photographed images of the two or more utterers, determining whether a case where utterers of a wake-up word and a continuous word included in the uttered speech are the same is a first case, or whether a case where the utterers of the wake-up and the continuous word included in the uttered speech are different is a second case, and determining a voice reception enhancement direction according to the first case or the second case.