Robot Voice Gesture Recognition Controller

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robots require users to utter predefined words or phrases for voice input, leading to unnatural interactions and limited utilization of voice input functions, as they often need explicit keywords to recognize commands and objects.

Innovation Solution

A robot equipped with a microphone for voice recognition and a camera for gesture analysis, which identifies pointed targets and processes commands without requiring direct keywords, using a controller to integrate voice and image data for accurate object recognition and control operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the robot uses predefined words or phrases for voice input, then the robot can recognize commands accurately, but the user interaction becomes unnatural and inconvenient

Engineering Contradiction:
Improvecommand recognition accuracyVSAvoiduser interaction naturalness
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces a gesture recognition system as an intermediary between the user and the voice command system. The gesture (pointing) serves as a mediator that bridges the gap between natural speech and predefined command structures, allowing users to indicate objects naturally while the system translates the combined gesture-voice input into actionable commands

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the command recognition process into two independent components: voice command recognition and gesture object indication. This segmentation allows each component to operate independently with its own optimization - voice processing handles the command intent while gesture processing handles the object identification, reducing the need for predefined phrase structures

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If the robot requires explicit keywords for object recognition, then the object identification is accurate, but the voice input function has limited utilization

Engineering Contradiction:
Improveobject identification accuracyVSAvoidvoice input flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The gesture recognition system acts as an intermediary that translates visual pointing information into object identification, replacing the need for explicit keywords. This mediator enables users to refer to objects using natural language pronouns or descriptions while the gesture provides the specific object identification

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the linguistic mechanism (keywords) with a visual-motor mechanism (gestures). Instead of requiring users to speak specific object names or keywords, the system captures the user's pointing gesture, processes the visual information, and identifies the intended object, substituting speech-based identification with gesture-based identification

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If the robot uses only voice input without gesture recognition, then the system complexity is low, but the user interaction convenience is reduced

Engineering Contradiction:
Improvesystem structure simplicityVSAvoiduser input convenience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The robot system is designed with multi-functionality, integrating both voice recognition and gesture recognition capabilities into a unified command processing system. This universal system can handle multiple input modalities (voice alone, gesture alone, or combined voice-gesture input), providing users with flexible interaction options while maintaining a cohesive system architecture

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11338429B2Robot
Publication Date: 2022.05.24 LG ELECTRONICS INC
  • US11338429B2 patent drawing
  • US11338429B2 patent drawing
  • US11338429B2 patent drawing

AI summary

Disclosed is a robot including a microphone configured to acquire a voice, a camera configured to acquire a first image including a gesture, and a controller configured to recognize the acquired voice, recognize a pointed position corresponding to the gesture included in the first image, control the camera to acquire a second image including the recognized pointed position, identify a pointed target included in the second image, and perform a control operation on the basis of the identified pointed target and a command included in the recognized voice.