Human-Robot Gesture Recognition With Lidar Tracking in Crowded Spaces

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Self-driving robots face challenges in densely populated environments where they struggle to differentiate between multiple people and objects, and in noisy conditions, making it difficult to accurately detect human gestures and voice commands.

Innovation Solution

The implementation of a system that uses vision-based human pose detection and Lidar tracking to identify and respond to human gestures, allowing for enhanced human-robot interaction (HRI) by associating gestures with commands and navigating through complex environments safely.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If self-driving robots operate in densely populated environments, then their application scope increases, but their ability to differentiate between multiple people and objects deteriorates

Engineering Contradiction:
Improveapplication scopeVSAvoidobject differentiation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system segments the detection task by using multiple specialized sensors (cameras, Lidar, microphones) that each handle specific aspects of environment perception. The camera captures visual information, Lidar provides depth data, and microphones detect audio signals, allowing the robot to differentiate objects more effectively in crowded environments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from two-dimensional camera images to three-dimensional spatial understanding by integrating Lidar depth data and spatial audio information. This multi-dimensional approach enables better object differentiation by adding depth and spatial context to the visual data

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If self-driving robots rely on voice command detection, then ease of operation improves, but detection accuracy deteriorates in noisy conditions

Engineering Contradiction:
Improvecommand input simplicityVSAvoidvoice command detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system merges multiple sensing modalities (visual gesture detection from cameras, spatial audio processing from microphones, and depth information from Lidar) to interpret human commands. This multi-modal approach allows the robot to cross-validate commands across different senses, improving accuracy in noisy environments while maintaining ease of operation

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces gesture recognition as an intermediary communication modality between human and robot. Gestures serve as a visual intermediary that complements voice commands, allowing the robot to disambiguate noisy audio signals through corresponding visual gestures, thereby maintaining command detection accuracy in noisy conditions

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If self-driving robots use multiple sensors for better detection, then measurement precision improves, but device complexity increases

Engineering Contradiction:
Improvegesture and object detection accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system designs sensors to serve multiple functions: cameras capture both visual appearance and gesture information, Lidar provides both depth mapping and spatial positioning, and microphones handle both ambient noise detection and voice command recognition. This multi-functionality reduces the need for separate specialized sensors, managing complexity while maintaining high detection precision

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables self-driving robots to effectively interpret and execute human commands in crowded and noisy settings, ensuring safe and socially acceptable operation by accurately identifying gestures and adapting to environmental obstacles.

Implementation Method 1

send one or more pulses; identify a cluster, based at least in part on one or more reflections associated with the one or more pulses

Methodology Applied
Scientific EffectLIDAR: LIDAR

Implementation Method 2

identify a cluster, based at least in part on one or more reflections associated with the one or more pulses

Methodology Applied
Scientific EffectReflection: Reflection

Implementation Method 3

determine, based at least in part on an image analysis of the image, a gesture associated with the object

Methodology Applied
Scientific EffectImage analysis: Image Processing

Data Source

PatentUS10948907B2Self-driving mobile robots using human-robot interactions
Publication Date: 2021.03.16 FORD GLOBAL TECH LLC
  • US10948907B2 patent drawing
  • US10948907B2 patent drawing
  • US10948907B2 patent drawing

AI summary

Systems, methods, and computer-readable media are disclosed for enhanced human-robot interactions. A device such as a robot may send one or more pulses. The device may identify one or more reflections associated with the one or more pulses. The device may determine, based at least in part on the one or more reflections, a cluster. The device may associate the cluster with an object identified in an image. The device may determine, based at least in part on an image analysis of the image, a gesture associated with the object. The device may determine, based at least in part on the gesture, a command associated with an action. The device may to perform the action.