Robot Engagement Estimation from 2D Human Keypoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robotic systems lack the ability to effectively determine the extent of engagement with humans, which is crucial for coordinated interaction, as existing methods are inefficient in interpreting human cues from 2D images.

Innovation Solution

A control system configured to identify visible and hidden keypoints in 2D images using a machine learning model, which determines the extent of engagement by processing 2D image data, allowing the robot to infer hidden keypoints' positions and adjust its interactions accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 3D imaging or depth sensors are used to determine engagement, then measurement precision improves, but device complexity and cost increase

Engineering Contradiction:
Improveengagement detection accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses a 2D image camera to capture visual information and creates a simplified representation (keypoints) that copies essential spatial relationships. This allows engagement detection using standard 2D imaging rather than requiring complex 3D sensors, resolving the contradiction by achieving sufficient measurement precision through intelligent processing of simpler data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces physical 3D sensing hardware with a computational approach using 2D images and machine learning models. The mechanical/optical complexity of depth sensors is substituted with algorithmic processing of standard camera data, reducing device complexity while maintaining engagement detection capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If more keypoints are processed to improve engagement detection accuracy, then measurement precision improves, but computational complexity increases

Engineering Contradiction:
Improveengagement detection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential spatial information needed for engagement detection by identifying and processing specific keypoints in the image. Rather than analyzing all pixels or comprehensive 3D data, it selectively processes key landmark points, reducing computational complexity while maintaining measurement precision

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the image processing task by dividing it into discrete keypoint detection and coordinate extraction steps. This segmentation allows the system to focus computational resources on critical features rather than processing the entire image uniformly, balancing accuracy with computational efficiency

Inventive Principle:
Principle #1Segmentation

3Device complexity

If 2D image data is used instead of 3D data, then device complexity decreases, but loss of information about depth and spatial relationships increases

Engineering Contradiction:
Improvesensor system complexityVSAvoiddepth information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent compensates for the loss of depth information in 2D images by utilizing the spatial dimension of keypoint coordinates within the 2D plane. By carefully selecting and processing keypoint positions, the system extracts sufficient spatial relationship information from 2D data to determine engagement, avoiding the need for 3D sensors

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11915523B2Engagement detection and attention estimation for human-robot interaction
Publication Date: 2024.02.27 GDM HOLDING LLC
  • US11915523B2 patent drawing
  • US11915523B2 patent drawing
  • US11915523B2 patent drawing

AI summary

A method includes receiving, from a camera disposed on a robotic device, a two-dimensional (2D) image of a body of an actor and determining, for each respective keypoint of a first subset of a plurality of keypoints, 2D coordinates of the respective keypoint within the 2D image. The plurality of keypoints represent body locations. Each respective keypoint of the first subset is visible in the 2D image. The method also includes determining a second subset of the plurality of keypoints. Each respective keypoint of the second subset is not visible in the 2D image. The method further includes determining, by way of a machine learning model, an extent of engagement of the actor with the robotic device based on (i) the 2D coordinates of keypoints of the first subset and (ii) for each respective keypoint of the second subset, an indicator that the respective keypoint is not visible.