Robot Engagement Estimation from 2D Human Keypoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robotic systems lack the ability to effectively determine the extent of engagement with humans, which is crucial for coordinated interaction, as existing methods are inefficient in interpreting human cues from 2D images.
Innovation Solution
A control system configured to identify visible and hidden keypoints in 2D images using a machine learning model, which determines the extent of engagement by processing 2D image data, allowing the robot to infer hidden keypoints' positions and adjust its interactions accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D imaging or depth sensors are used to determine engagement, then measurement precision improves, but device complexity and cost increase
Solution Approach 1:
The patent uses a 2D image camera to capture visual information and creates a simplified representation (keypoints) that copies essential spatial relationships. This allows engagement detection using standard 2D imaging rather than requiring complex 3D sensors, resolving the contradiction by achieving sufficient measurement precision through intelligent processing of simpler data
Solution Approach 2:
The patent replaces physical 3D sensing hardware with a computational approach using 2D images and machine learning models. The mechanical/optical complexity of depth sensors is substituted with algorithmic processing of standard camera data, reducing device complexity while maintaining engagement detection capability
2Measurement precision
If more keypoints are processed to improve engagement detection accuracy, then measurement precision improves, but computational complexity increases
Solution Approach 1:
The patent extracts only the essential spatial information needed for engagement detection by identifying and processing specific keypoints in the image. Rather than analyzing all pixels or comprehensive 3D data, it selectively processes key landmark points, reducing computational complexity while maintaining measurement precision
Solution Approach 2:
The patent segments the image processing task by dividing it into discrete keypoint detection and coordinate extraction steps. This segmentation allows the system to focus computational resources on critical features rather than processing the entire image uniformly, balancing accuracy with computational efficiency
3Device complexity
If 2D image data is used instead of 3D data, then device complexity decreases, but loss of information about depth and spatial relationships increases
Solution Approach 1:
The patent compensates for the loss of depth information in 2D images by utilizing the spatial dimension of keypoint coordinates within the 2D plane. By carefully selecting and processing keypoint positions, the system extracts sufficient spatial relationship information from 2D data to determine engagement, avoiding the need for 3D sensors
Data Source
AI summary
A method includes receiving, from a camera disposed on a robotic device, a two-dimensional (2D) image of a body of an actor and determining, for each respective keypoint of a first subset of a plurality of keypoints, 2D coordinates of the respective keypoint within the 2D image. The plurality of keypoints represent body locations. Each respective keypoint of the first subset is visible in the 2D image. The method also includes determining a second subset of the plurality of keypoints. Each respective keypoint of the second subset is not visible in the 2D image. The method further includes determining, by way of a machine learning model, an extent of engagement of the actor with the robotic device based on (i) the 2D coordinates of keypoints of the first subset and (ii) for each respective keypoint of the second subset, an indicator that the respective keypoint is not visible.


