Human-Robot Engagement Detection Using 2D Keypoints and Hidden Joints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Robotic systems lack the ability to effectively determine the extent of engagement with humans, which is crucial for coordinated interactions, as existing technologies struggle to accurately assess human intentions and poses using 2D image data efficiently.
Innovation Solution
A control system configured to identify visible and hidden keypoints in a 2D image, using a machine learning model to determine the extent of engagement based on the coordinates of visible keypoints and indicators of hidden keypoints, allowing the robot to infer human intentions and adjust its interactions accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D imaging or depth sensors are used to assess human engagement, then measurement precision may be improved, but device complexity and computational resources increase significantly
Solution Approach 1:
The patent uses a 2D image as a simplified copy or representation of the 3D scene, extracting sufficient engagement information without requiring full 3D reconstruction. The machine learning model processes 2D keypoint coordinates and visibility indicators to infer engagement states, achieving accurate assessment while avoiding the complexity of 3D imaging systems.
Solution Approach 2:
The patent extracts only the necessary features (keypoint coordinates and visibility indicators) from the 2D image that are sufficient for engagement assessment, rather than processing complete 3D spatial data. This selective extraction reduces computational requirements and system complexity while maintaining measurement precision for the specific task of engagement detection.
2Measurement precision
If complete 3D pose estimation is performed, then measurement precision is improved, but loss of time and computational resources increase
Solution Approach 1:
The patent extracts and processes only the essential elements needed for engagement assessment: 2D keypoint coordinates and visibility indicators. By eliminating the need for complete 3D pose estimation, the system achieves sufficient measurement precision for engagement detection while dramatically reducing processing time and computational resource requirements.
Solution Approach 2:
The patent applies partial action by performing only the minimum necessary processing (2D keypoint detection and visibility assessment) rather than complete 3D pose estimation. This partial processing approach provides adequate information for engagement assessment without the excessive computational burden of full 3D reconstruction.
3Measurement precision
If more comprehensive sensor data is collected, then measurement precision is improved, but device complexity and data processing requirements increase
Solution Approach 1:
The patent uses a 2D image copy from the camera as sufficient input data for engagement detection, rather than collecting comprehensive multi-sensor data. The machine learning model processes this simplified 2D representation to achieve accurate engagement assessment, avoiding the complexity of integrating multiple sensor types and processing their combined data streams.
Data Source
AI summary
A method includes receiving, from a camera disposed on a robotic device, a two-dimensional (2D) image of a body of an actor and determining, for each respective keypoint of a first subset of a plurality of keypoints, 2D coordinates of the respective keypoint within the 2D image. The plurality of keypoints represent body locations. Each respective keypoint of the first subset is visible in the 2D image. The method also includes determining a second subset of the plurality of keypoints. Each respective keypoint of the second subset is not visible in the 2D image. The method further includes determining, by way of a machine learning model, an extent of engagement of the actor with the robotic device based on (i) the 2D coordinates of keypoints of the first subset and (ii) for each respective keypoint of the second subset, an indicator that the respective keypoint is not visible.


