AR Hand-Object Pose Labeling Without Occlusion Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing 3D hand-object interaction datasets fail to adequately cover diversified real-world scenarios, leading to poor performance in specific applications due to hand-object occlusions and the need for task-specified datasets, and current dataset collection methods are inefficient and impractical for ordinary users.
Innovation Solution
A method using augmented reality (AR) to create virtual bounding boxes for physical objects, allowing users to record hand and object pose labels sequentially, decoupling the labeling process from physical interaction to overcome occlusions and enable efficient dataset collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and sensors are placed in laboratory environments to obtain 3D poses, then the occlusion issue is solved and both images and labels can be generated concurrently, but additional hardware setups are required which reduces accessibility for ordinary users
Solution Approach 1:
The patent uses a virtual copy (digital twin) of the physical object created through photogrammetry to perform the labeling task. Instead of requiring multiple cameras to capture real 3D poses, the system creates a virtual model and uses 2D image annotations to infer 3D poses on this virtual copy, eliminating the need for complex multi-camera hardware while maintaining labeling accuracy.
Solution Approach 2:
The patent introduces a virtual bounding box and virtual model as an intermediary between the physical object and the labeling process. The virtual bounding box serves as a mediator that can be manipulated in 2D space to represent 3D pose information, allowing annotators to work with simplified 2D interfaces while accurately capturing 3D interaction data.
2Ease of operation
If post-hoc 2D labeling interfaces are used to label 3D poses after image capture, then the process becomes feasible for ordinary users, but hand-object occlusions prevent annotators from labeling hand joints hidden behind objects
Solution Approach 1:
The patent creates a virtual copy of the hand-object interaction scene where occluded hand joints can be inferred and labeled without visual obstruction. Annotators work with the virtual model and bounding box representations where occlusion is not an issue, maintaining both user accessibility and labeling accuracy.
Solution Approach 2:
The patent transitions the labeling task from direct 2D image annotation to 3D virtual space annotation. By moving the labeling interface to 3D virtual space with virtual bounding boxes that can be manipulated spatially, the system allows annotators to label hand joints that would be occluded in 2D images, effectively adding a dimensional perspective that resolves the occlusion problem.
3Productivity
If pre-trained networks are used for 3D pose estimation, then the system can be deployed quickly, but performance deteriorates when specific application contexts or object types differ from training data
Solution Approach 1:
The patent performs preliminary actions by collecting and annotating domain-specific data through the virtual reality interface before training the final model. Users pre-collect interaction data for their specific objects and tasks using the virtual bounding box manipulation interface, creating a tailored dataset that preserves both deployment efficiency and application-specific accuracy.
Solution Approach 2:
The patent enables local quality by allowing customization of datasets for specific objects, tasks, and application domains. Instead of using generic pre-trained models, the system facilitates collection of domain-specific data with virtual reality interfaces tailored to particular interaction scenarios, ensuring high reliability for target applications while maintaining reasonable deployment speed through efficient annotation processes.
Data Source
AI summary
A method and system for hand-object interaction dataset collection is described herein, which is configured to support user-specified collection of hand-object interaction datasets. Such hand-object interaction datasets are useful, for example, for training 3D hand and object pose estimation model. The method and system adopt a sequential process of first recording hand and object pose labels by manipulating a virtual bounding box, rather than a physical object. Naturally, hand-object occlusions do not occur during the manipulation of the virtual bounding box, so these labels are provided with high accuracy. Subsequently, the images are separately captured of the hand-object interaction with the physical object. These images are paired with the previously recorded hand and object pose labels.


