Visuo-Tactile Object Pose Estimation Under Gripper Occlusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current robotics face challenges in accurately estimating the pose of objects due to reliance on open-loop grasp synthesis without feedback, especially in noisy and dynamic environments, leading to poor success rates, especially when visual data is occluded by grasp devices.
Innovation Solution
A system and method for visuo-tactile object pose estimation that combines image data, depth data, and tactile data to generate a visual estimate and a tactile estimate, which are then fused in a 3D space to estimate the object's six-dimensional pose, including location and orientation, using a network architecture that separates visual and tactile channels to account for occlusions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If open-loop grasp synthesis is used without feedback, then the system is simpler to implement, but the success rate of object manipulation deteriorates in noisy and dynamic environments
Solution Approach 1:
The patent implements a closed-loop system that uses tactile sensors to provide real-time feedback during grasp execution. The tactile information is processed to generate feedback signals that allow the system to detect and correct deviations from the planned grasp, thereby maintaining high success rates in dynamic environments while managing complexity through selective feedback implementation
2Ease of operation
If grasp devices occlude the object, then the grasp can be performed, but visual data of the object is lost for at least a portion of the object
Solution Approach 1:
The patent merges visual data from cameras with tactile data from sensors mounted on the grasp device. By combining these complementary data sources, the system compensates for visual information occluded by the grasp device, maintaining awareness of the object's full geometry and pose during manipulation
Solution Approach 2:
The patent introduces tactile sensors as an intermediary that indirectly measures properties of occluded object portions. These sensors act as mediators between the grasp device and the occluded object surfaces, providing information about contact forces, surface geometry, and object pose that would otherwise be invisible to visual sensors
3Device complexity
If only visual data is used for pose estimation, then the system is simpler, but the pose estimation accuracy deteriorates when objects are occluded or in noisy environments
Solution Approach 1:
The patent merges visual data from cameras with tactile data from sensors mounted on the grasp device. By combining these complementary data sources, the system compensates for visual information occluded by the grasp device, maintaining awareness of the object's full geometry and pose during manipulation
Solution Approach 2:
The patent creates a composite sensing system that integrates multiple sensor types (visual and tactile) with different measurement characteristics. This composite approach leverages the strengths of each sensor modality to achieve robust pose estimation that is insensitive to the weaknesses of individual sensors, particularly in occluded or noisy conditions
Data Source
AI summary
Systems and methods for visuo-tactile object pose estimation are provided. In one embodiment, a computer implemented method includes receiving image data, depth data, and tactile data about the object in the environment. The computer implemented method also includes generating a visual estimate of the object that includes an object point cloud. The computer implemented method further includes generating a tactile estimate of the object that includes a surface point cloud based on the tactile data. The computer implemented method yet further includes estimating a pose of the object based on the visual estimate and the tactile estimate by fusing the object point cloud and the surface point cloud in a 3D space. The pose is a six-dimensional pose.


