3D Object Recognition Using Pyramid View Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine vision systems face challenges in recognizing 3D objects with non-planar shapes imaged from unknown viewpoints, as they struggle with robust feature extraction under varying perspectives, occlusions, and clutter, and require extensive training data for view-based recognition, limiting their applicability in industrial settings.
Innovation Solution
A method that constructs a 3D model using a geometric representation of the object, samples views across different image resolutions, and organizes them in a tree structure, allowing for efficient matching and pose determination through a pyramid approach, incorporating camera calibration and texture augmentation for improved robustness and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If feature-based techniques are used to recognize 3D objects from unknown viewpoints, then the range of viewpoints can be extended, but the robustness to occlusions, clutter, and perspective distortions deteriorates
Solution Approach 1:
The patent segments the 3D object into multiple distinct features or landmarks that can be individually tracked. By dividing the object recognition task into feature detection, feature matching, and pose estimation stages, the system can handle occlusions and clutter more effectively while maintaining viewpoint flexibility.
Solution Approach 2:
The patent employs invariant feature detection methods that change parameters such as scale, rotation, and illumination to maintain feature detectability across different viewpoints. By using parameters that remain constant under perspective transformations, the system achieves both wide viewpoint adaptability and robustness to environmental variations.
2Reliability
If view-based techniques are used to recognize 3D objects, then robustness to viewpoint changes is improved, but the need for extensive training data increases
Solution Approach 1:
The patent performs preliminary action by pre-computing and storing invariant feature descriptors and their corresponding 3D coordinates in a database. This preprocessing step eliminates the need for extensive training data during actual recognition, as the system can directly match detected features against the pre-built database across various viewpoints.
Solution Approach 2:
The patent creates simplified 2D projections or renderings of the 3D object model from multiple predefined viewpoints and stores them as reference patterns. During recognition, these pre-generated views are compared with the input image, providing robustness to viewpoint changes without requiring large amounts of training data for each specific viewpoint.
3Measurement precision
If manual feature selection is used in template matching, then the matching accuracy can be improved, but the flexibility with regard to changing objects deteriorates
Solution Approach 1:
The patent develops a universal feature detection and description framework that can automatically extract meaningful features from any 3D object regardless of its specific geometry or texture. This multi-functional approach maintains high matching accuracy while providing flexibility to handle different object types without manual reconfiguration.
Solution Approach 2:
The patent implements self-service by enabling the system to automatically select and extract features from new objects without manual intervention. The invariant feature detection algorithms automatically identify salient points and compute their descriptors, allowing the system to adapt to changing objects while maintaining matching precision.
Data Source
AI summary
The present invention provides a system and method for recognizing a 3D object in a single camera image and for determining the 3D pose of the object with respect to the camera coordinate system. In one typical application, the 3D pose is used to make a robot pick up the object. A view-based approach is presented that does not show the drawbacks of previous methods because it is robust to image noise, object occlusions, clutter, and contrast changes. Furthermore, the 3D pose is determined with a high accuracy. Finally, the presented method allows the recognition of the 3D object as well as the determination of its 3D pose in a very short computation time, making it also suitable for real-time applications. These improvements are achieved by the methods disclosed herein.


