3-D Visual Phrases for Robust Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional object recognition techniques using 2-D visual phrases struggle with accurate recognition due to limited discriminative ability and inability to handle projective transformations caused by viewpoint changes, leading to poor performance in images taken from different angles or distances.
Innovation Solution
The development of 3-D visual phrases based on a 3-D object model, which characterizes the visual appearance and geometric relationships of triangular facets on the object's surface, enabling robust object recognition across varying viewpoints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If 2-D visual phrases are used for object recognition, then the recognition process is simple, but the discriminative ability is limited and accuracy deteriorates under viewpoint changes
Solution Approach 1:
The patent transitions from 2-D visual phrases to 3-D visual phrases by introducing depth information through a 3-D object model. Each visual phrase is represented as a triangular facet with three vertices in 3-D space, allowing the system to capture spatial structure and maintain accuracy under viewpoint changes while managing complexity through efficient 3-D to 2-D projection
2Productivity
If 2-D visual phrases consider only localized region co-occurrence, then the processing is computationally efficient, but the ability to handle projective transformations deteriorates
Solution Approach 1:
By representing visual phrases as 3-D triangular facets with vertices having (x, y, z) coordinates, the system captures spatial relationships that are invariant to viewpoint changes. The 3-D structure allows proper handling of projective transformations while maintaining computational efficiency through optimized 3-D to 2-D projection and matching algorithms
Solution Approach 2:
The patent changes the parameter representation from 2-D coordinates to 3-D coordinates, adding the depth dimension (z-coordinate) to each visual phrase vertex. This parameter expansion enables the system to distinguish between points that may appear similar in 2-D but have different spatial relationships in 3-D, improving robustness to viewpoint changes
Data Source
AI summary
The techniques discussed herein discover three-dimensional (3-D) visual phrases for an object based on a 3-D model of the object. The techniques then describe the 3-D visual phrases. Once described, the techniques use the 3-D visual phrases to detect the object in an image (e.g., object recognition).


