3D Shape Extraction From Unannotated Images Using Keypoint Correspondence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques for extracting 3D shape data from images rely on controlled and coherent image datasets with metadata, failing to effectively handle diverse and unstructured datasets with varying textures, backgrounds, and object variations.
Innovation Solution
A system that identifies robust keypoint correspondences across images using an image encoder, keypoint extractor, and machine learning models to predict an occupancy field, which is optimized with camera representation, enabling 3D shape extraction from unstructured datasets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional techniques (stereopsis, structure from motion, photogrammetry) are used to extract 3D shape data, then controlled input data with custom and highly coherent series of images can be processed, but the system cannot handle unstructured and unannotated image datasets with diverse textures, illuminations, shapes, and environments
Solution Approach 1:
The patent transforms the input data parameters from controlled, annotated images to unstructured, unannotated images by changing the fundamental assumptions about input data quality. The system adapts to varying illumination, texture, and viewpoint parameters without requiring controlled acquisition conditions.
Solution Approach 2:
The patent introduces an intermediary representation (occupancy field with positional encodings) that mediates between the unstructured image inputs and the 3D shape output. This intermediary allows the system to handle diverse inputs by encoding spatial information in a format that is invariant to illumination, texture, and viewpoint variations.
2Measurement precision
If multiple images from different perspectives are used to infer depth and reconstruct 3D structure, then 3D shape data can be extracted, but the system requires controlled input data with coherent image series for each shape
Solution Approach 1:
The patent enables the system to be self-sufficient by automatically learning keypoint correspondences and camera poses from unannotated images without requiring manual labeling or controlled data acquisition. The system serves itself by extracting structure from motion and geometry directly from the image set.
Solution Approach 2:
The patent performs preliminary actions by pre-processing images to extract features and identify keypoint correspondences before 3D reconstruction. This preliminary feature extraction and correspondence identification prepares the unstructured data in a format suitable for subsequent occupancy field prediction and 3D shape generation.
3Adaptability or versatility
If traditional 3D reconstruction methods are applied to unstructured images, then processing of diverse visual data becomes possible, but the system lacks mechanisms to identify keypoint correspondences and handle viewpoint variations
Solution Approach 1:
The patent transitions from 2D image space to 3D occupancy field space by introducing positional encodings that represent spatial coordinates in three dimensions. This dimensional transformation allows the system to reconstruct 3D shapes from 2D unstructured images by predicting occupancy values at each 3D point.
Solution Approach 2:
The patent replaces traditional mechanical 3D reconstruction methods (which rely on controlled camera movements and annotated features) with a machine learning-based occupancy network that automatically learns geometric relationships from unstructured image data through training on large datasets.
Data Source
AI summary
Systems and methods for extracting 3D shapes from unstructured and unannotated datasets are described. Embodiments are configured to obtain a first image and a second image, where the first image depicts an object and the second image includes a corresponding object of a same object category as the object. Embodiments are further configured to generate, using an image encoder, image features for portions of the first image and for portions of the second image; identify a keypoint correspondence between a first keypoint in the first image and a second keypoint in the second image by clustering the image features corresponding to the portions of the first image and the portions of the second image; and generate, using an occupancy network, a 3D model of the object based on the keypoint correspondence.


