3D Shape Extraction From Unannotated Images Using Keypoint Correspondence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for extracting 3D shape data from images rely on controlled and coherent image datasets with metadata, failing to effectively handle diverse and unstructured datasets with varying textures, backgrounds, and object variations.

Innovation Solution

A system that identifies robust keypoint correspondences across images using an image encoder, keypoint extractor, and machine learning models to predict an occupancy field, which is optimized with camera representation, enabling 3D shape extraction from unstructured datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional techniques (stereopsis, structure from motion, photogrammetry) are used to extract 3D shape data, then controlled input data with custom and highly coherent series of images can be processed, but the system cannot handle unstructured and unannotated image datasets with diverse textures, illuminations, shapes, and environments

Engineering Contradiction:
Improveability to handle diverse image datasetsVSAvoidcomplexity of image processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms the input data parameters from controlled, annotated images to unstructured, unannotated images by changing the fundamental assumptions about input data quality. The system adapts to varying illumination, texture, and viewpoint parameters without requiring controlled acquisition conditions.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary representation (occupancy field with positional encodings) that mediates between the unstructured image inputs and the 3D shape output. This intermediary allows the system to handle diverse inputs by encoding spatial information in a format that is invariant to illumination, texture, and viewpoint variations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple images from different perspectives are used to infer depth and reconstruct 3D structure, then 3D shape data can be extracted, but the system requires controlled input data with coherent image series for each shape

Engineering Contradiction:
Improveaccuracy of 3D shape extractionVSAvoidflexibility with unannotated datasets
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent enables the system to be self-sufficient by automatically learning keypoint correspondences and camera poses from unannotated images without requiring manual labeling or controlled data acquisition. The system serves itself by extracting structure from motion and geometry directly from the image set.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary actions by pre-processing images to extract features and identify keypoint correspondences before 3D reconstruction. This preliminary feature extraction and correspondence identification prepares the unstructured data in a format suitable for subsequent occupancy field prediction and 3D shape generation.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If traditional 3D reconstruction methods are applied to unstructured images, then processing of diverse visual data becomes possible, but the system lacks mechanisms to identify keypoint correspondences and handle viewpoint variations

Engineering Contradiction:
Improvehandling of unstructured image dataVSAvoidaccuracy of shape reconstruction
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transitions from 2D image space to 3D occupancy field space by introducing positional encodings that represent spatial coordinates in three dimensions. This dimensional transformation allows the system to reconstruct 3D shapes from 2D unstructured images by predicting occupancy values at each 3D point.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent replaces traditional mechanical 3D reconstruction methods (which rely on controlled camera movements and annotated features) with a machine learning-based occupancy network that automatically learns geometric relationships from unstructured image data through training on large datasets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12633057B2Extracting 3D shapes from large-scale unannotated image datasets
Publication Date: 2026.05.19 ADOBE INC
  • US12633057B2 patent drawing
  • US12633057B2 patent drawing
  • US12633057B2 patent drawing

AI summary

Systems and methods for extracting 3D shapes from unstructured and unannotated datasets are described. Embodiments are configured to obtain a first image and a second image, where the first image depicts an object and the second image includes a corresponding object of a same object category as the object. Embodiments are further configured to generate, using an image encoder, image features for portions of the first image and for portions of the second image; identify a keypoint correspondence between a first keypoint in the first image and a second keypoint in the second image by clustering the image features corresponding to the portions of the first image and the portions of the second image; and generate, using an occupancy network, a 3D model of the object based on the keypoint correspondence.