Polarization Imaging for Clean Shape Estimation in Computer Vision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing visual datasets for training computer vision models, particularly for object pose detection, often suffer from noise and inaccuracies in surface normals maps and depth maps, which can lead to suboptimal performance in shape estimation tasks.

Innovation Solution

The system generates visual datasets by capturing images using an imaging system, estimating the pose of objects, rendering shape estimates from 3D models, and creating data points with clean ground truth shape information, including surface normals maps and depth maps, using polarization imaging and physics-based deep learning techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing visual datasets are used for training computer vision models, then models can be trained on large amounts of data, but the datasets contain noise and inaccuracies in surface normals maps and depth maps that lead to suboptimal performance

Engineering Contradiction:
Improveaccuracy of shape estimatesVSAvoidnoise in surface normals maps and depth maps
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system performs preliminary actions by capturing multiple polarization images at different angles before processing. These pre-captured polarization images serve as the foundation for computing accurate surface normals maps and depth maps, allowing the system to prepare high-quality training data in advance and avoid the noise and inaccuracies present in existing datasets

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical imaging systems with polarization imaging technology. By using polarization cameras to capture images at multiple angles and applying physics-based deep learning techniques, the system generates cleaner surface normals maps and depth maps without the noise and inaccuracies inherent in conventional imaging approaches

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If polarization imaging and physics-based deep learning techniques are used, then cleaner shape estimates can be generated, but the system complexity increases

Engineering Contradiction:
Improveaccuracy of shape estimatesVSAvoidimaging system and processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The polarization camera system serves multiple functions simultaneously: it captures images at different polarization angles, provides data for surface normals computation, and enables depth estimation. This multi-functionality reduces the need for separate specialized devices, thereby managing system complexity while achieving high measurement precision

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system changes the parameter space by capturing images at multiple polarization angles rather than using a single traditional image. This parameter transformation allows the same physical data to be processed through physics-based deep learning techniques to generate clean surface normals and depth maps, achieving improved accuracy without proportionally increasing hardware complexity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If 3D models are used for pose estimation, then accurate shape information can be obtained, but the system cannot handle unknown objects without pre-existing 3D models

Engineering Contradiction:
Improveaccuracy of shape estimatesVSAvoidability to handle unknown objects
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

Instead of relying on pre-existing 3D models, the system creates new 3D representations by copying and synthesizing surface normals maps and depth maps from polarization images. This allows the system to generate accurate shape information for unknown objects by constructing their 3D models from captured image data rather than requiring pre-programmed models

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary computation of surface normals maps and depth maps from polarization images before final pose estimation. This preliminary action creates the necessary shape information on-the-fly, enabling the system to handle unknown objects by generating their 3D representations in advance without requiring pre-existing models

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach results in datasets with lower noise and higher accuracy in shape estimates, enabling computer vision models to more accurately predict the shapes of objects, even for unknown objects without pre-existing 3D models.

Implementation Method 1

The imaging system may include a polarization camera system, and the one or more input images may include one or more polarization images

Methodology Applied
Scientific EffectPolarization: Polarisation

Implementation Method 2

The one or more polarization images may include a plurality of spectral channels corresponding to different portions of an electromagnetic spectrum

Methodology Applied
Scientific EffectElectromagnetic radiation detection: Absorption (EM radiation)

Data Source

PatentUS12340538B2Systems and methods for generating and using visual datasets for training computer vision models
Publication Date: 2025.06.24 INTRINSIC INNOVATION LLC
  • US12340538B2 patent drawing
  • US12340538B2 patent drawing
  • US12340538B2 patent drawing

AI summary

A system for collecting data for training a computer vision model for shape estimation includes: an imaging system configured to capture one or more images; and a processing system including a processor and memory storing instructions that, when executed by the processor, cause the processor to: receive one or more input images from the imaging system; estimate a pose of an object depicted in the one or more images; render a shape estimate from a 3-D model of the object posed in accordance with the pose of the object; and generate a data point of a training dataset, the data point including one or more images based on the one or more input images and a label corresponding to the one or more images, the label including the shape estimate.