2D Imaging System Using ML Feature Maps for Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Industrial imaging systems face challenges in reliably detecting and tracking objects with varying structures and poses, particularly when using conventional 2D image data, as they require large manually labeled datasets and can become outdated due to environmental changes, and often rely on expensive depth-sensing cameras.

Innovation Solution

A computer-implemented method and system that uses a machine learning-based feature generation model to identify corresponding points in images by generating pixel feature vectors and feature maps, allowing for efficient detection and tracking without depth data, using conventional 2D image capture devices and enabling ongoing updates with environmental feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If conventional 2D imaging systems are used for object detection and tracking, then the system cost is reduced, but the reliability of detection and tracking deteriorates due to object variation and pose variation

Engineering Contradiction:
Improvesystem costVSAvoiddetection reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transforms 2D image data into a pseudo-3D representation by generating depth maps and 3D point clouds from conventional 2D images using machine learning models. This allows the system to achieve 3D-like detection and tracking capabilities while using inexpensive 2D cameras, resolving the contradiction between cost and reliability.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces intermediate representations including depth maps, normal maps, and 3D point clouds as mediators between the 2D image input and the final detection/output. These intermediate representations encode spatial and geometric information that improves detection reliability while maintaining compatibility with 2D imaging systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If machine learning models are trained with large manually labeled datasets to improve reliability, then detection accuracy improves, but the complexity and time required for training increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidtraining complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-supervised learning where the system generates its own training labels by exploiting geometric constraints and temporal consistency in the video data. The system automatically creates supervision signals from the data itself without requiring manual annotation, reducing training complexity while maintaining high detection accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses feedback from the environment where the system operates to continuously improve the machine learning models. Real-world operational data is used to retrain and update models, allowing the system to adapt to changes in objects, poses, and environmental conditions without extensive manual relabeling.

Inventive Principle:
Principle #23Feedback

3Reliability

If depth-sensing cameras are used to enable object detection and tracking, then detection reliability improves, but the system cost increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidsystem cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent creates a synthetic copy of depth information and 3D geometric data from 2D images using trained machine learning models. Instead of using expensive depth-sensing hardware, the system generates equivalent depth maps and 3D point clouds computationally, achieving the same detection reliability at lower cost.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical/optical depth-sensing system with a computational approach using machine learning models that infer depth and 3D structure from 2D images. This substitution eliminates the need for expensive depth cameras while maintaining detection and tracking reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If the imaging system operates in environments with varying object structures and poses, then the system versatility improves, but the measurement precision deteriorates

Engineering Contradiction:
Improvesystem versatilityVSAvoidpoint correspondence precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adaptation of the machine learning models to handle varying object structures and poses. The system uses temporal consistency across video frames and continuous learning to maintain precise point correspondence measurements despite changes in object appearance, pose, and structure throughout the operational environment.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240296647A9Universal visual correspondence imaging system and method
Publication Date: 2024.09.05 MANEVA INC
  • US20240296647A9 patent drawing
  • US20240296647A9 patent drawing
  • US20240296647A9 patent drawing

AI summary

System and method for detecting a target object within an environment, including obtaining a two-dimensional input image of a scene within the environment; generating, using a machine learning based feature generation model, a feature map of respective feature vectors for the input image; comparing the feature vectors included in the feature map with reference feature vectors generated by the feature generation model based on reference points within a reference image, wherein the reference image includes an reference object instance that corresponds to the target object; based on the comparing, identifying points of interest in the input image that correspond to the reference points; and determining a presence of the target object in the environment based on the comparing.