Multi-Modal Sensor Annotation for Computer Vision

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for annotating images or video files in computer vision applications are time-consuming and computationally expensive, requiring intense human effort and processing power, especially when dealing with large datasets from moving imaging devices like unmanned aerial vehicles.

Innovation Solution

The use of multi-modal sensor data, where calibrated sensors like digital cameras and thermographic cameras capture synchronized images, allowing attributes from one modality to enhance the detection and annotation accuracy in another modality, such as transposing thermal attributes onto visual data to improve object detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation methods are used, then annotation accuracy can be maintained, but annotation time and labor costs increase significantly

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The annotation process is segmented into multiple stages: initial automatic annotation using computer vision algorithms, followed by selective manual verification and refinement. This division allows the system to leverage automated speed while maintaining accuracy through targeted human intervention only where needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A semi-automatic annotation system acts as an intermediary between fully automatic and fully manual methods. The system uses computer vision algorithms to generate preliminary annotations, then presents them to human annotators for verification and correction, combining the speed of automation with the accuracy of manual review.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automatic annotation methods are used, then annotation speed increases, but computational resources and processing power requirements increase significantly

Engineering Contradiction:
Improveannotation speedVSAvoidcomputational resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

Instead of applying computationally intensive automatic annotation to every single image, the system applies automatic annotation selectively based on confidence thresholds and image characteristics. Low-confidence or complex images are routed to manual annotation, optimizing the balance between speed and resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically adjusts annotation processing parameters based on image characteristics, complexity, and confidence scores. This allows the computational resources to be allocated efficiently, applying heavy processing only where necessary and using lighter processing for straightforward cases.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If traditional annotation methods are used for large datasets, then complete coverage can be achieved, but the process becomes prohibitively expensive and time-consuming

Engineering Contradiction:
Improvedataset coverageVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system uses automatically generated annotations from computer vision algorithms to annotate the majority of the dataset without human intervention. This self-service approach handles routine annotation tasks autonomously, reserving human annotators for edge cases and quality assurance, thereby achieving complete dataset coverage at scale.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary automatic annotation on the entire dataset before manual review. This preliminary action pre-processes the data, creating a baseline annotation set that can be quickly refined later, rather than starting from scratch with manual annotation for each image.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10691943B1Annotating images based on multi-modal sensor data
Publication Date: 2020.06.23 AMAZON TECH INC
  • US10691943B1 patent drawing
  • US10691943B1 patent drawing
  • US10691943B1 patent drawing

AI summary

Imaging data or other data captured using a camera may be classified based on data captured using another sensor that is calibrated with the camera and operates in a different modality. Where a digital camera configured to capture visual images is calibrated with another sensor such as a thermal camera, a radiographic camera or an ultraviolet camera, and such sensors capture data simultaneously from a scene, the respectively captured data may be processed to detect one or more objects therein. A probability that data depicts one or more objects of interest may be enhanced based on data captured from calibrated sensors operating in different modalities. Where an object of interest is detected to a sufficient degree of confidence, annotated data from which the object was detected may be used to train one or more classifiers to recognize the object, or similar objects, or for any other purpose.