Object Tracking via Affine-Invariant Feature Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image processing technologies face challenges in tracking movable objects using temporally-spaced images due to limited data availability, inability to provide reliable and invariant object features, and difficulties in managing differences in image capture platforms and resolutions, as well as accurately segmenting objects from backgrounds and distinguishing them from similar objects.

Innovation Solution

A method and apparatus for tracking objects using temporally-spaced images by aligning and comparing features between images, assigning descriptors to detected features, and correlating them to determine matches, enabling image recognition and reacquisition of objects after temporary losses in tracking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If standard frame-to-frame data association mechanisms are used for visual object recognition, then object recognition can be performed in contiguous images, but these mechanisms are not usable when multiple images are separated by time intervals

Engineering Contradiction:
Improveobject recognition reliabilityVSAvoidapplicability to temporally-spaced images
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the object recognition approach by changing parameters from frame-to-frame temporal continuity to feature-based spatial and temporal invariance. It uses affine transformation parameters to model pose changes and appearance changes over time intervals, enabling recognition across temporally-spaced images by adjusting the temporal parameter from contiguous to spaced intervals

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary actions by pre-establishing invariant feature representations and transformation models during a learning phase before actual tracking. It pre-computes affine transformation parameters and feature descriptors that remain invariant under various transformations, so that when temporally-spaced images are encountered, the system can directly apply these pre-established models without requiring contiguous frame data

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If limited data from the first learning image is used for training, then the learning sequence can be established, but reliable and invariant object features cannot be provided to overcome drastic pose changes, aspect changes, appearance changes and occlusions

Engineering Contradiction:
Improvelearning sequence establishmentVSAvoidfeature representation reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent addresses limited training data by changing the parameter of feature representation from raw pixel data to affine-invariant feature descriptors. These descriptors are designed to be invariant under affine transformations including scale, rotation, shear, and translation, allowing the system to maintain reliability even with limited training data by focusing on transformation-invariant characteristics

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the object recognition problem into multiple invariant feature components that can be independently extracted and transformed. By dividing the object into key feature points and describing their relationships through affine-invariant metrics, the system can build reliable representations from limited data by focusing on structural relationships rather than complete object appearance

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If features are extracted to represent object characteristics, then object identification is possible, but features are not invariant under drastic pose changes, aspect changes, appearance changes and occlusions

Engineering Contradiction:
Improveobject feature detectionVSAvoidfeature invariance under transformation
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

The patent transforms feature parameters from standard image coordinates to affine-invariant coordinate systems. It applies normalization techniques that remove the effects of affine transformations, converting pose-dependent features into pose-independent descriptors that maintain stability under drastic transformations while preserving measurement precision through invariant metric calculations

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If images captured by differing platforms and resolutions are processed, then multi-platform tracking is enabled, but differences in platform and resolution create challenges for accurate object recognition

Engineering Contradiction:
Improvemulti-platform compatibilityVSAvoidobject recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent resolves multi-platform resolution differences by changing the parameter scale through affine transformation normalization. It establishes a common reference frame that accounts for different resolutions by applying scale-invariant transformations, allowing accurate object recognition across platforms with differing resolutions by treating resolution as a transformable parameter rather than a fixed constraint

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7929728B2Method and apparatus for tracking a movable object
Publication Date: 2011.04.19 SRI INTERNATIONAL
  • US7929728B2 patent drawing
  • US7929728B2 patent drawing
  • US7929728B2 patent drawing

AI summary

A method and apparatus for tracking a movable object using a plurality of images, each of which is separated by an interval of time is disclosed. The plurality of images includes first and second images. The method and apparatus include elements for aligning the first and second images as a function of (i) at least one feature of a first movable object captured in the first image, and (ii) at least one feature of a second movable object captured in the second image; and after aligning the first and second images, comparing at least one portion of the first image with at least one portion of the second image.