Point Tracking Using Trained Neural Networks for Occlusion Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Fiducial elements in imaging systems can be distracting, difficult to track due to occlusion, and limit the ability to capture information about points in space, especially in live performances or dynamic environments where they may be temporarily occluded or not visible.

Innovation Solution

A trained neural network, specifically a directed graph function approximator, is used to track points in images by generating a three-dimensional model of the locale or object, allowing the network to identify and locate tracked points even when fiducial elements are occluded or absent, using supervised or unsupervised learning with synthesized training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If fiducial elements are used as reference points for tracking, then point tracking can be achieved, but the fiducial elements become distracting and visible in the final output

Engineering Contradiction:
Improvepoint tracking accuracyVSAvoidvisual distraction
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent creates virtual fiducial elements through synthesized training data that replicate the appearance and geometric properties of physical fiducials. These virtual elements are embedded in 3D models and rendered into images, allowing the neural network to learn fiducial detection without requiring actual physical tags to be visible in the final output. The network learns to detect and track points based on the synthesized fiducial patterns during training, then applies this knowledge to real scenes without the distracting physical markers.

Inventive Principle:
Principle #26Copying

2Measurement precision

If physical fiducial elements are attached to structures, then point tracking reference points are established, but the ability to track floating points in space is limited

Engineering Contradiction:
Improvereference point detectionVSAvoidtracking point flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transitions from 2D image processing to 3D spatial reasoning by creating three-dimensional models of locales and rendering fiducial elements within these 3D spaces. The neural network is trained on synthesized images generated from 3D models, enabling it to understand and track points in three-dimensional space rather than just on two-dimensional surfaces. This allows tracking of floating points in space by establishing geometric relationships and spatial constraints within the 3D model framework.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If fiducial elements are used for tracking, then anchor points can be identified, but tracking fails when fiducials are temporarily occluded

Engineering Contradiction:
Improvetracking stabilityVSAvoidpoint location accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by pre-training the neural network on extensive synthesized training data that includes numerous scenarios with occlusions, different lighting conditions, and various camera angles. The 3D models are used to generate training images where fiducial elements appear in diverse configurations and occlusion states. This preliminary training enables the network to learn robust feature extraction and tracking algorithms that maintain accuracy even when fiducials are temporarily occluded during actual operation, as the network has already encountered similar scenarios during training.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If multiple fiducials are distributed throughout a locale, then tracking coverage is improved, but the complexity of the system increases

Engineering Contradiction:
Improvetracking coverageVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a universal training framework using 3D models and synthesized images that can represent any locale or environment. Instead of developing separate tracking systems for different locales, the neural network is trained on a diverse set of synthesized images generated from 3D models of various spaces, objects, and configurations. This universal training approach enables the single trained network to handle multiple fiducials distributed throughout any locale, reducing system complexity while maintaining comprehensive tracking coverage across diverse environments.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11080884B2Point tracking using a trained network
Publication Date: 2021.08.03 COSTAR REALTY INFORMATION INC
  • US11080884B2 patent drawing
  • US11080884B2 patent drawing
  • US11080884B2 patent drawing

AI summary

A trained network for point tracking includes an input layer configured to receive an encoding of an image. The image is of a locale or object on which the network has been trained. The network also includes a set of internal weights which encode information associated with the locale or object, and a tracked point therein or thereon. The network also includes an output layer configured to provide an output based on the image as received at the input layer and the set of internal weights. The output layer includes a point tracking node that tracks the tracked point in the image. The point tracking node can track the point by generating coordinates for the tracked point in an input image of the locale or object. Methods of specifying and training the network using a three-dimensional model of the locale or object are also disclosed.