Temporal Deformable Kernels for Accurate Video Object Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learning models for object identification and tracking in image data, such as those used in autonomous vehicles, suffer from inaccuracies and inefficiencies due to their inability to effectively utilize temporal information and deformable kernels, leading to flawed object detection and segmentation.

Innovation Solution

Implementing temporal-based deformable convolutions in neural networks that learn and apply offset fields to regular kernels based on image content from past and current images, allowing for the generation of deformed kernels for improved object detection, classification, and segmentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine-learning models are used for object detection, then the system is simple to implement, but the detection accuracy and precision are flawed

Engineering Contradiction:
Improveobject detection accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies deformable convolutions that dynamically adjust kernel shapes and positions based on learned offset fields, allowing the model to adapt to object deformations and movements in video sequences. This dynamic adaptation improves detection accuracy by matching the actual object geometry rather than using fixed kernels

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces temporal dimension by incorporating video sequences and using temporal convolutions to process information across multiple time frames. This additional temporal dimension enables the model to track object motion and improve detection precision through temporal context

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If temporal information is not utilized, then the computational process is simple, but object tracking and trajectory planning are inaccurate

Engineering Contradiction:
Improveobject tracking accuracyVSAvoidtemporal processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing video frames and extracting temporal features before main detection. The model learns offset fields and deformation patterns from temporal sequences in advance, preparing the data for more accurate object tracking and reducing complexity during real-time detection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediate representations such as offset fields and deformation vectors that mediate between raw temporal video data and final detection results. These intermediaries capture temporal motion patterns without requiring complex direct processing of entire video sequences

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If regular kernels are used for convolution, then the computational process is efficient, but the ability to handle object deformation and motion is limited

Engineering Contradiction:
Improvedeformation handling capabilityVSAvoidcomputational speed
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent applies local quality by using deformable convolutions that adjust kernel properties locally based on learned offset fields. Each kernel can have different shapes and positions adapted to local object characteristics, improving deformation handling while maintaining computational efficiency through selective adaptation rather than global complexity

Inventive Principle:
Principle #3Local quality

4Measurement precision

If accurate object segmentation is achieved through complex models, then detection precision improves, but computational time increases

Engineering Contradiction:
Improvesegmentation precisionVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by focusing computational resources on relevant regions. Deformable convolutions with learned offset fields concentrate processing on areas with object deformations or motions, rather than uniformly processing entire images. This partial focus achieves accurate segmentation while reducing overall computational time

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12118445B1Temporal-based deformable kernels
Publication Date: 2024.10.15 ZOOX INC
  • US12118445B1 patent drawing
  • US12118445B1 patent drawing
  • US12118445B1 patent drawing

AI summary

Techniques are disclosed for implementing a convolutional neural network that determines an offset field for deforming a kernel to be used in a convolution. The offset field is temporally-based, at least in part, on data generated at an earlier time. Furthermore, techniques are disclosed for using sensor data to train a neural network to learn shapes or configurations of such deformed kernels. The temporal-based deformable convolutions may be used for object identification, object matching, object classification, segmentation, and/or object tracking, in various examples.