Video Tracking via Key Point Posture Geometry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing person tracking technologies fail to accurately track multiple individuals across video frames, especially when conditions such as congestion, view angle, distance, and frame rate differ from learned data, and require reference images for each posture stored in a database.

Innovation Solution

A tracking apparatus that detects targets in video frames, extracts key points, generates posture information, and tracks targets based on position and orientation across frames, allowing continuous tracking without the need for reference images or specific learning conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning-based posture tracking is used, then tracking accuracy under learned conditions is improved, but tracking reliability under varying conditions (congestion, angle, distance, frame rate) deteriorates

Engineering Contradiction:
Improveposture tracking accuracyVSAvoidtracking continuity under varying conditions
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the tracking parameters from deep learning model predictions to geometric relationships between key points. By representing posture as relative positions and orientations of key points rather than relying on learned patterns, the system adapts to varying conditions (congestion, angle, distance, frame rate) without requiring retraining or reference images.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If reference images for each posture are stored in a database, then person identification accuracy is improved, but device complexity and storage requirements increase

Engineering Contradiction:
Improveperson identification accuracyVSAvoiddatabase storage requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential geometric features (key point positions and orientations) needed for tracking, eliminating the need to store complete reference images. By taking out only the necessary positional and orientational data rather than storing full image references, the system reduces storage complexity while maintaining tracking capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of storing actual reference images, the patent creates simplified geometric copies represented by key point coordinates and relative orientations. These geometric representations serve as lightweight substitutes that capture the essential posture information without the complexity of storing full image data.

Inventive Principle:
Principle #26Copying

3Measurement precision

If three-dimensional posture estimation from two-dimensional joint positions is performed, then posture information quality is improved, but ability to track multiple persons deteriorates

Engineering Contradiction:
Improveposture information qualityVSAvoidmulti-person tracking capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the tracking problem by independently tracking key points for each person rather than estimating complete three-dimensional postures. By dividing the multi-person tracking task into independent key point tracking segments, the system maintains posture information quality while scaling to multiple persons without the computational complexity of full 3D estimation for each individual.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230386049A1Tracking apparatus, tracking system, tracking method, and recording medium
Publication Date: 2023.11.30 NEC CORP
  • US20230386049A1 patent drawing
  • US20230386049A1 patent drawing
  • US20230386049A1 patent drawing

AI summary

A tracking apparatus that includes a detection unit that detects a tracked target from at least two frames constituting video data; an extraction unit that extracts at least one key point from the tracked target having been detected, a posture information generation unit that generates posture information of the tracked target based on the at least one key point, and a tracking unit that tracks the tracked target based on a position and an orientation of the posture information of the tracked target detected from each of the at least two frames.