Video Object Tracking with Kinetic Model Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing tracking methods, such as those using Siamese region proposal networks, are vulnerable to loss of sight or transfer, making long-term tracking of objects in videos challenging due to deformation, shielding, or passing of the tracking target.
Innovation Solution
An image processing apparatus and method that utilizes image feature calculation, object identification, kinetic model selection, and object tracking means to robustly track objects in time series by integrating image features and tracking results, selecting appropriate kinetic models, and predicting object positions using stored models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If tracking methods use only image features from each frame (e.g., Siamese region proposal network), then tracking speed is fast, but tracking accuracy deteriorates due to vulnerability to loss of sight or transfer
Solution Approach 1:
The patent applies preliminary action by pre-extracting and storing kinetic models of objects before tracking occurs. These kinetic models capture temporal patterns and motion characteristics that enable the system to predict object positions and maintain tracking accuracy even when the object is temporarily occluded or transfers between regions, thus resolving the contradiction between fast tracking and accurate long-term tracking.
Solution Approach 2:
The patent implements continuity of useful action by maintaining a kinetic model that continuously updates and stores temporal information about object motion. This continuous model allows the tracking system to bridge gaps caused by occlusion or transfer, ensuring uninterrupted and accurate tracking throughout the video sequence without relying solely on instantaneous frame features.
2Device complexity
If tracking methods use only image features from each frame, then device complexity is low, but reliability deteriorates due to vulnerability to deformation, shielding, or passing
Solution Approach 1:
The system performs preliminary extraction of kinetic models that encode temporal motion patterns and object characteristics before actual tracking occurs. This pre-computed kinetic information provides robust cues for tracking through occlusion, deformation, or transfer events, significantly improving reliability without requiring complex real-time processing during tracking.
Solution Approach 2:
The kinetic model acts as an intermediary between the input video frames and the tracking output. It mediates the tracking process by providing temporal context and motion predictions that bridge the gap between consecutive frames, enabling reliable tracking even when direct visual observation is interrupted by occlusion or transfer.
3Measurement precision
If tracking methods use abundant image features from deep learning, then detection accuracy improves, but vulnerability to loss of sight or transfer increases
Solution Approach 1:
The patent merges spatial image features extracted from deep learning with temporal kinetic models. This combination allows the system to leverage the high detection accuracy of deep learning while the temporal dimension provided by kinetic models compensates for vulnerability to occlusion and transfer, achieving both high accuracy and robustness simultaneously.
Solution Approach 2:
The system performs preliminary extraction of temporal kinetic patterns before tracking occurs. These pre-computed temporal models provide robustness against loss of sight or transfer by establishing expected motion trajectories in advance, allowing the system to maintain reliable tracking even when spatial features become unreliable due to occlusion.
Data Source
AI summary
According to the present disclosure, it is an object to provide, in particular, an image processing apparatus, an image processing system, an image processing method, and a non-transitory computer-readable medium storing an image processing program therein, capable of tracking an object in a video with high accuracy. An image processing apparatus according to the present disclosure includes: image feature calculation means for outputting an image feature of an object using a time-series image of the object; object identification means for outputting an identification result obtained by identifying an object using the image feature and a tracking result of the object; kinetic model selection means for selecting an appropriate kinetic model from a plurality of kinetic models based on the identification result and the tracking result; and object tracking means for tracking the object in time series and calculating the tracking result, from the identified object.


