Sports Video Action Classification Using Play State Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing action recognition methods in video analysis face challenges in accurately classifying actions in sports videos due to subtle visual differences and limitations in high-quality human and object tracking, especially under environmental factors like camera location and occlusions.
Innovation Solution
The method involves classifying actions in sports videos by utilizing play state information, player roles, and play-object position, specifically by receiving and analyzing video segments, determining play-object locations, and classifying actions based on these factors to improve recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If visual features are used for action recognition, then the method can be applied to various video types, but the accuracy is reduced due to subtle visual differences in action performance
Solution Approach 1:
The patent transitions from analyzing actions within a single video segment to utilizing temporal context across multiple segments. By examining play states and player roles in subsequent segments, the system adds a temporal dimension to action recognition, enabling more accurate classification despite subtle visual differences in the primary segment.
2Measurement precision
If tracking-based features are used to provide additional information about human activities, then action recognition can be improved, but the reliability is reduced due to limitations in high-quality tracking
Solution Approach 1:
The patent performs action classification by examining future play states and player roles before finalizing the action recognition. This preliminary examination of subsequent segments provides contextual information that compensates for tracking limitations, allowing the system to make more reliable action classifications even when tracking quality is compromised by environmental factors.
3Device complexity
If only visual features are used for action recognition, then the system complexity is low, but the measurement precision of action classification is insufficient
Solution Approach 1:
The patent combines visual features from the primary video segment with contextual features from subsequent segments, including play states and player roles. This merging of multiple feature types creates a more comprehensive classification system that achieves higher accuracy without requiring complex additional hardware, leveraging existing video data more effectively.
Data Source
AI summary
An apparatus for classifying an action in a video of a sports event including receiving a first video segment and a second video segment in the sports video, wherein the second video segment is the segment after next of the first video segment; receiving a predetermined play state of the first video segment and a different predetermined play state of the second video segment, locating a play object in one of a predetermined set of regions in a field of the sports event in the second segment; determining a role of a player interacting with the play object when the play object is located one of the predetermined set of regions in the second video segment; and classifying the action in the first video segment based on the located play object and the determined role.


