Action Recognition Using Poselet Keyframes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video-based action recognition models are computationally expensive and sensitive to changes in action duration and dropped frames, as they rely on computing features over long temporal trajectories.
Innovation Solution
Modeling human actions as a sequence of discriminative keyframes, where poselets are selected and encoded to represent key states of actions, with correlations between poselets used to determine action occurrence, reducing computational resources and sensitivity to duration changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If long temporal trajectories with many frames are used for action recognition, then descriptive accuracy is improved, but computational cost increases
Solution Approach 1:
The patent extracts only the most informative frames (keyframes) from the video sequence based on poselet detection confidence scores, rather than processing all frames. This selective extraction maintains action recognition accuracy while significantly reducing computational cost by focusing only on discriminative moments in the action sequence.
Solution Approach 2:
The patent processes only a subset of frames (keyframes) rather than the complete frame sequence. By identifying and processing only those frames that contain critical action information, the system achieves adequate action recognition with reduced computational effort, applying partial action processing to solve the contradiction.
2Loss of information
If long temporal trajectories with many frames are used for action recognition, then action description quality is improved, but sensitivity to action duration changes increases
Solution Approach 1:
The patent dynamically adapts the number of keyframes and their temporal spacing based on the detected action characteristics. The system adjusts the frame selection process to account for varying action durations, making the recognition model more robust to temporal variations while maintaining descriptive quality through adaptive keyframe sampling.
Solution Approach 2:
The patent changes the temporal parameters of frame sampling based on detected action patterns. By adjusting the time intervals and number of keyframes according to the specific action being recognized, the system maintains high description quality while becoming less sensitive to duration changes through parameter adaptation.
3Loss of information
If long temporal trajectories with many frames are used for action recognition, then action modeling completeness is improved, but sensitivity to dropped frames increases
Solution Approach 1:
The patent performs preliminary poselet detection and confidence scoring on all frames before selecting keyframes for detailed processing. This preliminary action identifies which frames contain critical action information, allowing the system to maintain modeling completeness while being robust to dropped frames by focusing computational resources on the most informative moments.
Solution Approach 2:
The patent extracts only the most critical frames for detailed action modeling, rather than requiring complete processing of all frames. This extraction approach maintains action modeling completeness by focusing on discriminative keyframes while naturally providing robustness to dropped frames, as the system does not depend on every intermediate frame being present.
Data Source
AI summary
Methods and systems for video action recognition using poselet keyframes are disclosed. An action recognition model may be implemented to spatially and temporally model discriminative action components as a set of discriminative keyframes. One method of action recognition may include the operations of selecting a plurality of poselets that are components of an action, encoding each of a plurality of video frames as a summary of the detection confidence of each of the plurality of poselets for the video frame, and encoding correlations between poselets in the encoded video frames.


