Action Classification via Manipulated Object Movement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning algorithms for video classification require large amounts of labeled training data and struggle to distinguish between target objects and background elements, leading to inefficient action classification, especially when fine-scale actions and details are involved.
Innovation Solution
A computing device processor is configured to identify target regions and surrounding regions within video frames, extract features, generate manipulated object identifiers, and classify actions based on object movements, using techniques like hand detectors and grasp classifiers to reduce the need for extensive training data and improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If machine learning algorithms are used for video classification, then action classification can be automated, but large amounts of labeled training data are required
Solution Approach 1:
The video frame is segmented into a target region containing the target object and a surrounding region. This segmentation allows the system to focus computational resources and attention only on relevant areas, reducing the need for extensive training data while maintaining accurate action classification.
Solution Approach 2:
The patent extracts features specifically from the surrounding region of the target object, separating this information from the entire frame. This extraction approach concentrates relevant visual information, enabling effective action classification with reduced training data requirements.
2Productivity
If machine learning algorithms analyze entire video frames, then action classification can be performed, but the system struggles to distinguish between target objects and background elements
Solution Approach 1:
The video frame is divided into a target region containing the target object and a surrounding region. This segmentation enables the system to focus on relevant areas, improving the distinction between target objects and background elements while maintaining classification efficiency.
Solution Approach 2:
The patent applies different processing qualities to different regions: the target region receives focused analysis for object identification, while the surrounding region is processed to extract contextual features. This local quality approach enhances object-background distinction accuracy without compromising overall productivity.
3Measurement precision
If human operators view and analyze video data, then accurate understanding can be achieved, but the process is costly and time-consuming
Solution Approach 1:
The system performs self-service by automatically detecting target objects, extracting features from surrounding regions, generating manipulated object identifiers, and classifying actions without human intervention. This automation maintains accurate understanding while eliminating the time and cost associated with human analysis.
Solution Approach 2:
The patent replaces the mechanical process of human visual analysis with an automated computational system that processes video frames, extracts features, and classifies actions. This substitution maintains measurement precision while significantly reducing time loss.
Data Source
AI summary
A computing device, including a processor configured to receive a first video including a plurality of frames. For each frame, the processor may determine that a target region of the frame includes a target object. The processor may determine a surrounding region within which the target region is located. The surrounding region may be smaller than the frame. The processor may identify one or more features located in the surrounding region. From the one or more features, the processor may generate one or more manipulated object identifiers. For each of a plurality of pairs of frames, the processor may determine a respective manipulated object movement between a first manipulated object identifier of the first frame and a second manipulated object identifier of the second frame. The processor may classify at least one action performed in the first video based on the plurality of manipulated object movements.


