Object Dataset Creation Using Labeled Action-Object Videos
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning models face performance issues when applied to data with different feature distributions than their training data, and existing object datasets lack sufficient examples of objects in various states, hindering action recognition and compliance verification.
Innovation Solution
A method for creating or modifying object datasets using labeled action-object videos, where a subset of frames with bounding boxes is pruned to identify sufficiently distinct objects, and their information addition scores are assessed to determine inclusion in the dataset, ensuring diverse training examples and separate state categories for objects like 'open' vs. 'closed' boxes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing object datasets are used for training deep learning models, then the training process can be completed, but the models face performance issues when applied to data with different feature distributions
Solution Approach 1:
The patent changes the parameters of training data by creating multiple state categories for objects (e.g., open/closed boxes, on/off switches). This transforms the training dataset to include diverse feature distributions across different object states, enabling models to adapt to varying input characteristics while maintaining reliable performance
2Ease of manufacture
If existing object datasets are used, then dataset creation is simple, but they lack sufficient examples of objects in various states, hindering action recognition
Solution Approach 1:
The patent segments object representations by creating distinct state categories (e.g., open box, closed box, on switch, off switch). This segmentation divides the object data into multiple meaningful groups, each representing a different state, thereby increasing the quantity and diversity of training examples without complicating the overall dataset creation process
Solution Approach 2:
The patent performs preliminary action by automatically generating state category labels and organizing objects into different states before training. This pre-processing step creates a structured dataset with diverse object state examples, enabling action recognition models to learn from pre-organized state information rather than raw unprocessed data
3Ease of operation
If existing object datasets are used, then training data preparation is straightforward, but compliance verification is hindered due to lack of state diversity
Solution Approach 1:
The patent performs preliminary action by pre-organizing training data into state categories (e.g., open/closed, on/off) before compliance verification. This pre-structuring enables the verification system to easily compare predicted states against known state categories, improving compliance verification accuracy while maintaining straightforward data preparation procedures
Solution Approach 2:
The patent changes the organizational parameters of the dataset by introducing state categories as a new dimension for data structuring. This parameter change transforms the dataset from simple object collections to state-aware organized structures, enabling reliable compliance verification through state comparison while keeping data preparation systematic and manageable
Data Source
AI summary
An object dataset creation or modification mechanism is provided for object dataset creation or modification using a labeled action-object video. For a plurality of frames of the labeled action-object video, an identification is made of a subset of frames where a bounding box object (BBO) exists. BBOs in the subset of frames where a BBO exists are pruned to identify sufficiently distinct BBOs thereby forming a set of pruned BBOs. For each pruned BBO in the set of pruned BBOs: an information addition score is determined; the information addition score is assessed; responsive to the information addition score being positively assessed, the pruned BBO is added to an object dataset; and, responsive to the information addition score being negatively assessed, the pruned BBO is discarded.


