Self-Supervised Image Curation Using Visual Templates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image curation methods for applications like Autonomous Driving and CCTV surveillance are manual, time-consuming, and costly, especially when dealing with diverse scenes containing various objects, making it challenging to automate the data curation process effectively.
Innovation Solution
A system utilizing a deep neural network architecture with a pre-trained DNN and feature database to iteratively refine and curate images based on pre-defined visual templates, employing feature extraction, similarity search, patch generation, and self-supervised learning to filter and refine image patches without manual labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual image curation is used, then labeling accuracy can be ensured, but time consumption and cost increase significantly
Solution Approach 1:
The system enables self-service through automated image curation using unsupervised learning algorithms. The model automatically identifies and labels objects of interest in images without requiring manual human intervention for each image, thereby maintaining labeling accuracy while dramatically reducing time consumption and costs associated with manual curation processes
Solution Approach 2:
The patent replaces the mechanical manual labeling process with an automated computational system. Instead of human operators manually examining and labeling images, the system uses machine learning models to automatically detect, segment, and label objects of interest, substituting human mechanical effort with automated algorithmic processing
2Productivity
If automated image curation is implemented, then time and cost are reduced, but handling diverse scenes with various objects becomes challenging
Solution Approach 1:
The system achieves universality by designing a multi-functional automated curation pipeline that can handle diverse scene types and object categories through a single unified framework. The unsupervised learning model is capable of adapting to different domains (e.g., autonomous driving, surveillance) and various object types without requiring domain-specific manual configuration, thereby maintaining high productivity across diverse应用场景
Solution Approach 2:
The system employs parameter changes by adjusting model hyperparameters and processing thresholds dynamically based on the characteristics of different scene types. This allows the automated curation system to adapt its behavior to suit different domains and object types, improving versatility while maintaining efficient automated processing
3Reliability
If labeled data is used for training, then supervised learning performance improves, but the cost and time for data preparation increase
Solution Approach 1:
The system performs preliminary action by automatically generating labeled training data through unsupervised learning before the actual model training phase. This pre-labeling process prepares high-quality training data without manual intervention, enabling subsequent supervised learning to achieve high performance while avoiding the time-consuming manual data preparation process
Solution Approach 2:
The system creates copies of unlabeled images with automatically generated labels that can be used for training. These synthesized labeled copies serve as training data, allowing the model to learn from diverse scenarios without requiring manual creation of labeled datasets, thereby reducing data preparation time while maintaining training effectiveness
Data Source
AI summary
A method and system for curating images containing specific objects specified by visual templates are described. A set of visual image templates representing an input object is provided to a pre-trained deep neural network (DNN). A feature extraction module is configured to extract a set of feature vectors representing a set of visual image templates and a list of images stored in a DNN-based feature database. A patch generation module is configured to generate image patches representing the input object, from a set of relevant neighbor images providing probable regions where the visual image templates are getting matched. The patch generation module is further configured to train a new network model in a self-supervised manner, with image patches generation by the patch generation module, in an iterative manner.

