Automated Training Data Generation for Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenges in object recognition include the lack of valid training data, high cost, and privacy concerns associated with manually labeled data collection.
Innovation Solution
Automatically generating training data sets by leveraging search graphs and knowledge graphs through computer systems, utilizing techniques like named entity extraction and clustering to collect and filter images, and combining relevant data pairs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data collection and labeling is used, then data quality can be ensured through human review, but the cost and time consumption increase significantly
Solution Approach 1:
The system performs preliminary automated filtering and clustering of images before human review. By pre-processing the data to group similar images and remove obvious duplicates, the system reduces the volume of data requiring manual labeling while maintaining quality control, thus decreasing time consumption without sacrificing data quality
Solution Approach 2:
The patent introduces automated algorithms (clustering, similarity detection) as intermediary steps between raw data collection and final labeling. These intermediaries pre-process the data to identify groups of similar images, allowing human annotators to work on consolidated sets rather than individual images, thereby reducing overall time consumption while preserving quality through selective human review
2Measurement precision
If manual data collection is used, then data accuracy can be maintained, but the cost increases significantly
Solution Approach 1:
The system performs preliminary automated filtering and clustering to reduce the dataset size before human review. By pre-identifying and grouping similar images algorithmically, the system minimizes the portion of data requiring expensive manual processing, thereby maintaining accuracy through targeted human validation while significantly reducing overall cost
Solution Approach 2:
The patent uses automated image matching and clustering to create representative copies or proxies for groups of similar images. Instead of manually processing every unique image, the system identifies representative samples from clusters, reducing the number of manual operations needed while preserving data accuracy through the representative nature of the selected samples
3Productivity
If large amounts of data are collected automatically, then productivity increases, but data quality control becomes more difficult
Solution Approach 1:
The patent segments the large dataset into clusters of similar images using automated algorithms. By dividing the massive dataset into manageable groups based on similarity metrics, the system maintains productivity through automated processing while improving quality control by enabling focused human review on cluster representatives rather than individual images across the entire dataset
Solution Approach 2:
The system introduces automated clustering algorithms as intermediary processing steps between data collection and quality control. These intermediaries organize large volumes of automatically collected data into structured groups, maintaining high productivity while facilitating quality control by presenting organized, manageable units for human validation rather than unstructured raw data
4Measurement precision
If diverse images are collected for each object, then recognition accuracy improves, but the complexity of data management increases
Solution Approach 1:
The patent merges multiple images of the same object into clusters based on similarity analysis. By combining diverse images into organized groups representing each object, the system maintains recognition accuracy through exposure to varied examples while reducing data management complexity by consolidating related images into unified clusters that can be managed as single units
Data Source
AI summary
The present disclosure provides method and apparatus for automatically generating a training data set for object recognition. Profile information of a plurality of objects may be obtained. For each object among the plurality of objects, a group of initial images associated with the object may be collected based on identity information of the object included in profile information of the object. The group of initial images may be filtered to obtain a group of filtered images associated with the object. A group of training data pairs corresponding to the object may be generated through labeling each of the group of filtered images with the identity information of the object. The group of training data pairs may be added into the training data set.


