Automated Image Dataset Curation via Feature Vector Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image recognition systems face inefficiencies due to large, noisy datasets that require excessive computing resources and human curation, which is impractical for managing image coherence and reducing memory usage.
Innovation Solution
An automated and unsupervised system generates feature vectors for pixel regions, clusters them, computes representative indices, and modifies scores to create a unified dataset with a confidence index, optimizing dataset organization and resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human curation is used to verify each image, then reliability of the dataset is improved, but productivity deteriorates as datasets expand
Solution Approach 1:
The system uses automated algorithms to curate and verify images without human intervention. The algorithm processes images, extracts features, compares them against the target individual, and automatically determines relevance, allowing the system to serve itself rather than requiring human curators
Solution Approach 2:
The patent replaces the mechanical human verification process with an automated computational system that uses feature extraction, clustering, and comparison algorithms to verify images, substituting human labor with machine-based processing
2Quantity of substance
If all images are stored in the dataset, then quantity of data is improved, but loss of information increases due to redundant and noisy data
Solution Approach 1:
The system extracts only the relevant features from images (such as facial features) and separates them from irrelevant information. It then extracts only the necessary images that contain the target individual, removing redundant and noisy data while preserving essential information
Solution Approach 2:
The patent applies different quality standards to different parts of the dataset. Relevant images containing the target individual are maintained with high quality and detail, while irrelevant or redundant images are removed or marked as low priority, creating a heterogeneous quality structure optimized for recognition tasks
3Speed
If images are organized for easy access, then speed of retrieval is improved, but device complexity increases due to organization requirements
Solution Approach 1:
The system performs preliminary organization of images during the dataset creation phase. Images are pre-processed, features are extracted and stored, and images are organized by relevance to the target individual before any retrieval operation occurs, so that when retrieval is needed, the work is already done
Solution Approach 2:
The patent divides the dataset into distinct segments or clusters based on image relevance and features. Images are segmented into groups (such as relevant vs. irrelevant, or different feature clusters), allowing for efficient retrieval by directly accessing the appropriate segment without searching the entire dataset
Data Source
AI summary
Techniques for dataset processing are provided. A plurality of feature vectors is generated for a plurality of pixel regions from a plurality of images, and a plurality of clusters is generated based on the plurality of feature vectors. A score is assigned to each respective pixel region in the plurality of pixel regions based at least in part on a cluster the respective pixel region is associated with. A unified dataset is created by, for each respective cluster in the plurality of clusters, computing a representative index for each pixel region in the respective cluster by comparing the pixel regions in the cluster, and modifying the score of the pixel regions in the cluster based on the computed representative indices. A confidence index is generated for the unified dataset based at least in part on a level of fragmentation in the unified dataset.


