Automated Image Dataset Curation via Feature Vector Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image recognition systems face inefficiencies due to large, noisy datasets that require excessive computing resources and human curation, which is impractical for managing image coherence and reducing memory usage.

Innovation Solution

An automated and unsupervised system generates feature vectors for pixel regions, clusters them, computes representative indices, and modifies scores to create a unified dataset with a confidence index, optimizing dataset organization and resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If human curation is used to verify each image, then reliability of the dataset is improved, but productivity deteriorates as datasets expand

Engineering Contradiction:
Improvedataset reliabilityVSAvoidcuration productivity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system uses automated algorithms to curate and verify images without human intervention. The algorithm processes images, extracts features, compares them against the target individual, and automatically determines relevance, allowing the system to serve itself rather than requiring human curators

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human verification process with an automated computational system that uses feature extraction, clustering, and comparison algorithms to verify images, substituting human labor with machine-based processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If all images are stored in the dataset, then quantity of data is improved, but loss of information increases due to redundant and noisy data

Engineering Contradiction:
Improvedata quantityVSAvoidinformation quality
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The system extracts only the relevant features from images (such as facial features) and separates them from irrelevant information. It then extracts only the necessary images that contain the target individual, removing redundant and noisy data while preserving essential information

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality standards to different parts of the dataset. Relevant images containing the target individual are maintained with high quality and detail, while irrelevant or redundant images are removed or marked as low priority, creating a heterogeneous quality structure optimized for recognition tasks

Inventive Principle:
Principle #3Local quality

3Speed

If images are organized for easy access, then speed of retrieval is improved, but device complexity increases due to organization requirements

Engineering Contradiction:
Improveretrieval speedVSAvoidorganization complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system performs preliminary organization of images during the dataset creation phase. Images are pre-processed, features are extracted and stored, and images are organized by relevance to the target individual before any retrieval operation occurs, so that when retrieval is needed, the work is already done

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides the dataset into distinct segments or clusters based on image relevance and features. Images are segmented into groups (such as relevant vs. irrelevant, or different feature clusters), allowing for efficient retrieval by directly accessing the appropriate segment without searching the entire dataset

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10943098B2Automated and unsupervised curation of image datasets
Publication Date: 2021.03.09 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10943098B2 patent drawing
  • US10943098B2 patent drawing
  • US10943098B2 patent drawing

AI summary

Techniques for dataset processing are provided. A plurality of feature vectors is generated for a plurality of pixel regions from a plurality of images, and a plurality of clusters is generated based on the plurality of feature vectors. A score is assigned to each respective pixel region in the plurality of pixel regions based at least in part on a cluster the respective pixel region is associated with. A unified dataset is created by, for each respective cluster in the plurality of clusters, computing a representative index for each pixel region in the respective cluster by comparing the pixel regions in the cluster, and modifying the score of the pixel regions in the cluster based on the computed representative indices. A confidence index is generated for the unified dataset based at least in part on a level of fragmentation in the unified dataset.