Pre-labeled Stock Image Repository for Machine Vision Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of machine vision systems is hindered by the computational and labor-intensive process of creating training data sets for machine learning algorithms, which requires manual labeling or computationally intensive automated labeling, leading to high costs and inefficiencies.
Innovation Solution
The use of pre-existing libraries of stock images, which are pre-labeled, to create a label-rich training dataset by combining these labels with item taxonomy information, reducing the need for de novo data set generation and minimizing computational loads.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual labeling or computationally intensive automated labeling algorithms are used to create training data sets, then the machine learning algorithms can be trained to operate machine vision systems, but the processing time and computational resources required increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-labeling stock images before they are needed for training. Image repositories contain pre-labeled images with metadata (categories, attributes, bounding boxes) that were labeled in advance using automated algorithms. When training data is needed, these pre-labeled images are immediately available without requiring real-time labeling, thus reducing processing time while maintaining data quality.
Solution Approach 2:
The patent uses copying by creating synthetic training images through digital manipulation. Instead of manually labeling each image, the system generates copies of existing stock images with various transformations (cropping, resizing, color adjustments, overlaying multiple images) to create diverse training datasets. This automated copying process eliminates manual labeling while preserving the quality of training data.
2Reliability
If manual labeling or computationally intensive automated labeling algorithms are used to create training data sets, then the machine learning algorithms can be trained to operate machine vision systems, but the computational resources and costs increase
Solution Approach 1:
The system performs labeling computations in advance when creating and maintaining the image repository, rather than when training is needed. Stock images are pre-labeled and stored with metadata, so the computational burden is distributed over time and amortized. This reduces the computational resources required at any single moment while maintaining high training data quality.
Solution Approach 2:
The patent reduces computational resources by using automated image copying and transformation to generate training data. Instead of computationally intensive manual labeling or complex automated labeling algorithms, the system creates synthetic training images through digital copying, cropping, resizing, and combining operations, which are computationally efficient while producing sufficient training data quality.
3Adaptability or versatility
If de novo data set generation is performed, then the training data can be customized for specific applications, but the process becomes labor-intensive and expensive
Solution Approach 1:
The patent applies universality by creating a multi-functional image repository that serves both as a stock image library and as a training data source. The same pre-labeled stock images are used for multiple purposes: commercial licensing, design inspiration, and machine learning training. This eliminates the need to create separate customized datasets for different applications, reducing labor and costs while maintaining adaptability through the repository's diverse, pre-categorized content.
Solution Approach 2:
The system achieves customization efficiently through automated copying and transformation of stock images. Rather than manually creating customized datasets, the system generates application-specific training data by copying, cropping, filtering, and transforming pre-labeled stock images using automated algorithms. This maintains the ability to customize datasets for specific applications while eliminating manual labor and reducing costs.
Data Source
AI summary
Systems and methods including one or more processors and one or more non-transitory storage devices storing computing instructions configured to run on the one or more processors and perform acts of receiving one or more digital images from a repository of digital images; annotating the one or more digital images from the repository of digital images; digitally altering the one or more digital images, as annotated, from the repository of digital images; digitally combining the one or more digital images, as annotated and digitally altered, with at least one or more portions of one or more other digital images of the repository of digital images to create one or more combined digital images; training a machine learning algorithm on the one or more combined digital images; and storing the machine learning algorithm, as trained, in the one or more non-transitory computer readable storage devices.


