Crowdsourced Data Marketplace for ML Training Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning algorithms lack sufficient high-quality data sets for training, as existing data sets are often insufficient in size and quality, and the generation of synthetic data introduces biases and over-fitting.

Innovation Solution

A system and method for creating high-quality data sets through crowdsourced curation, utilizing a data marketplace that incentivizes contributors, with data automatically scored for quality and provenance, and human curation of low-scoring data, combined with synthetic data generation using generative adversarial networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic data is generated to increase data supply, then the quantity of training data is improved, but the quality deteriorates due to over-fitting and inherited quality problems

Engineering Contradiction:
Improvequantity of training dataVSAvoidquality of training data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the data curation process into multiple independent components: automated ingestion from disparate sources, reputation scoring subsystem, human steward verification queue, and synthetic data generation. Each component operates independently with specific quality checks, allowing the system to process large volumes of data while maintaining quality through distributed verification rather than centralized bottlenecks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a feedback loop where data stewards verify and correct synthetic data, which then feeds back into improving the reputation scoring algorithm. The system continuously learns from human verification outcomes, adjusting scoring metrics and thresholds to better identify high-quality data. This closed-loop feedback ensures that synthetic data generation improves over time while maintaining reliability.

Inventive Principle:
Principle #23Feedback

2Reliability

If manual curation is used to ensure data quality, then the reliability of data sets is improved, but the productivity deteriorates due to time-consuming manual processes

Engineering Contradiction:
Improvequality of training dataVSAvoiddata creation speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies local quality by implementing different verification strategies for different data sources and types. High-risk sources undergo rigorous human steward review, while trusted sources with established reputations receive automated processing. Data entries are evaluated individually based on their specific characteristics rather than applying uniform manual review to all data, optimizing both quality and throughput.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary automated filtering and reputation scoring before data reaches human stewards. The system pre-processes incoming data through multiple automated quality checks, source reputation evaluation, and initial validation rules, so that human curators only need to review data that requires expert judgment. This preliminary action significantly reduces the manual workload while maintaining high quality standards.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated ingestion is used to increase data processing speed, then the productivity is improved, but the measurement precision deteriorates due to difficulties in detecting and measuring data quality

Engineering Contradiction:
Improvedata processing speedVSAvoiddata quality assessment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a universal reputation scoring system that evaluates data from multiple disparate sources using a consistent multi-dimensional framework. The same scoring algorithm assesses structured data, unstructured text, images, and other data types, applying universal quality metrics across diverse formats. This multi-functional approach maintains measurement precision while handling varied data types at high speed through automated processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary reputation scoring layer between automated ingestion and final data acceptance. Rather than directly accepting or rejecting data based on single metrics, the system uses the reputation score as an intermediary assessment that guides further processing decisions. This intermediary evaluation layer refines quality measurement by combining multiple automated indicators into a comprehensive score that predicts data reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240223615A1System and method for data set creation with crowd-based reinforcement
Publication Date: 2024.07.04 QOMPLX INC
  • US20240223615A1 patent drawing
  • US20240223615A1 patent drawing
  • US20240223615A1 patent drawing

AI summary

A system and method for creation, augmentation, and expansion of high-quality data set collections for training of machine learning algorithms via crowdsourced curation that utilizes a data marketplace which incentivizes data gatherers, publishers, and users to contribute to the creation of a vast resource of reliable data set and knowledge collections and classifications. Data is automatically ingested from disparate sources and autonomously checked for data quality, provenance, uncertainty, and risks and subsequently given a score for a given use context. Data stewards curate a queue of low scoring real data as well as synthetically generated data for specific applications. All reputable data is stored for user (machine or human) consumption and further iterative data or model generation or utilization with appropriate provenance, curation, and license/use limitations and terms.