Crowdsourced Data Marketplace for ML Training Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning algorithms lack sufficient high-quality data sets for training, as existing data sets are often insufficient in size and quality, and the generation of synthetic data introduces biases and over-fitting.
Innovation Solution
A system and method for creating high-quality data sets through crowdsourced curation, utilizing a data marketplace that incentivizes contributors, with data automatically scored for quality and provenance, and human curation of low-scoring data, combined with synthetic data generation using generative adversarial networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic data is generated to increase data supply, then the quantity of training data is improved, but the quality deteriorates due to over-fitting and inherited quality problems
Solution Approach 1:
The patent segments the data curation process into multiple independent components: automated ingestion from disparate sources, reputation scoring subsystem, human steward verification queue, and synthetic data generation. Each component operates independently with specific quality checks, allowing the system to process large volumes of data while maintaining quality through distributed verification rather than centralized bottlenecks.
Solution Approach 2:
The patent implements a feedback loop where data stewards verify and correct synthetic data, which then feeds back into improving the reputation scoring algorithm. The system continuously learns from human verification outcomes, adjusting scoring metrics and thresholds to better identify high-quality data. This closed-loop feedback ensures that synthetic data generation improves over time while maintaining reliability.
2Reliability
If manual curation is used to ensure data quality, then the reliability of data sets is improved, but the productivity deteriorates due to time-consuming manual processes
Solution Approach 1:
The patent applies local quality by implementing different verification strategies for different data sources and types. High-risk sources undergo rigorous human steward review, while trusted sources with established reputations receive automated processing. Data entries are evaluated individually based on their specific characteristics rather than applying uniform manual review to all data, optimizing both quality and throughput.
Solution Approach 2:
The patent performs preliminary automated filtering and reputation scoring before data reaches human stewards. The system pre-processes incoming data through multiple automated quality checks, source reputation evaluation, and initial validation rules, so that human curators only need to review data that requires expert judgment. This preliminary action significantly reduces the manual workload while maintaining high quality standards.
3Productivity
If automated ingestion is used to increase data processing speed, then the productivity is improved, but the measurement precision deteriorates due to difficulties in detecting and measuring data quality
Solution Approach 1:
The patent implements a universal reputation scoring system that evaluates data from multiple disparate sources using a consistent multi-dimensional framework. The same scoring algorithm assesses structured data, unstructured text, images, and other data types, applying universal quality metrics across diverse formats. This multi-functional approach maintains measurement precision while handling varied data types at high speed through automated processing.
Solution Approach 2:
The patent introduces an intermediary reputation scoring layer between automated ingestion and final data acceptance. Rather than directly accepting or rejecting data based on single metrics, the system uses the reputation score as an intermediary assessment that guides further processing decisions. This intermediary evaluation layer refines quality measurement by combining multiple automated indicators into a comprehensive score that predicts data reliability.
Data Source
AI summary
A system and method for creation, augmentation, and expansion of high-quality data set collections for training of machine learning algorithms via crowdsourced curation that utilizes a data marketplace which incentivizes data gatherers, publishers, and users to contribute to the creation of a vast resource of reliable data set and knowledge collections and classifications. Data is automatically ingested from disparate sources and autonomously checked for data quality, provenance, uncertainty, and risks and subsequently given a score for a given use context. Data stewards curate a queue of low scoring real data as well as synthetically generated data for specific applications. All reputable data is stored for user (machine or human) consumption and further iterative data or model generation or utilization with appropriate provenance, curation, and license/use limitations and terms.


