Crowdsourced Data Discovery for Enterprise Curation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The proliferation of big data with disparate sources, varying volumes, and rapid velocity makes traditional data curation methods unsustainable, as they struggle to maintain consistency, scalability, and accuracy in an ever-changing data landscape.
Innovation Solution
A data discovery solution that leverages crowdsourcing through anonymized multi-tenancy and an internet-scale matching and validation gaming platform, utilizing mobile application games to improve automated data curation and create a comprehensive library of semantic-technical mappings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional manual data curation methods are used, then data quality and accuracy can be maintained, but scalability and productivity deteriorate as data volume increases
Solution Approach 1:
The system enables data to curate itself through automated classification algorithms and machine learning models that automatically categorize, tag, and organize data without human intervention, allowing the system to serve itself at scale while maintaining consistency and quality standards
Solution Approach 2:
Manual mechanical curation processes are replaced with automated computational systems including AI/ML algorithms, natural language processing, and automated classification engines that process data at machine speed, eliminating the bottleneck of human manual work while preserving data quality through programmable validation rules
2Productivity
If automated machine learning models are used for data discovery, then productivity increases, but measurement precision deteriorates due to false positives
Solution Approach 1:
The system implements feedback loops where automated classification results are continuously validated, refined, and improved through performance monitoring, error correction mechanisms, and retraining on validated data, allowing the system to learn from its mistakes and improve precision over time while maintaining high productivity
Solution Approach 2:
The system performs preliminary automated classification to quickly identify potential matches and patterns, then applies more rigorous validation and verification steps only to promising candidates, enabling fast initial processing followed by precision-focused validation of high-priority results
3Measurement precision
If data curators manually evaluate and categorize data from diverse sources, then measurement precision is maintained, but loss of time increases due to the volume and velocity of big data
Solution Approach 1:
The curation process is segmented into distinct automated stages including initial data ingestion, classification, tagging, validation, and archiving, with each stage handled by specialized algorithms that process different aspects of data curation in parallel, dramatically reducing total processing time while maintaining comprehensive quality control
Solution Approach 2:
The system implements continuous automated curation processes that operate without interruption, with data flowing through classification and validation pipelines in real-time as it is generated, eliminating idle time and ensuring continuous value extraction from incoming data streams without sacrificing accuracy
4Adaptability or versatility
If enterprises use disparate data sources with varying volumes and velocities, then adaptability increases, but device complexity increases making consistent curation difficult
Solution Approach 1:
The system implements a universal data curation platform with standardized interfaces and processing pipelines that can handle multiple data sources, formats, and volumes through a single unified system, eliminating the need for separate curation systems for each data source and reducing overall complexity through consolidation
Data Source
AI summary
Disclosed are methods and systems for a data discovery solution which harnesses the power of crowdsourcing to improve automated data curation. This is done in two complimentary ways: (a) large scale collective curation through anonymized multi-tenancy, and (b) and through internet scale matching and validation gaming platform using mobile application game. The result is the most extensive library of semantic-technical mappings of the enterprise data, which are immediately at hand to provide a fast, easy and a good understanding of the enterprise data. The data discovery solution forms a gateway for governing and unlocking value from big data.


