Data Curation System Resolving Ambiguous Values via Source Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Insufficient data curation resources lead to uncurated or partially curated data, which can result in unreliable computer-implemented services due to ambiguous values in datasets, causing disruptions and impacting the quality of downstream applications.
Innovation Solution
A method and system that identify ambiguous values in data, generate potential replacement values, and interact with the data source for validation to obtain a final replacement value, thereby reducing the burden on curation resources and ensuring data reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data curation resources are limited, then productivity is improved by processing less data, but reliability deteriorates due to uncurated or partially curated data
Solution Approach 1:
The data source system performs self-curation by automatically validating and replacing ambiguous values in its own data without requiring external curation resources. The system identifies ambiguous values, generates potential replacement values, interacts with itself to validate replacements, and updates its data pipeline autonomously, thereby maintaining high reliability without consuming additional curation resources.
Solution Approach 2:
An intermediary validation mechanism is introduced between data extraction and data pipeline population. This intermediary layer handles the ambiguous value resolution process, acting as a buffer that ensures data quality is maintained while allowing the main data flow to continue uninterrupted. The intermediary system coordinates between the data source and the data pipeline to ensure reliable data transfer.
2Productivity
If ambiguous values are replaced without interaction with data source, then productivity is improved by faster data processing, but reliability deteriorates due to potential incorrect replacements
Solution Approach 1:
The system implements a feedback loop where potential replacement values are presented back to the data source for validation. The data source provides feedback on whether the suggested replacements are accurate, allowing the system to learn from previous replacements and improve future validation decisions. This feedback mechanism ensures high reliability while maintaining efficient automated processing.
Solution Approach 2:
The system performs preliminary actions by generating potential replacement values before finalizing data pipeline population. Ambiguous values are identified and potential replacements are generated in advance, allowing the system to prepare cleaned data structures ahead of time. This preliminary processing enables faster overall data pipeline operation while maintaining accuracy through subsequent validation.
Data Source
AI summary
Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, the data may be screened for ambiguous values. A potential replacement value for each ambiguous value may be provided to the data source and the data source may indicate whether the potential replacement value should be used in the data pipeline as a final replacement value for the ambiguous value.


