Prioritizing Data Curation Targets by Downstream Impact
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Insufficient data curation resources lead to uncurated or partially curated data, which can result in untrustworthy data being provided to downstream consumers, potentially disrupting computer-implemented services.
Innovation Solution
A method and system for prioritizing data curation targets based on impact scores, ensuring that high-priority targets are curated first using available resources, thereby generating more trustworthy partially curated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data curation resources are increased to curate all identified curation targets, then data trustworthiness is improved, but resource consumption and cost increase
Solution Approach 1:
The patent applies partial action by curating only a prioritized subset of curation targets rather than all targets. The system identifies and curates high-priority curation targets first based on impact scores, accepting that not all data will be fully curated due to resource constraints. This resolves the contradiction by achieving acceptable data trustworthiness for critical targets while limiting resource consumption.
Solution Approach 2:
The patent implements local quality by applying different curation priorities to different data targets. High-impact curation targets receive full curation resources and attention, while lower-impact targets receive reduced or deferred curation. This differentiated approach optimizes the balance between data trustworthiness and resource allocation by focusing resources where they have the greatest impact.
2Manufacturing precision
If data curation is performed on all identified targets, then data quality is improved, but processing time and productivity are reduced
Solution Approach 1:
The system performs partial curation on high-priority targets within a target time period rather than attempting to curate all targets. This approach maintains acceptable data quality for critical targets while ensuring that the curation process completes within the specified time frame, thus preserving productivity.
Solution Approach 2:
The patent segments the curation process by dividing curation targets into priority groups based on impact scores. High-priority targets are curated first within the target time period, while lower-priority targets are deferred to future periods. This segmentation enables the system to maintain data quality for critical targets without compromising overall processing speed and productivity.
3Reliability
If high-priority curation targets are selected and curated first, then downstream service reliability is improved, but data completeness is reduced
Solution Approach 1:
The patent applies local quality by ensuring high-quality curation for high-priority targets that directly impact downstream services, while accepting lower or partial curation quality for lower-priority targets. This approach prioritizes service reliability by guaranteeing data quality where it matters most, while minimizing information loss in less critical areas.
Data Source
AI summary
Methods and systems for curating data by a data manager are disclosed. Data may be curated from various data sources before being provided to downstream consumers that may rely on the trustworthiness of the curated data in order to provide desired computer-implemented services. During the data curation process, data curation resources are used to improve the trustworthiness and/or value of the collected data. However, data curation resources (e.g., data curators, computing resources) may be limited and/or insufficient to perform the data curation process as desired, which may result in unusable and/or uncurated (e.g., untrustworthy) data. Thus, portions of the data (e.g., curation targets) may be prioritized (e.g., relative to other curation targets). The curation targets may be curated with the available data curation resources based on their relative priority in order to reduce the likelihood of providing untrustworthy data to entities (e.g., downstream consumers) that facilitate the computer-implemented services.


