High-Trust Dataset Governance for Conflicting Data Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems lack efficient methods for building high trust datasets due to varying data accuracy, completeness, and quality across different data sources, leading to the use of stale and inadequate information, which can impact planning and resource efficiency, particularly in healthcare.
Innovation Solution
A method and system that retrieve data from multiple sources with associated trust scores, identify conflicting data subsets, select the most reliable source based on trust scores, and update the high trust dataset and source scores, using AI and data analytics for real-time or continuous updates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data is retrieved from multiple data sources without centralized governance, then data availability is improved, but data quality and reliability deteriorate due to varying accuracy and completeness
Solution Approach 1:
The patent introduces a centralized data governance system that acts as an intermediary between multiple data sources and end users. This system retrieves data from various sources, evaluates their trust scores, resolves conflicts, and provides a unified high-trust dataset, thereby maintaining data availability while ensuring quality through centralized control
Solution Approach 2:
The system dynamically changes the trust score parameter for each data source based on continuous evaluation of data quality metrics. By adjusting these parameters, the system can adaptively select higher-quality data sources while maintaining access to multiple sources, thus improving reliability without sacrificing availability
2Reliability
If individual entities evaluate data accuracy independently, then data reliability is improved, but resource consumption increases due to duplicate evaluation efforts
Solution Approach 1:
The centralized data governance system provides a universal data evaluation service that multiple entities can access. Instead of each entity performing independent evaluations, the system performs the evaluation once and makes results available to all, thereby maintaining data reliability while dramatically improving resource efficiency through shared functionality
3Speed
If data sources update at varying rates, then data freshness is improved for time-sensitive applications, but data consistency deteriorates due to conflicts between updated and stale data
Solution Approach 1:
The system implements continuous feedback loops that monitor data conflicts between different sources. When conflicts are detected, the system uses trust scores to identify the most reliable data source and resolves conflicts by prioritizing data from high-trust sources, thereby maintaining consistency while allowing rapid updates from time-sensitive sources
Solution Approach 2:
The system performs preliminary conflict detection and resolution before data is used by end applications. By proactively identifying and resolving conflicts between updated and stale data, the system ensures consistency is maintained while still allowing data sources to update at their optimal frequencies
Data Source
AI summary
Described herein are systems, methods, and non-transitory computer readable medium for building a high trust dataset. The method for building a high trust dataset may comprise. repeatedly, retrieving data from a plurality of data sources each having an associated trust score, identifying at least one subset of the retrieved data which conflicts with at least one of another subset of the retrieved data and the high trust dataset, for each identified subset of the retrieved data, selecting one of the plurality of data sources from which the subset will be included in the high trust dataset based on the associated trust score, based on the trust score, updating the high trust dataset to comprise the subset from the selected one of the plurality of data sources, and updating the associated trust score of at least one of the plurality of data sources.


