Data Veracity Scoring for Data Lake Query Trust
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data lakes face challenges in ensuring the veracity of data from diverse sources, leading to unreliable query results, which can impact business decisions and attract regulatory penalties.
Innovation Solution
Implementing data lineage techniques to associate metadata with data sets, calculating veracity scores based on trust attributes like ancestry, signatures, retention, hash values, and immutability, and combining these scores with query results to provide a framework for trusted queries and models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data from diverse sources is stored in a data lake to enable agile business queries, then the versatility and business insight capability are improved, but the reliability and trustworthiness of query results deteriorate due to potential data inaccuracy
Solution Approach 1:
The system performs preliminary actions by calculating veracity scores for data sets before they are queried. Metadata representing veracity scores is pre-computed and stored with the data sets, so that when queries are executed, the reliability information is already available without needing to recalculate it at query time. This allows the system to maintain both versatility in handling diverse data sources and reliability in query results through advance preparation of trust metrics.
Solution Approach 2:
The invention introduces veracity scores as an intermediary element between the diverse data sources and the query results. These scores act as a mediator that quantifies the trustworthiness of data from different sources, allowing the system to handle diverse data with varying reliability levels. The veracity metadata serves as a bridge that enables agile queries while providing transparency about data quality, thus resolving the contradiction between versatility and reliability.
2Reliability
If veracity metadata is stored with each data set to indicate trustworthiness, then the reliability of data usage is improved, but the device complexity and storage requirements worsen
Solution Approach 1:
The system simplifies complexity by changing the representation of veracity from a complex multi-dimensional assessment to a single numerical parameter - the veracity score. This parameter change allows the system to track data reliability without managing complex metadata structures. The veracity score condenses multiple trust attributes into one manageable metric, reducing the complexity of metadata storage and retrieval while maintaining reliable veracity tracking.
3Measurement precision
If veracity scores are calculated and stored for all data sets, then the measurement precision of data quality assessment is improved, but the loss of time and computational resources worsen
Solution Approach 1:
The system applies preliminary action by pre-calculating veracity scores when data sets are ingested or updated, rather than calculating them at query time. This advance computation stores the measurement results in metadata, eliminating repeated calculation overhead. The precision of veracity assessment is maintained through accurate initial measurement, while time loss is reduced by avoiding redundant calculations during query operations.
Data Source
AI summary
Techniques for determining and representing the veracity of data stored in a data repository and results of queries directed to the stored data by utilizing information lineage that is indicative of the veracity of the stored data. For example, in one example, one or more data repositories are maintained. The one or more data repositories comprise metadata representative of the veracity of one or more data sets stored in the one or more data repositories. In response to a query to at least one data set of the one or more data sets stored in the one or more data repositories, a result of the query for the at least one data set is returned in combination with corresponding metadata representing the veracity of the at least one data set.


