Incremental Dataset Maintenance via Delta Joins
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data cleansing techniques are computationally taxing and inefficient, leading to long processing times and stale datasets due to the need to cleanse large volumes of transaction data infrequently, which hinders data analytics and increases call center volumes for resolving transaction issues.
Innovation Solution
A data cleansing platform that periodically cleanses raw transaction records by generating a delta dataset with new and updated merchant keys, allowing for a data join operation to maintain a cleansed dataset in a timely manner, reducing computational resources required and keeping the merchant lookup table up-to-date.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data cleansing techniques are used to process large volumes of transaction data, then data accuracy is improved, but processing time increases significantly and computational resources are excessively consumed
Solution Approach 1:
The patent segments the data cleansing process into two distinct phases: (1) an initial comprehensive cleanse of historical transaction data to establish a baseline cleansed dataset, and (2) subsequent incremental updates using only new transactions since the last cleanse. This segmentation allows the system to maintain data accuracy while dramatically reducing processing time for ongoing operations, as the incremental phase processes only new data rather than re-processing the entire dataset.
Solution Approach 2:
The patent performs a preliminary comprehensive data cleanse operation to establish a baseline cleansed dataset containing entity keys for all historical transactions. This preliminary action creates a reusable reference dataset that enables subsequent incremental updates to proceed efficiently by comparing only new transactions against the pre-established baseline, thereby reducing ongoing processing time and resource consumption.
2Reliability
If comprehensive data cleansing is performed on all transaction data, then data completeness is improved, but computational resource consumption increases
Solution Approach 1:
The patent segments the data processing workload by separating the comprehensive initial cleanse of historical data from subsequent incremental updates. The initial phase ensures complete data cleansing and entity key generation for all historical transactions, while the incremental phase processes only new transactions by comparing them against the pre-established baseline. This segmentation maintains data completeness while significantly reducing computational resource consumption for ongoing operations.
Solution Approach 2:
The patent applies partial action by performing complete data cleansing only on the initial historical dataset, and then applying incremental cleansing only to new transactions since the last cleanse. Rather than repeatedly processing the entire dataset, the system processes only the necessary subset (new transactions), thereby maintaining data completeness while reducing computational resource consumption for periodic updates.
3Loss of energy
If data cleansing is performed infrequently to reduce processing load, then computational resources are conserved, but dataset freshness deteriorates
Solution Approach 1:
The patent segments the data maintenance operation into an initial comprehensive cleanse and subsequent incremental updates. The incremental update mechanism allows the system to process only new transactions since the last cleanse by comparing them against the baseline dataset, enabling frequent updates with minimal computational overhead. This segmentation enables the system to maintain dataset freshness while conserving computational resources.
Solution Approach 2:
The patent implements periodic incremental data cleansing operations that run at scheduled intervals to update the dataset with new transactions. By using the baseline dataset as a reference, the periodic incremental updates can be performed quickly and efficiently, maintaining dataset freshness without requiring extensive computational resources, thus enabling frequent updates rather than infrequent comprehensive cleanses.
Data Source
AI summary
In some implementations, a data cleaning platform may determine a respective entity key for each data record in a cleansed dataset based on a combination of fields, in each data record, that contain information that uniquely identifies an entity associated with a respective data record. The data cleaning platform may generate a delta dataset based on a set of uncleansed data records related to transactions that occurred after a time when the cleansed dataset was first generated. For example, in some implementations, each uncleansed data record in the delta dataset may be associated with a corresponding entity key based on the combination of fields. The data cleaning platform may perform a data join to update the cleansed dataset to include data records related to the transactions that occurred after the time when the cleansed dataset was first generated.


