Reference Data Gap Detection Using Version-Difference Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting reference data standardization gaps in data sets are labor-intensive, costly, and prone to errors, and existing data profiling technologies fail to distinguish between outdated reference data versions and other data quality issues, requiring significant computing resources and often only provide partial solutions.
Innovation Solution
A computer-implemented method and system that identifies reference data candidates using an index, determines differences between earlier and current reference data sets, and compares these differences with an inverted index to identify standardization gaps, while eliminating false positives and allowing for user-defined thresholds to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data profiling is performed against the current set of reference data sets, then data quality issues can be detected, but computing resources are significantly required and the process cannot distinguish between outdated reference data versions and other data quality issues
Solution Approach 1:
The patent segments the reference data comparison process into two distinct phases: first comparing operational data against the current reference data set version to identify potential issues, then separately comparing against historical reference data set versions to determine if outdated versions are the root cause. This segmentation allows precise identification of standardization gaps while optimizing computing resource usage by avoiding unnecessary full-profile comparisons against all historical versions.
2Measurement precision
If data management experts manually identify standardization gaps, then accurate detection is possible, but the approach is labor-intensive, costly, and error-prone
Solution Approach 1:
The patent implements a self-service automated system that performs standardization gap detection without requiring manual intervention from data management experts. The system automatically compares operational data against both current and historical reference data set versions, identifies standardization gaps, and generates reports. This self-service approach maintains high detection accuracy while eliminating the labor-intensive, costly, and error-prone manual processes previously required.
3Reliability
If full set of data records is used in data profiling projects, then all standardization issues can be caught, but a lot of computing resources are required
Solution Approach 1:
The patent applies preliminary action by first comparing operational data against the current reference data set version to identify potential standardization issues before proceeding to historical version comparisons. This preliminary filtering step reduces the data set that requires extensive processing against historical versions, thereby maintaining detection completeness while significantly improving processing efficiency and reducing computing resource requirements.
4Adaptability or versatility
If reference data sets are standardized across all business units, then data aggregation into useful business content is enabled, but companies with acquired business units have areas where reference data has not yet been standardized
Solution Approach 1:
The patent implements feedback by automatically detecting standardization gaps in acquired or newly integrated business units and providing detailed reports that highlight specific areas requiring standardization. This feedback mechanism enables data management teams to prioritize and systematically address standardization needs across the organization, facilitating gradual integration of diverse reference data sets while maintaining the ability to aggregate data into useful business content.
Data Source
AI summary
A computer-implemented method for detecting reference data standardization gaps in data sets is disclosed. The method comprises identifying at least one reference data candidate in a data set, using an index for values of the identified at least one reference data candidate, and determining a difference between an earlier version of a reference data set relating to the reference data candidate and a current version of the reference data set. Furthermore, the method comprises comparing the determined difference with values of the index, and identifying entries in the at least one reference data candidate having a value identical to a value of the difference as reference data standardization gap.


