Data Deviation Detection via Relation Pattern Algorithms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data inconsistency detection tools require manual configuration and struggle with complex data models, lack of documentation, and dynamic data changes, leading to inefficiencies in identifying deviations between data sources.
Innovation Solution
A method and apparatus that automatically detect deviations by identifying data post pairs with matching attributes, applying relation pattern algorithms, determining conformity levels, and selecting the best-suited algorithm to analyze data value combinations for non-conformance, thereby identifying potential data deviations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual configuration is used to instruct tools what to look for in data sources, then detection accuracy can be improved, but operation complexity and time consumption increase significantly
Solution Approach 1:
The system automatically discovers data models, relationships, and deviation patterns without requiring manual configuration. The tool self-adapts to different data sources by analyzing their structures and learning what constitutes a deviation, eliminating the need for operators to manually instruct the tool what to look for while maintaining high detection accuracy
Solution Approach 2:
The system dynamically adjusts detection parameters and algorithms based on the specific characteristics of each data source. By automatically adapting parameters such as relation patterns, data models, and deviation thresholds to match the observed data structures, the system achieves high accuracy without manual configuration
2Reliability
If comprehensive manual configuration is provided for each data source combination, then detection completeness improves, but time consumption and resource usage increase
Solution Approach 1:
The system performs preliminary automatic analysis of data sources to discover their structures, relationships, and characteristics before actual deviation detection begins. This preliminary configuration phase is executed automatically and stored for reuse, ensuring comprehensive detection coverage while minimizing the time required for actual detection operations
Solution Approach 2:
The system creates a universal detection framework that automatically adapts to multiple data source combinations. By discovering and storing data models and relationships in a reusable format, the same framework can detect deviations across different data sources without requiring separate manual configuration for each combination, reducing overall time consumption
3Measurement precision
If detailed documentation of data models is required for each vendor system, then detection accuracy improves, but accessibility and ease of operation deteriorate
Solution Approach 1:
The system automatically discovers and learns data models from vendor systems without requiring external documentation. By analyzing the actual data structures, relationships, and patterns directly from the sources, the system achieves accurate detection while remaining accessible to systems from any vendor, eliminating the barrier of requiring detailed documentation
4Ease of operation
If static deviation criteria are used for data comparison, then detection simplicity is maintained, but adaptability to changing data patterns deteriorates
Solution Approach 1:
The system uses dynamic deviation criteria that automatically adapt to changing data patterns while maintaining operational simplicity. By continuously learning from observed data relationships and adjusting detection thresholds and patterns accordingly, the system remains both simple to operate and highly adaptable to evolving data characteristics without requiring manual reconfiguration
Data Source
AI summary
The present disclosure describes a method and an apparatus for detecting deviations in data sources, each data source comprising a plurality of data posts, each data post comprising a number of data values. The method comprises identifying (102) data post pairs, each pair comprising a first data post in a first data source and a second data post in a second data source, wherein, for a unique matching data attribute of the first data post and the second data post in a data post pair, a subset of the data value is equal. The method further comprises determining (104) whether individual of a plurality of combinations of data values of the first data post with data values of the second data post within each of the plurality of data post pairs fulfill individual of a plurality of relation pattern algorithms, and determining (106) a conformity level for the determined fulfillment of relation pattern algorithms for the plurality of data post pairs. The method further comprises selecting (108) relation pattern algorithm from the plurality of relation pattern algorithms based on the determined conformity level, and analyzing (110) data value combinations of individual data post pairs in relation to the selected relation pattern algorithm in order to detect data value combinations of individual data post pairs that does not conform to the selected relation pattern algorithm, a non-conformance indicating (114) a possible deviation in data of the individual data post pair.


