Outlier Detection Algorithm for Transaction Data Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex computer-implemented workflow processes face challenges in ensuring data quality and consistency due to user errors and the limitations of pre-defined business rules, which are costly and often only applied to critical fields, leading to potential erroneous data and system exceptions.
Innovation Solution
An outlier detection algorithm is implemented to identify anomalous entries in graphical user interfaces by analyzing the number of similar records, distinct values, and same values, providing alerts and modifying workflows to address outliers, thereby reducing dependency on human expertise and enabling adaptive data governance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If pre-defined business rules are used to verify data quality, then data accuracy in critical fields is improved, but implementation cost and complexity increase
Solution Approach 1:
The system performs self-service by automatically detecting outliers through statistical analysis of data patterns without requiring manual definition of business rules. The outlier detection algorithm autonomously identifies anomalous entries by comparing data against historical patterns and statistical thresholds, eliminating the need for expensive manual rule configuration while maintaining data quality verification.
2Reliability
If pre-defined business rules are implemented for data verification, then data quality in critical fields is improved, but the scope is limited to only critical fields
Solution Approach 1:
The outlier detection algorithm provides universal applicability across all data fields regardless of their criticality. Unlike traditional business rules that must be individually configured for each field, the statistical outlier detection mechanism uniformly applies to all numerical and categorical data, automatically adapting to different field types and providing consistent quality verification system-wide.
3Reliability
If manual approval processes are used to verify user-generated data, then data consistency is improved, but processing time and human expertise requirements increase
Solution Approach 1:
The system replaces the mechanical manual approval process with an automated computational outlier detection mechanism. The algorithm instantly analyzes data entries using statistical methods to identify inconsistencies, eliminating the time-consuming human review process while maintaining or improving detection accuracy through systematic mathematical analysis of data patterns.
4Productivity
If user-generated entries are accepted without verification, then processing speed is improved, but user errors and typographical mistakes increase data quality issues
Solution Approach 1:
The outlier detection algorithm performs preliminary verification of data entries immediately upon submission, before further processing occurs. By automatically detecting and flagging anomalous values in real-time, the system prevents erroneous data from propagating through subsequent workflow stages, thereby maintaining high processing speed while ensuring data quality through proactive error detection.
Data Source
AI summary
Transaction data is received from a remote client computing device that includes user-generated entries in each of a plurality of fields. Thereafter, it can be determined, using an outlier detection algorithm, that values for one or more of the entries is an outlier. Data can then be provided (e.g., displayed in a visual display, loaded into memory, stored in physical persistence, transmitted to a remote computing system, etc.). The outlier algorithm can be based on a number of similar records g, a number of distinct values d(g) in the similar records, and a number of same values s in the similar records. Related apparatus, systems, techniques and articles are also described.


