Source Data Integrity Checks Using Statistical Forecasts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing processes rely on data from sources that may lack integrity, leading to poor quality outputs, especially when using externally or internally sourced data, which can undermine stakeholder confidence and lead to errors that are not detected until later.
Innovation Solution
A system and method for examining data integrity and quality using statistical models, including constant, linear, and quadratic models, to identify unexpected values and flag or halt processes until issues are resolved.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data from external or internal sources is used without integrity controls, then the process can operate without additional verification steps, but the output quality deteriorates due to poor data inputs
Solution Approach 1:
The patent applies preliminary action by performing data integrity checks and statistical model validations before the main processing occurs. The system generates forecasts using multiple statistical models (constant, linear, quadratic) and compares actual data against these forecasts to identify unexpected values before they propagate through the processing system, thereby preventing poor quality outputs without significantly impacting overall process speed
Solution Approach 2:
The patent introduces an intermediary verification layer between data input and processing. This intermediary system uses statistical models and anomaly detection to mediate between raw source data and the main processing pipeline, filtering out unexpected values and ensuring data quality before downstream operations, thus maintaining both productivity and reliability
2Reliability
If statistical models and data integrity checks are implemented, then data quality and reliability improve, but the process complexity increases
Solution Approach 1:
The patent segments the data verification process into distinct, modular components: data reception, historical data comparison, multiple statistical model generation (constant, linear, quadratic), forecast comparison, and unexpected value identification. Each component performs a specific function and can be independently maintained, reducing overall system complexity while ensuring comprehensive data integrity checking
Solution Approach 2:
The patent manages complexity by dynamically adjusting model parameters and selecting appropriate statistical models based on data characteristics. The system uses parameter-based approaches (constant, linear, quadratic models) that can be selected and configured without fundamentally changing the system architecture, allowing flexible data integrity verification with controlled complexity
3Measurement precision
If multiple statistical models are generated and compared, then measurement precision of unexpected values improves, but the time required for data examination increases
Solution Approach 1:
The patent applies partial action by generating multiple statistical models (constant, linear, quadratic) only when necessary for accurate detection, rather than always using the most complex approach. The system selectively applies different levels of model complexity based on the data characteristics and required detection precision, optimizing the balance between accuracy and examination time
Solution Approach 2:
The system performs preliminary analysis to determine the appropriate level of model complexity needed. By pre-assessing data patterns and selecting suitable statistical models before full examination, the system avoids unnecessary computational overhead while maintaining high detection accuracy for unexpected values
Data Source
AI summary
A system and method are provided for examining data from a source. The method is executed by a device having a processor and includes receiving a set of historical data and a set of current data to be examined, from the source. The method also includes generating multiple statistical models based on the historical data and a forecast for each model. The method also includes selecting one of the multiple statistical models based on at least one criterion, and generating a new forecast using the selected model. The method also includes comparing the set of current data against the new forecast to identify any data points in the set of current data with unexpected values. The method also includes outputting a result of the comparison, the result comprising any data points with unexpected values.


