Dynamic Outlier Bias Reduction in Statistical Model Development
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for removing outlier data in statistical and mathematical model development are often subjective and prone to bias, lacking an objective, dynamic statistical process for data quality and validation operations.
Innovation Solution
A computer-implemented method that selects bias criteria, generates predicted values, error sets, and error threshold values to create a censored data set, iteratively refining model coefficients using linear or non-linear optimization until performance termination criteria are met, ensuring objective outlier bias reduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If outlier data is removed using a semi-quantitative process requiring subjective input, then the process is simple to implement, but the analysis results are biased and lack objectivity
Solution Approach 1:
The patent transforms the outlier removal process from a subjective manual operation to an objective automated process by changing the parameters from human judgment criteria to statistical metrics (error thresholds, standard deviations, confidence intervals). The system dynamically adjusts these parameters based on data characteristics, enabling objective outlier identification without manual intervention while maintaining process simplicity through algorithmic automation.
Solution Approach 2:
The patent replaces the mechanical/manual process of subjective outlier removal with an automated computational system that uses statistical algorithms and mathematical models. The system automatically calculates error thresholds, identifies outliers based on predefined statistical criteria, and removes them from the dataset, eliminating human subjectivity while maintaining ease of use through automated execution.
2Reliability
If outlier data is removed to reduce bias, then the fairness of analysis is improved, but the risk of changing calculation results through data censoring increases
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors the impact of outlier removal on calculation results. By tracking changes in statistical metrics and model performance before and after outlier removal, the system can detect when removal operations are excessively altering results. This feedback loop allows the system to adjust outlier removal criteria to maintain fairness while preserving valid data, thus reducing the risk of information loss.
Solution Approach 2:
The patent applies partial outlier removal by using dynamic thresholds that adapt to data characteristics. Instead of removing all potential outliers or using fixed aggressive criteria, the system removes only those outliers that meet dynamically calculated thresholds based on statistical significance. This partial action approach ensures sufficient bias reduction while minimizing the loss of valid data points that would be incorrectly identified as outliers.
3Measurement precision
If iterative refinement of model coefficients is performed, then the accuracy of statistical calculations is improved, but the computation time increases
Solution Approach 1:
The patent performs preliminary actions by pre-calculating and storing statistical properties of the data (mean, standard deviation, error distributions) before the iterative refinement process. These pre-computed statistics are used to initialize model coefficients and set appropriate convergence criteria, reducing the number of iterative cycles needed. The system also pre-identifies potential outlier patterns, allowing the iterative process to focus only on refining coefficients rather than detecting outliers during each iteration, thus improving accuracy while reducing computation time.
Data Source
AI summary
A system and method is described herein for data filtering to reduce functional, and trend line outlier bias. Outliers are removed from the data set through an objective statistical method. Bias is determined based on absolute, relative error, or both. Error values are computed from the data, model coefficients, or trend line calculations. Outlier data records are removed when the error values are greater than or equal to the user-supplied criteria. For optimization methods or other iterative calculations, the removed data are re-applied each iteration to the model computing new results. Using model values for the complete dataset, new error values are computed and the outlier bias reduction procedure is re-applied. Overall error is minimized for model coefficients and outlier removed data in an iterative fashion until user defined error improvement limits are reached. The filtered data may be used for validation, outlier bias reduction and data quality operations.


