Dynamic Outlier Bias Reduction in Data Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for removing outlier data in data-driven model development are subjective and prone to introducing bias, affecting the fairness and representativeness of calculations, particularly in complex analyses like greenhouse gas emissions standards.
Innovation Solution
A dynamic statistical process that selects bias criteria, generates error thresholds, and iteratively refines model coefficients to objectively remove outliers, using absolute and relative errors, and optimization models to minimize prediction errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If semi-quantitative processes with subjective input are used to remove outliers, then the process is simpler to implement, but the objectivity and reliability of the analysis are reduced
Solution Approach 1:
The patent replaces the manual, subjective mechanical process of outlier identification with an automated statistical system. The system uses quantitative metrics (z-scores, interquartile ranges, standard deviations) to objectively identify and remove outliers, eliminating human bias while maintaining implementation feasibility through algorithmic automation.
Solution Approach 2:
The patent introduces specific statistical parameters (z-score thresholds, interquartile range multiples, standard deviation thresholds) to define outlier boundaries. By changing from subjective judgment to parameter-based definitions, the system achieves both objectivity and repeatability while remaining computationally straightforward to implement.
2Reliability
If outlier data is removed to reduce bias, then the representativeness of the analysis is improved, but the risk of introducing new bias through data censoring increases
Solution Approach 1:
The patent implements feedback mechanisms where the outlier removal process is iteratively applied, with each iteration using the results of the previous iteration to refine the analysis. The system continuously monitors whether outlier removal is improving or degrading the statistical quality metrics, and adjusts the removal threshold accordingly to prevent over-censoring.
Solution Approach 2:
The patent applies partial outlier removal by using configurable thresholds that remove only the most extreme values rather than all values beyond arbitrary boundaries. The system allows for graduated removal strategies where less extreme outliers are retained, balancing bias reduction with preservation of data representativeness.
3Measurement precision
If dynamic iterative processes are used to refine model coefficients, then the accuracy of the model is improved, but the computational complexity and time required increase
Solution Approach 1:
The patent performs preliminary outlier removal and data filtering before the main model coefficient refinement process. By pre-processing the data to eliminate obvious outliers and establish initial coefficient estimates, the system reduces the computational burden of subsequent iterative refinement while maintaining final model accuracy.
Solution Approach 2:
The patent implements periodic iterative refinement where model coefficients are updated at discrete intervals rather than continuously. Each iteration refines the coefficients based on current data quality assessments, and the process terminates when convergence criteria are met, balancing accuracy improvement against computational cost.
Data Source
AI summary
A system and method is described herein for data filtering to reduce functional, and trend line outlier bias. Outliers are removed from the data set through an objective statistical method. Bias is determined based on absolute, relative error, or both. Error values are computed from the data, model coefficients, or trend line calculations. Outlier data records are removed when the error values are greater than or equal to the user-supplied criteria. For optimization methods or other iterative calculations, the removed data are re-applied each iteration to the model computing new results. Using model values for the complete dataset, new error values are computed and the outlier bias reduction procedure is re-applied. Overall error is minimized for model coefficients and outlier removed data in an iterative fashion until user defined error improvement limits are reached. The filtered data may be used for validation, outlier bias reduction and data quality operations.


