Dynamic outlier bias reduction system and method

The method addresses the issue of subjective outlier removal in data analysis by using iterative error-based filtering and optimization to achieve accurate and unbiased data analysis, enhancing the reliability of statistical and mathematical models.

JP2025166116APending Publication Date: 2025-11-05HARTFORD STEAM BOILER INSPECTION & INSURANCE CO
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025133286
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-02-20
Filing Date
2025-08-08
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing data analysis methods fail to objectively remove outlier data during preliminary model development, leading to biased and unrepresentative calculations, particularly in greenhouse gas standards, due to subjective and data-driven approaches that alter calculation results.

Method used

A computer-implemented method for reducing outlier bias through iterative processes involving error threshold setting, model coefficient adjustment, and optimization techniques to ensure data quality and accuracy, using relative and absolute error calculations to filter out outliers.

Benefits of technology

This approach provides an objective and dynamic method to reduce outlier bias, ensuring fair and representative data analysis by iteratively refining model coefficients until performance criteria are met, thereby improving the accuracy and reliability of statistical and mathematical models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025166116000001_ABST
    Figure 2025166116000001_ABST
Patent Text Reader

Abstract

To provide functional system and method for data filtering for reducing outlier bias of a trend line.SOLUTION: The present invention is directed to an objective, statistic method of eliminating an outlier from a data set. The method has the steps of determining a bias based on an absolute error, a relative error or both of them, calculating an error value from calculation of data, model coefficients or trend lines, and eliminating an outlier data record if the error value exceeds a reference set by a user. For the purpose of a repeated calculation such as an optimization method, the eliminated data is re-applied to the model for use in calculation of a new result upon each repeated calculation. With use of a model value for a completed data set, a new error value is calculated, and an outlier bias reduction process is re-applied to minimize the total error for the model coefficient and the outlier eliminated data repeated until the value reaches a user definition error improvement limit. Filtered data is used for the purpose of authentication, the outlier bias reduction, and data quality operation.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to data analysis in which outlier components are removed (or filtered) from the analytical development. The analysis may involve the calculation of simple statistics or complex operations involving mathematical models that use the data in the development. The filtering of outlier data is for the purposes of data quality and data validation operations, or for the purposes of calculating representative standards, statistics, data sets that are applied to subsequent analysis, regression analysis, time series analysis, or qualifying data for mathematical model development.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This continuation-in-part patent application claims priority to U.S. Provisional Patent Application No. 13 / 213,780, entitled "Dynamic Outlier Bias Reduction System and Method," filed August 19, 2011, which is incorporated herein by reference in its entirety. [Background technology]

[0003] Removing outlier data in baseline or data-driven model development is an important part of the preliminary analytical work to ensure that a representative and fair analysis is developed from the underlying data. For example, carbon dioxide (CO2), ozone (O3), water vapor (H2O), hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), chlorofluorocarbons (CFCs), sulfur hexafluoride (SF6), methane (CH4), nitrous oxide (N2O), carbon monoxide (CO), nitrogen oxides (NO xDeveloping fair benchmarks for greenhouse gas standards for CO2 emissions and non-methane volatile organic compounds (NMVOCs) requires that the collected industrial data used in standard development exhibit certain characteristics. Extremely good or bad performance by a few industrial locations should not bias the standards calculated for other industrial locations. Inclusion of such performance results in the standard calculation would be deemed unfair or unrepresentative. In the past, performance outliers were removed through a semi-quantitative process requiring subjective input. Current systems and methods are data-driven approaches, which perform this task as an integral part of model development, rather than during the preliminary analysis or preliminary model development stage.

[0004] Bias removal can be a subjective process in which justification is documented in a prescribed format to justify the data modification. However, all forms of outlier removal are forms of data truncation that potentially alter the calculation results. Such data filtering may or may not reduce bias or error in the calculation, and in the context of full analytical disclosure, rigorous data removal guidelines and documentation for outlier removal should be included in the analytical results. Therefore, there is a need in the industry to provide new systems and methods for objectively removing outlier data bias using a dynamic statistical process useful for data quality operations, data validation, statistical calculations, mathematical model development, etc. The outlier bias removal system and method can also be used to classify data into representative categories. In this case, the data is applied to develop a mathematical model customized for each group. In a preferred embodiment, coefficients are defined as multiplicative and additive factors of the mathematical model and other inherently nonlinear numerical parameters. For example, f(x, y, z) = a*x + b*y c In the mathematical model of +d*sin(ez)+f, a, b, c, d, e, and f are all defined as coefficients. The values ​​of these terms are either fixed or become part of the development of the mathematical model. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2007 / 117233(A1) Pamphlet Summary of the Invention

[0006] A preferred embodiment is a computer-implemented method for reducing outlier bias, comprising the steps of selecting a bias criterion, providing a dataset, providing a set of model coefficients, selecting a set of target values, (1) generating a set of predicted values ​​for the completed dataset, (2) generating an error set for the dataset, (3) generating a set of error thresholds based on the error set and the bias criterion, (4) generating a truncated dataset based on the error set and the set of error thresholds, (5) generating a set of new model coefficients using the set of new model coefficients, and (6) repeating steps (1)-(5) using the set of new model coefficients until a censored performance termination criterion is met. In a preferred embodiment, the set of predicted values ​​is generated based on the dataset and the set of model coefficients. In a preferred embodiment, the error set includes a set of absolute errors and a set of relative errors generated based on the set of predicted values ​​and the set of target values. In another embodiment, the error set includes values ​​calculated as differences between the set of predicted values ​​and the set of target values. In another embodiment, generating the set of new coefficients further includes minimizing the set of errors between the set of predicted values ​​and the set of actual values, which may be achieved using a linear or non-linear optimization model. In a preferred embodiment, the truncated performance termination criterion is based on a standard error and a coefficient of determination.

[0007] Another embodiment includes a computer-implemented method for reducing outlier bias, comprising the steps of selecting an error criterion, selecting a dataset, selecting a set of actual values, selecting an initial set of model coefficients, generating a set of model predictions based on the completed dataset and the initial set of model coefficients, (1) generating a set of errors based on the model predictions and the set of actual values ​​for the completed dataset, (2) generating a set of error thresholds based on the completed set of errors for the completed dataset and the error criterion, (3) generating an outlier-removed dataset, the filtering of which is based on the completed dataset and the set of error thresholds, and (4) filtering the filtered dataset and the set of old coefficients. (5) generating a set of outlier-bias-reduced model predictions based on the filtered data set and the set of new model coefficients, wherein generating the set of outlier-bias-reduced model predictions is performed by a computer processor; (6) generating a set of model performance values ​​based on the model predictions and the set of actual values, repeating steps (1) through (6) using the new set of coefficients in place of the set of coefficients from the previous iteration until a performance termination criterion is met; and storing the set of model predictions on a computer data medium.

[0008] Another embodiment includes a computer-implemented method for reducing outlier bias, comprising the steps of selecting a target variable for a facility, selecting a set of actual values ​​for the target variable, identifying variables for the facility associated with the target variable, obtaining a dataset for the facility, the dataset including values ​​for the variables, selecting a bias criterion, selecting a set of model coefficients, (1) generating a set of predicted values ​​based on the completed dataset and the set of model coefficients, (2) generating a set of censored model performance values ​​based on the set of predicted values ​​and the set of actual values, (3) generating an error set based on the set of predicted values ​​and the set of actual values ​​for the target variable, and (4) calculating a censored model performance value based on the error set and the bias criterion. (5) generating a censored data set based on the data set and the set of error thresholds with the processor; (6) generating a set of new model coefficients based on the censored data set and the set of model coefficients with the processor; (7) generating a set of new predicted values ​​based on the data set and the set of new model coefficients with the processor; (8) generating a set of new censored model performance values ​​based on the set of new predicted values ​​and the set of actual values, repeating steps (1) through (8) using the set of new coefficients until a censored performance termination criterion is met; and storing the set of new model predicted values ​​on a computer data medium.

[0009] Another embodiment includes a computer-implemented method for reducing outlier bias, comprising determining a target variable for a facility, the target variable being a metric for an industrial facility related to its production, financial performance, or emissions, identifying a plurality of variables for the facility, the plurality of variables being direct variables for the facility that affect the target variable and a set of transformed variables for the facility, each of which is a function of at least one direct facility variable that affects the target variable, selecting an error metric comprising an absolute error and a relative error, obtaining a dataset for the facility, the dataset comprising values ​​for the variables, selecting a set of actual values ​​for the target variable, selecting an initial set of model coefficients, generating a set of model predictions based on the completed dataset and the initial set of model coefficients, and generating a completed set of errors based on the completed set of model predictions and the set of actual values, wherein the relative error is calculated according to the formula: Relative Error m =((predicted value m - Actual value m ) / actual value m ) 2 (where "m" is a reference number), and the absolute error is calculated using the formula: Absolute Error m =(predicted value m - Actual value m ) 2generating a set of model performance values ​​based on the set of model predictions and the set of actual values, the set of overall model performance values ​​including a first standard error and a first coefficient of determination; (1) generating a set of errors based on the model predictions and the set of actual values ​​for the completed dataset; (2) generating a set of error thresholds based on the completed set of errors and the error criterion for the completed dataset; (3) generating an outlier-removed dataset by removing data having error values ​​equal to or greater than the error thresholds, the filtering based on the completed dataset and the set of error thresholds; (4) minimizing the error between the set of predictions and the set of actual values ​​using at least one of a linear optimization model and a nonlinear optimization model. (5) generating a set of outlier-bias-reduced model predictions based on the outlier-removed data set and the set of model coefficients, wherein generating the new model predictions is performed by a computer processor; (6) generating a set of overall model performance values ​​based on the set of new predicted model values ​​and the set of actual values, wherein the set of model performance values ​​includes a second standard error and a second coefficient of determination; repeating steps (1)-(6) using the new set of coefficients in place of the set of coefficients from a previous iteration until a performance termination criterion is met, wherein the performance termination criterion includes a standard error termination value and a coefficient of determination termination value; andSatisfying the performance termination criteria includes the standard error termination value being greater than the difference between the first and second standard errors and the coefficient of determination termination value being greater than the difference between the first and second coefficients of determination, and storing the set of new model predictions on a computer data medium.

[0010] Another embodiment includes a computer-implemented method for reducing outlier bias, the method including the following steps. (2) generating an outlier-removed data set, the filtering of which is based on the data set and the set of error thresholds; (3) generating a set of outlier-bias-reduced model forecasts based on the outlier-removed data set and a plurality of previous model forecasts, wherein generating the set of outlier-bias-reduced model forecasts is performed by a computer processor; (4) determining a set of errors based on the set of new model forecasts and the set of actual values; repeating steps (1)-(4) using the new model forecasts in place of the set of model forecasts from the previous iteration until a performance termination criterion is met; and storing the set of outlier-bias-reduced model forecasts on a computer data medium.

[0011] Another embodiment includes a computer-implemented method for reducing outlier bias, comprising: determining a target variable for a facility; identifying a plurality of variables for the facility, the plurality of variables being direct variables for the facility that affect the target variable and a set of transformed variables for the facility, each of which is a function of at least one direct facility variable that affects the target variable; selecting an error metric comprising an absolute error and a relative error; obtaining a dataset comprising values ​​for the plurality of variables; selecting a set of actual values ​​for the target variable; selecting an initial set of model coefficients; generating a set of model predicted values ​​by applying the set of model coefficients to the dataset; determining a set of performance values ​​based on the set of model predicted values ​​and the set of actual values, the set of performance values ​​comprising a first standard error and a first coefficient of determination; (1) generating a set of errors based on the set of model predicted values ​​and the set of actual values ​​for the completed dataset, the relative error being calculated according to the formula: Relative Error m =((predicted value m - Actual value m ) / actual value m ) 2 (where "m" is a reference number), and the absolute error is calculated using the formula: Absolute Error m =(predicted value m - Actual value m ) 2(2) generating a set of error thresholds based on the completed set of errors for the completed dataset and the error criterion; (3) generating an outlier-removed dataset by removing data having error values ​​equal to or greater than the set of error thresholds, where the filtering is based on the dataset and the set of error thresholds; (4) generating a set of new coefficients based on the outlier-removed dataset and the set of old coefficients; (5) generating a set of outlier-bias-reduced model predictions based on the outlier-removed dataset and the set of new model coefficients by minimizing the error between the set of predicted values ​​and the set of actual values ​​using at least one of a linear optimization model and a nonlinear optimization model, where the model generating the predicted values ​​by a computer processor; (6) generating a set of updated performance values ​​based on the set of outlier bias-reduced model predicted values ​​and the set of actual values, the set of updated performance values ​​including a second standard error and a second coefficient of determination; repeating steps (1) through (6) using the set of new coefficients in place of the set of coefficients from the previous iteration until a performance termination criterion is met, the performance termination criterion including a standard error termination value and a coefficient of determination termination value, and satisfying the performance termination criterion includes the standard error termination value being greater than the difference between the first and second standard errors and the coefficient of determination termination value being greater than the difference between the first and second coefficients of determination; and storing the set of outlier bias reduction factors on a computer data medium.

[0012] Another embodiment includes a computer-implemented method for evaluating the viability of a dataset for use in developing a model, including the steps of providing a target dataset including a plurality of data values, generating a random target dataset based on the target dataset, selecting a set of bias metric values, generating an outlier-bias-reduced target dataset by a processor based on the dataset and the selected bias metric values, generating an outlier-bias-reduced random dataset by the processor based on the random dataset and the selected bias metric values, calculating a set of error values ​​for the outlier-bias-reduced dataset and the outlier-bias-reduced random dataset, calculating a set of correlation coefficients for the outlier-bias-reduced dataset and the outlier-bias-reduced random dataset, generating bias scale curves for the dataset and the random dataset based on the selected bias metric values ​​and corresponding error values ​​and correlation coefficients, and comparing the bias scale curve for the dataset with the bias scale curve for the random dataset. The outlier bias-reduced target dataset and the outlier bias-reduced random target dataset are generated using a dynamic outlier bias removal method. The random target dataset may comprise a plurality of randomly selected data values ​​expanded from a plurality of values ​​within the range of the plurality of data values. The set of error values ​​may include a set of standard errors. Here, the set of correlation coefficients includes a set of coefficient of determination values. Another embodiment further includes generating automated advice regarding the target dataset's feasibility of supporting the developed model, and vice versa, based on a comparison of the bias reference curve for the target dataset with the bias reference curve for the random target dataset. Advice can be generated based on analyst-selected parameters, such as a correlation coefficient threshold and / or an error threshold. Yet another embodiment further includes the following steps:generating, by a processor, an outlier bias-reduced actual data set based on the actual data set and selected bias metric values; generating, by the processor, an outlier bias-reduced random actual data set based on the random actual data set and selected bias metric values; generating, by the processor, a random data plot based on the outlier bias-reduced random target data set and the outlier bias-reduced random actual data for each selected bias metric; generating, by a processor, a real data plot based on the outlier bias-reduced target data set and the outlier bias-reduced actual target data for each selected bias metric; and comparing the random data plots with the real data plots corresponding to each selected bias metric.

[0013] A preferred embodiment includes a system comprising a server including a processor and a storage subsystem, a database including a dataset and stored by the storage subsystem, and a computer program stored by the storage subsystem. The computer program includes instructions that, when executed, cause the processor to: select a bias criterion, provide a set of model coefficients, select a set of target values, (1) generate a set of predicted values ​​for the dataset, (2) generate an error set for the dataset, (3) generate a set of error thresholds based on the error set and the bias criterion, (4) generate a censored dataset based on the error set and the set of error thresholds, (5) generate a set of new model coefficients, and (6) repeat steps (1)-(5) using the set of new model coefficients until a censored performance termination criterion is met. In a preferred embodiment, the set of predicted values ​​is generated based on the dataset and the set of model coefficients. In a preferred embodiment, the error set includes a set of absolute errors and a set of relative errors generated based on the set of predicted values ​​and the set of target values. In another embodiment, the error set includes values ​​calculated as the differences between the set of predicted values ​​and the set of target values. In another embodiment, generating the set of new coefficients further includes minimizing the set of errors between the set of predicted values ​​and the set of actual values. This can be accomplished using a linear or nonlinear optimization model. In a preferred embodiment, the truncated performance termination criterion is based on a standard error and a coefficient of determination.

[0014] Another embodiment of the invention includes a system comprising a server including a processor and a storage subsystem, a database including a dataset and stored by the storage subsystem, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to: select an error criterion, select a set of actual values, select an initial set of coefficients, generate a completed set of model predictions from the dataset and the initial set of coefficients, (1) generate a set of errors based on the model predictions and the set of actual values ​​for the completed dataset, (2) generate a set of error thresholds based on the completed set of errors for the completed dataset and the error criterion, (3) generate an outlier-removed dataset, the filtering of which is based on the completed dataset and the set of error thresholds, and (4) generate outlier-bias-reduced model predictions based on the outlier-removed dataset and the set of coefficients. (5) generating a set of new coefficients based on the outlier-removed data set and the set of old coefficients, wherein generating the set of new coefficients is performed by the computer processor; (6) generating a set of model performance values ​​based on the outlier-bias-reduced model predictors and the set of actual values, repeating steps (1)-(6) using the new set of coefficients in place of the set of coefficients from the previous iteration until a performance termination criterion is met; and storing the overall set of outlier-bias-reduced model predictors on a computer data medium.

[0015] Yet another embodiment includes a system including a server including a processor and a storage subsystem, a database stored by the storage subsystem including a target variable for a facility, a set of actual values ​​for the target variable, variables for the facility associated with the target variable, and a data set for the facility including values ​​for the variables, and a computer program stored by the storage subsystem, the computer program including instructions, when executed, that cause the processor to: select a bias criterion, select a set of model coefficients, (1) generate a set of predicted values ​​based on the data set and the set of model coefficients, (2) generate a set of censored model performance values ​​based on the set of predicted values ​​and the set of actual values, (3) generate an error set based on the set of predicted values ​​and the set of actual values ​​for the target variable, (4) generate a set of error thresholds based on the error set and the bias criterion, and (5) generate a censored model performance value based on the data set and the set of error thresholds. (6) generating a set of new model coefficients based on the censored data set and the set of model coefficients; (7) generating a set of new predicted values ​​based on the data set and the set of new model coefficients; (8) generating a set of new censored model performance values ​​based on the set of new predicted values ​​and the set of actual values, repeating steps (1) through (8) using the set of new coefficients until a censored performance termination criterion is met, and storing the set of new model predicted values ​​in the storage subsystem.

[0016] Another embodiment includes a system comprising a server including a processor and a storage subsystem, a database including a dataset for a facility and stored by the storage subsystem, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to: determine a target variable, identify a plurality of variables, the plurality of variables being direct variables for the facility that affect the target variable and a set of transformed variables for the facility, each of which is a function of at least one direct variable that affects the target variable, select an error metric including an absolute error and a relative error, select a set of actual values ​​for the target variable, select an initial set of coefficients, generate a set of model predictions from the dataset and the initial set of coefficients, and determine a set of errors based on the set of model predictions and the set of actual values, where the relative error is expressed by the formula: Relative Error m =((predicted value m - Actual value m ) / actual value m ) 2 (where "m" is a reference number), and the absolute error is calculated using the formula: Absolute Error m =(predicted value m - Actual value m ) 2determining a set of performance values ​​based on the set of model predictions and the set of actual values, the set of performance values ​​including a first standard error and a first coefficient of determination; (1) generating a set of errors based on the model predictions and the set of actual values; (2) generating a set of error thresholds based on the completed set of errors for the completed dataset and the error criterion; (3) generating an outlier-removed dataset by filtering data having error values ​​outside the set of error thresholds, the filtering based on the dataset and the set of error thresholds; (4) generating a set of new model predictions based on the outlier-removed dataset and the set of coefficients by minimizing an error between the set of predictions and the set of actual values ​​using at least one of a linear optimization model and a nonlinear optimization model, the outlier bias being minimized. (5) generating a set of new coefficients based on the outlier-removed data set and the set of old coefficients, where generating the set of new coefficients is performed by the computer processor; (6) generating a set of performance values ​​based on the set of new model forecasts and the set of actual values, where the set of model performance values ​​includes a second standard error and a second coefficient of determination; repeating steps (1)-(6) using the set of new coefficients in place of the set of coefficients from a previous iteration until a performance termination criterion is met, where the performance termination criterion includes a standard error termination value and a coefficient of determination, and where satisfying the performance termination criterion includes the standard error termination value being greater than a difference between the first and second standard errors and the coefficient of determination termination value being greater than a difference between the first and second coefficients of determination; and storing the set of new model forecasts on a computer data medium.

[0017] Another embodiment of the invention includes a system comprising a server including a processor and a storage subsystem, a database including a data set and stored by the storage subsystem, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to: (2) generating an outlier-removed dataset, the filtering of which is based on the dataset and the set of error thresholds; (3) generating a set of outlier-bias-reduced model forecasts based on the outlier-removed dataset and the completed set of model forecasts, wherein generating the set of outlier-bias-reduced model forecasts is performed by a computer processor; (4) determining a set of errors based on the set of outlier-bias-reduced model forecasts and the corresponding set of actual values; repeating steps (1)-(4) using the set of outlier-bias-reduced model forecasts in place of the set of model forecasts until a performance termination criterion is met; and storing the set of outlier bias reduction factors on a computer data medium.

[0018] Another embodiment of the invention includes a system comprising a server including a processor and a storage subsystem, a database including a data set and stored by the storage subsystem, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to: the set of model coefficients is applied to the data set to generate a set of model predictions; determining a set of performance values ​​based on the set of model predictions and the set of actual values, the set of performance values ​​including a first standard error and a first coefficient of determination; (1) determining a set of errors based on the set of model predictions and the set of actual values, the relative error being calculated using the formula: Relative Error k =((predicted value k - Actual value k ) / actual value k ) 2 (where "k" is a reference number), and the absolute error is calculated using the formula: Absolute Error k =(predicted value k - Actual value k ) 2(2) determining a set of error thresholds based on the set of errors for the completed dataset and the error criterion; (3) generating an outlier-removed dataset by removing data having error values ​​equal to or greater than the error thresholds, the filtering being based on the dataset and the set of error thresholds; (4) generating a set of new coefficients based on the outlier-removed dataset and the set of old coefficients; (5) minimizing an error between the set of predicted values ​​and the set of actual values ​​using at least one of a linear optimization model and a nonlinear optimization model and generating a set of outlier-bias-reduced model predictions based on the outlier-removed dataset and the set of coefficients. (5) generating a set of updated performance values ​​based on the set of outlier bias-reduced model predicted values ​​and the set of actual values, the set of updated performance values ​​including a second standard error and a second coefficient of determination; repeating steps (1) through (5) using the set of new coefficients in place of the set of coefficients from a previous iteration until a performance termination criterion is met, the performance termination criterion including a standard error termination value and a coefficient of determination termination value, and satisfying the performance termination criterion includes the standard error termination value being greater than the difference between the first and second standard errors and the coefficient of determination termination value being greater than the difference between the first and second coefficients of determination; and storing the set of outlier bias reduction factors on a computer data medium.

[0019] Yet another embodiment includes a system for assessing the viability of a dataset for use in developing a model, the system including a server including a processor and a storage subsystem, a database including a target dataset including a plurality of model predictions and stored by the storage subsystem, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to: The method includes generating a random target data set, selecting a set of bias metric values, generating a plurality of outlier bias-reduced data sets based on the target data set and the selected bias metric values, generating a single outlier bias-reduced random target data set based on the random target data set and the selected bias metric values, calculating a set of error values ​​for the outlier bias-reduced target data set and the outlier bias-reduced random target data set, calculating a set of correlation coefficients for the outlier bias-reduced target data set and the outlier bias-reduced random target data set, generating a plurality of bias basis curves for the target data set and the random target data set based on the corresponding error values ​​and correlation coefficients for the selected bias metric, and comparing the bias basis curve for the target data set with the bias basis curve for the random target data set. The processor generates the outlier bias-reduced target data set and the outlier bias-reduced random target data set using a dynamic outlier bias removal method. The random target data set may consist of a plurality of randomly selected data values ​​extracted from a plurality of values ​​within the range of the plurality of data values. Additionally, the set of error values ​​may include a set of standard errors. The set of correlation coefficients may include a set of coefficient of determination values. In another embodiment, the program further includes instructions that, when executed, cause the processor to generate automated advice based on a comparison of the bias reference curve for the target dataset versus the bias reference curve for the random target dataset.The advice can be generated based on analyst-selected parameters, such as a correlation coefficient threshold and / or an error threshold. In yet another embodiment, the system's database further includes an actual data set comprising a plurality of actual data values ​​corresponding to the model predictions, and the program further includes instructions that, when executed, cause the processor to: generate a random actual data set based on the actual data set, generate an outlier-bias-reduced actual data set based on the actual data set and selected bias metric values, generate an outlier-bias-reduced random actual data set based on the random actual data set and selected bias metric values, generate a random data plot based on the outlier-bias-reduced random target data set and the outlier-bias-reduced random actual data for each selected bias metric, generate a realistic data plot based on the outlier-bias-reduced target data set and the outlier-bias-reduced actual target data set for each selected bias metric, and compare the random data plot with the realistic data plot corresponding to each selected bias metric.

[0020] Another embodiment includes a system for reducing outlier bias in a measured target variable for a facility, the system including a computer unit for processing a data set, the computer unit including a processor and storage subsystem, an input unit for inputting the data set to be processed, the input unit including a measurement device for measuring a given target variable and providing a corresponding data set, an output unit for outputting a processed data set, and a computer program stored by the storage subsystem, the computer program including instructions that, when executed, cause the processor to perform the steps of: (3) generating a set of error thresholds based on the error set and the bias criterion; (4) generating a truncated dataset based on the error set and the set of error thresholds; (5) generating a set of new model coefficients; and (6) repeating steps (1) through (5) using the set of new model coefficients until a censored performance termination criterion is met.

[0021] Yet another embodiment includes a system for reducing outlier bias in a target variable measured for a financial instrument such as an equity security (e.g., common stock) or a derivative contract (e.g., forward, futures, option, swap, etc.). The system includes a computer unit for processing a data set, the computer unit including a processor and a storage subsystem, an input unit for receiving the data set to be processed, the input unit including a storage device for storing data on the target variable (e.g., stock price) and providing a corresponding data set, an output unit for outputting the processed data set, and a computer program stored by the storage subsystem. The computer program includes instructions that, when executed, cause the processor to perform the steps of: (3) generating a set of error thresholds based on the error set and the bias criterion; (4) generating a truncated data set based on the error set and the set of error thresholds; (5) generating a set of new model coefficients; and (6) repeating steps (1) through (5) using the set of new model coefficients until a truncated performance termination criterion is met. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a flowchart illustrating one embodiment of a method for identifying and removing data outliers. [Figure 2]1 is a flowchart illustrating one embodiment of a method for identifying and removing data outliers for data quality operations. [Figure 3] 1 is a flow chart illustrating one embodiment of a method for identifying and removing data outliers for data validation. [Figure 4] 1 is an exemplary node for implementing the method of the present invention; [Figure 5] 1 is an exemplary graph for quantitative evaluation of a data set. [Figure 6] 6A and 6B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the randomized and realistic data sets relative to the entire data set, respectively. [Figure 7] 7A and 7B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the unselected data set and the realistic data set, respectively, after removing 30% of the data as outliers. [Figure 8] 8A and 8B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the unselected data set and the realistic data set, respectively, after removing 50% of the data as outliers. [Figure 9] 1 illustrates an example system used to reduce outlier bias in a target variable measured for a facility. DETAILED DESCRIPTION OF THE INVENTION

[0023] The following disclosure provides many different embodiments or examples that implement different features of the systems and methods for accessing and managing structured content. Specific examples of components, processes, and implementations are described to help clarify the invention. These are merely examples and are not intended to limit the invention from what is set forth in the claims. Well-known elements are presented without detailed descriptions so as not to obscure the preferred embodiments of the invention with unnecessary detail. For the most part, details unnecessary to gain a complete understanding of the preferred embodiments of the invention are omitted, insofar as such details are within the skill of those skilled in the art.

[0024] A mathematical description of one embodiment of dynamic outlier bias reduction is as follows.

number

number

[0025]

number

[0026] Another mathematical description of one embodiment of dynamic outlier bias reduction is as follows:

number

number

[0027]

number

[0028] After each iteration in which new model coefficients are calculated from the current truncated data set, the removed data from the previous iteration plus the current censored data are recombined. This combination encompasses all data values ​​in the completed data set. The current model coefficients are then applied to the completed data set to calculate a completed set of predicted values. Absolute and relative errors are calculated for the completed set of predicted values, and a new bias-based percentile threshold is calculated. A new truncated data set is created by removing all data values ​​with absolute or relative errors greater than the threshold, and then a nonlinear optimization model is applied to the newly censored data set to calculate new model coefficients. This process allows the possibility of including all data values ​​in the model data set to be examined at each iteration. When the model coefficients converge to values ​​that best fit the data, some data values ​​excluded in a previous iteration may be included in subsequent iterations.

[0029] In one embodiment, variations in greenhouse gas emissions can result in overestimation or underestimation of emission results, leading to bias in model predictions. These non-industrial influences, such as errors in environmental conditions and calculation procedures, can cause results for a particular facility to differ radically from similar facilities unless the bias in the model predictions is removed. Bias in model predictions also exists due to unique operating conditions.

[0030] If the analyst is convinced that a facility's calculations are erroneous or have unique extenuating characteristics, bias can be manually removed by simply removing the facility's data from the calculations. However, when measuring facility performance from many different companies, regions, and countries, precise a priori knowledge of the data details is impractical. Therefore, any analyst-based data removal procedure has the potential to add undocumented and unsupported bias to the data.

[0031] In one embodiment, dynamic outlier bias reduction is applied to a procedure that uses data and a predetermined overall error criterion to determine statistical outliers to be removed from the model coefficient calculation. This is a data-driven process that uses a global error criterion driven by the data to identify outliers, for example, using a percentile function. The use of dynamic outlier bias reduction is not limited to reducing bias in model predictions; its use in this embodiment is merely illustrative and exemplary. Dynamic outlier bias reduction is also used, for example, to remove outliers from any statistical data set. This includes, but is not limited to, use in, for example, arithmetic means, linear regressions, and trend line calculations. Outlier facilities are still ranked in the calculation results, but the outliers are not used in the filtered data set applied to calculate model coefficients or statistical results.

[0032] A commonly used standard procedure for removing outliers is to simply calculate the standard deviation (σ) of a data set and define all data that fall outside a 2σ interval from the mean as, for example, an outlier. This procedure has statistical assumptions that are generally untestable in practice. A description of the dynamic outlier bias reduction method applied in one embodiment of the present invention is summarized in Figure 1 and uses both relative error and absolute error. For example, for facility "m": Relative Error m =((predicted value m - Actual value m ) / actual value m ) 2 (1) Absolute Error m =(predicted value m - Actual value m ) 2 (2) This becomes:

[0033] In step 110, the analyst specifies error threshold criteria that define outliers to be removed from the calculation. For example, an 80th percentile value for the relative and absolute error may be established using a percentile operation as the error function. This means that calculations include data values ​​below the 80th percentile value for relative error and the 80th percentile value for absolute error, and the remaining values ​​are removed or considered outliers. In this example, for a data value to avoid removal, the data value must be below the 80th percentile value for both relative and absolute error. However, because the percentile thresholds for both relative and absolute error can vary independently, in other embodiments, only one percentile threshold is used.

[0034] In step 120, the model standard error and the coefficient of determination (r 2 ) percent change criteria are specified. While the values ​​of these statistics will vary from model to model, the percent change in the previous iteration procedure can be preset, for example, 5 percent. These values ​​can be used to terminate the iteration procedure. Another termination criterion could be a simple number of iterations.

[0035] In step 130, an optimization calculation is performed to generate model coefficients and predictions for each facility.

[0036] In step 140, both the relative and absolute errors for all sites are calculated using equations (1) and (2).

[0037] In step 150, an error function with the threshold criteria specified in step 110 is applied to the data calculated in step 140 to determine an outlier threshold.

[0038] In step 160, the data is filtered to include only facilities whose relative error, absolute error, or both, depending on the configuration selected, are less than the error threshold calculated in step 150.

[0039] In step 170, an optimization calculation is performed using the outlier-removed data set.

[0040] In step 180, the standard error and r 2 The percent change in is compared to the criteria identified in step 120. If the percent change is greater than the criteria, the process is repeated by returning to step 140. Otherwise, the iterative procedure ends in step 190, and the resulting model calculated from this dynamic outlier bias reduction criteria is finalized. The model results are applied to all sites for that current iteration, regardless of the status of previously removed or accepted data.

[0041] In another embodiment, the process begins with the selection of predetermined iterative parameters, specifically: (1) absolute error and relative error percentile values, one or both of which are used in the iterative process; (2) the coefficient of determination (r 2 (also known as the mean squared error), and (3) the standard error improvement.

[0042] The process begins with an original data set, a set of actual data, and either at least one coefficient or factor used to calculate predicted values ​​based on the original data set. The coefficient or set of coefficients is applied to the original data set to produce a set of predicted values. The set of coefficients includes, but is not limited to, scalar, exponential, parametric, and periodic functions. The set of predicted data is then compared to the set of actual data. A standard error and coefficient of determination are calculated based on the difference between the predicted and actual data. The absolute and relative errors associated with each data point are used to remove data outliers based on user-selected absolute and relative error percentile values. Ranking of the data is not required, as all data outside the ranges associated with the percentile values ​​for absolute and / or relative error are removed from the original data set. The use of absolute and relative errors to filter data is exemplary and for illustrative purposes only, as the method can be performed on absolute or relative errors alone or on other functions.

[0043] The data associated with absolute and relative errors within a user-selected percentile range is the outlier-removed dataset, with each iteration of the process having its own filtered dataset. This first outlier-removed dataset is used to determine predicted values ​​that are compared to actual values. At least one coefficient is determined by optimizing the error, and the coefficient is then used to generate predicted values ​​based on the first outlier-removed dataset. The outlier-bias reduced coefficients serve as a mechanism by which knowledge is transferred from one iteration to the next.

[0044] After the first outlier-removed data set is created, the standard error and coefficient of determination are calculated and compared to the standard error and coefficient of determination of the original data set. If both the standard error difference and the coefficient of determination difference are less than their respective improvement values, the process stops. However, if at least one of the improvement criteria is not met, the process continues for another iteration. The use of standard error and coefficient of determination as checks on the iterative process is illustrative and exemplary only, as such checks could be performed using only standard error or only coefficient of determination, different statistical checks, or other performance termination criteria (such as the number of iterations).

[0045] If the first iteration fails to meet the improvement criteria, a second iteration begins by applying the first outlier-bias-reduced data coefficients to the original data to determine a new set of predicted values. In this case, the original data is again processed, and absolute and relative errors for the data points, as well as standard errors and coefficient of determination values ​​for the original data set, are established while the coefficients of the first outlier-removed data set are used. The data is then filtered to form a second outlier-removed data set, and coefficients based on the second outlier-removed data set are determined.

[0046] However, the second outlier-removed data set is not necessarily a subset of the first outlier-removed data set, and is associated with a second set of outlier-bias-reduced model coefficients, a second standard error, and a second coefficient of determination. Once these values ​​are determined, the second standard error is compared to the first standard error, and the second coefficient of determination is compared to the first coefficient of determination.

[0047] If the improvement (in standard error and coefficient of determination) exceeds the difference in these parameters, the process ends. Otherwise, another iteration begins by processing the original data again. This time, the second outlier-bias-reduced coefficients are used to process the original data set and generate a new set of predicted values. Filtering based on user-selected percentile values ​​for absolute and relative error creates a third outlier-removed data set that is optimized to determine a set of third outlier-bias-reduced coefficients. The process continues until error improvement or other termination criteria (such as convergence criteria or a specified number of iterations) are met.

[0048] The output of this process is a set of coefficients or model parameters, where a coefficient or model parameter is a mathematical value (or set of values) such as, but not limited to, a model prediction for comparing data, slope and intercept values ​​of a linear equation, an exponent, or coefficients of a polynomial. The output of Dynamic Outlier Bias Reduction is not an output value in its own right, but rather the coefficients that modify the data to determine the output value.

[0049] In another embodiment illustrated in FIG. 2, dynamic outlier bias reduction is applied as a data quality method to evaluate the consistency and accuracy of data to ensure that the data is suitable for a particular use. The method does not involve an iterative procedure for data quality operations. Other data quality methods can also be used in conjunction with dynamic outlier bias reduction during this process. The method is applied to the arithmetic mean calculation of a given data set. A data quality criterion, for example, is that consecutive data values ​​fall within the same range; that is, any values ​​that are too far apart constitute poor quality data. In this case, the error term is composed of consecutive values ​​of the function, and dynamic outlier bias reduction is applied to these error values.

[0050] In step 210, the initial data is listed in any order.

[0051] Step 220 configures a function or operation to be performed on the data set. In this example embodiment, the function or operation is an ascending ranking of the data followed by a successive arithmetic mean calculation where each line corresponds to the average of all the data above that line.

[0052] Step 230 uses the successive values ​​from the result of step 220 to calculate relative and absolute errors from the data.

[0053] Step 240 allows the analyst to input a desired outlier removal error metric (%). The quality metric value is the resulting value from the error calculation in step 230 based on the data in step 220.

[0054] Step 250 shows the data quality outlier filtered data set. If the relative and absolute errors exceed the specified error criteria given in step 240, the specified values ​​are removed.

[0055] Step 260 depicts the comparison of the arithmetic mean of the completed data set versus the outlier-removed data set. The analyst is the final step in determining whether the identified outlier-removed data components are in fact of inferior quality in any applied mathematical or statistical calculations. The dynamic outlier bias reduction system and method eliminates direct data removal by the analyst, and best practice guidelines prompt the analyst to review and check results for implementation validity.

[0056] In another embodiment, illustrated in FIG. 3, dynamic outlier bias reduction is applied as a data validation method to test the reasonable accuracy of a data set to determine whether the data is appropriate for a particular use. This method does not involve an iterative procedure for data validation operations. In this example, dynamic outlier bias reduction is applied to calculate the Pearson correlation coefficient between two data sets. The Pearson correlation coefficient is highly sensitive to values ​​that are different relative to other data points in the data set. Validating a data set against this statistic is important to ensure that the results are representative of what the majority of the data suggests, without the influence of extreme values. The data validation process in this example ensures that consecutive data values ​​fall within a specified range. That is, any values ​​that are too far apart (e.g., values ​​outside the specified range) indicate poor quality data. This is achieved by constructing error terms for consecutive values ​​of the function. Dynamic outlier bias reduction is applied to these error values, resulting in the outlier-removed data set being validated data.

[0057] In step 310, the pairs of data are listed in any order.

[0058] Step 320 calculates the relative and absolute errors for each aligned pair in the data set.

[0059] Step 330 allows the analyst to input desired data validation criteria. In this example, both relative and absolute error thresholds of 90% are selected. The quality metric values ​​entered in step 330 are the resulting absolute and relative error percentile values ​​for the data presented in step 320.

[0060] Step 340 illustrates the outlier removal process, in which potentially invalid data is removed from the data set using a criterion that both relative and absolute error values ​​exceed values ​​corresponding to the user-selected percentile values ​​entered in step 330. In practice, other error criteria can be used, so that when multiple criteria are applied as shown in this example, any combination of error values ​​can be applied to determine the outlier removal rules.

[0061] Step 350 calculates statistical results for the validated and raw data values, in this case the Pearson correlation coefficient. These results are then validated by an analyst.

[0062] In another embodiment, dynamic outlier bias reduction is used to validate the entire dataset. Standard error improvement values, coefficient of determination improvement values, and absolute and relative error thresholds are selected, and then the dataset is filtered according to the error criteria. Even if the original dataset is of high quality, there will still be some data with error values ​​outside the absolute and relative error thresholds. Therefore, it is important to determine whether any removal of data is necessary. If the outlier-removed dataset passes the standard error improvement and coefficient of determination improvement criteria after the first iteration, the original dataset is validated because the filtered dataset produces standard errors and coefficients of determination that are too small to be considered significant (e.g., less than the selected improvement values).

[0063] In another embodiment, dynamic outlier bias reduction is used to provide insight into how iterations of data outlier removal affect the calculation. Graphs or data tables are provided so the user can observe the progress of the data outlier removal calculation as each iteration occurs. This step-by-step approach allows the analyst to observe unique characteristics of the calculation that can add value and knowledge to the results. For example, the speed and nature of convergence indicate the impact of dynamic outlier bias reduction on calculating representativeness factors for multidimensional datasets.

[0064] As an example, consider a linear regression calculation on a poor quality dataset of 87 records. The equation to be regressed is of the form y=mx+b. Table 1 shows the results of the iterative process for five iterations. It is noteworthy that convergence is achieved in three iterations, using a 95% relative and absolute error criterion. The change in the regression coefficients can be observed. The dynamic outlier bias reduction method reduced the calculation dataset based on 79 records. A relatively low coefficient of determination (r 2 =39%) is r 2 This indicates that a lower (<95%) criterion needs to be tested to consider the additional effect of outlier removal on the statistics and on the calculated regression coefficients. [Table 1]

[0065] Table 2 shows the results of applying dynamic outlier bias reduction using 80% relative and absolute error criteria. Notably, a 15 percentage point change in the outlier error criterion (from 95% to 80%) resulted in an additional 35% reduction in the allowed data (from 79 to 51 records). 2 Analysts can use the graphical depiction of the change in the regression line along with the outlier-removed data and numerical results in Tables 1 and 2 in the analytical process to communicate the outlier-removed results to a wide audience and to provide insight into the effect of data variability on the analytical results. [Table 2]

[0066] As illustrated in FIG. 4 , one embodiment of a system used to perform the method includes a computer system. The hardware consists of a processor 410 containing sufficient system memory 420 to perform the necessary numerical calculations. The processor 410 executes a computer program residing in the system memory 420 to perform the method. A video and storage controller 430 is used to enable operation of a display 440. The system includes various data storage devices for data input, such as a floppy disk unit 450, an internal / external disk drive 460, an internal CD / DVD 470, a tape unit 480, and other types of electronic storage media 490. The data storage devices described above are illustrative and exemplary only. These storage media are used to input data sets and outlier removal criteria into the system, store outlier-removed data sets, store calculation factors, and store system-generated trend lines and trend line iteration graphs. Calculations can be applied in a statistical software package or can be performed from data entered in a spreadsheet format, for example, using Microsoft Excel. Calculations are performed using customized software programs designed for enterprise-specific system implementations or using commercially available software compatible with database and spreadsheet programs such as Excel. The system may also interface with proprietary or public external storage media 300 that interfaces with other databases to provide data for use with the dynamic outlier bias reduction system and method calculations. Output devices may be telecommunications devices 510 that transmit system-generated graphs and reports, such as calculation worksheets, via an intranet or the Internet to administrative personnel, printers 520, electronic storage media similar to those described above as input devices 450, 460, 470, 480, 490, and proprietary storage databases 530. These output devices are merely illustrative and exemplary.

[0067] As illustrated in Figures 5, 6A, 6B, 7A, 7B, 8A, and 8B, in one embodiment, dynamic outlier bias reduction can be used to quantitatively and qualitatively assess the quality of a dataset. This is based on comparing the error and correlation of the data values ​​of a dataset with the error and correlation of a benchmark dataset consisting of random data values ​​drawn from within an appropriate range. In one embodiment, the error can be specified to be the standard error of the dataset. The correlation can be calculated using the coefficient of determination (r 2 ) In other embodiments, the correlation can be specified to be Kendall's rank correlation coefficient, commonly referred to as Kendall's tau (τ) coefficient. In yet other embodiments, the correlation can be specified to be Spearman's rank correlation coefficient or Spearman's ρ (rho) coefficient. As described above, dynamic outlier bias reduction is used to systematically remove data values ​​identified as outliers. It is not intended to describe the representativeness of the underlying model or process. Typically, outliers are associated with a relatively small number of data values. However, in reality, data sets can be insidiously contaminated with spurious values ​​or random noise. The graphical illustrations in Figures 5, 6A, 6B, 7A, 7B, 8A, and 8B illustrate how dynamic outlier bias reduction systems and methods can be applied to identify situations where the underlying model is not supported by the data. Outlier reduction is performed by removing data values ​​where the calculated relative and / or absolute error between the predictive model and the actual data values ​​is greater than a percentile-based bias criterion, such as 80%. This means that a data value is removed if its relative or absolute error percentile value is greater than the percentile threshold associated with the 80th percentile (80% of the data values ​​have an error less than this value).

[0068] As illustrated in Figure 5, both a realistic model development dataset developed within a real dataset and a dataset of random values ​​are contrasted. In practice, analysts typically have no prior knowledge of any dataset contamination, so such understanding must be based on observing replicate results from several model runs using the dynamic outlier bias reduction system and method. Figure 5 illustrates the results of an example model development run for both datasets. The standard error, a measure of the amount of error unexplained by the model, is used to calculate the coefficient of determination (%) or r, which represents how much data variability is explained by the model. 2 The percentile value next to each point represents the bias criterion. For example, 90% indicates that data values ​​for relative or absolute error values ​​greater than the 90th percentile are removed as outliers from the model. This corresponds to removing 10% of the data values ​​with the highest error per iteration.

[0069] As illustrated in Figure 5, for both the random and realistic data set models, increasing the bias criterion reduces the error; that is, the standard error and coefficient of determination improve for both data sets. However, the standard error for the random data set is two to three times larger than the realistic model data set. An analyst can use an 80% coefficient of determination requirement as an acceptable level of precision for determining model parameters, for example. Figure 5 shows that an 80% r is obtained at a 70% bias criterion for the random data set and at approximately an 85% bias criterion for the realistic data set. 2is achieved. However, the corresponding standard error for the random dataset is more than twice as large as that for the realistic dataset. That is, by systematically performing the model dataset analysis with different bias criteria and repeating the calculations with a representative pseudo dataset and plotting the results as shown in FIG. 5, an analyst can assess the acceptable bias criteria for the dataset (i.e., the acceptable percentage of data values ​​removed) and, therefore, the overall dataset quality. Furthermore, such systematic model dataset analysis can be used to automatically provide advice regarding the viability of a dataset used in model development based on a configurable set of parameters. For example, in one embodiment in which a model is developed using dynamic outlier bias removal for the dataset, the error and correlation coefficient values ​​for the model dataset and for the representative pseudo dataset calculated under different bias criteria can be used to automatically provide advice regarding the viability of the dataset in supporting the developed model, and, essentially, the viability of the deployed model in supporting the dataset.

[0070] Observing the behavior of these model performance values ​​for several cases provides a quantitative basis for determining whether the data values ​​are representative of the process being modeled, as illustrated in Figure 5. For example, referring to Figure 5, the standard error for a realistic data set at a 100% bias criterion (i.e., no bias reduction) corresponds approximately to the standard error for a random data set at a 65% bias criterion (i.e., the highest error is 35% of the removed data value). Such a finding supports the conclusion that the data are not contaminated.

[0071] In addition to the quantitative analysis described above, facilitated by the exemplary graph of Figure 5, dynamic outlier bias reduction can be utilized in an equally, if not more powerful, subjective procedure to aid in the quality assessment of a data set by plotting model predicted values ​​against actual target values ​​given the data, for both outliers and included outcomes.

[0072] Figures 6A and 6B illustrate such plots for the 100% point for both the realistic and random curves in Figure 5. The large scatter in Figure 6A is consistent with the model's inability to fit arbitrary target values ​​and, consequently, intentional randomness. Figure 6B is consistent with and typical of a collection of real data, with model predictions and actual values ​​clustering around a line where the model-predicted value equals the actual target value (hereafter referred to as the actual=predicted line).

[0073] Figures 7A and 7B illustrate results from 70% of the points in Figure 5 (i.e., 30% of the data are removed as outliers). Although outlier bias reduction in Figures 7A and 7B is shown to remove points furthest from the actual = predicted line, the large variation in model accuracy between Figures 7A and 7B illustrates the process by which this data set is modeled.

[0074] Figures 8A and 8B show results from the 50% point in Figure 5 (i.e., 50% of the data removed as outliers). In this case, approximately half of the data are identified as outliers; even with this much variability removed from the dataset, the model in Figure 8A still does not accurately describe a random dataset. The general variability around the Actual = Predicted line is similar to that in Figures 6A and 7A when considering the removed data in each case. Figure 8B shows that when 50% of the variability was removed, the model was able to generate predicted results that closely matched the actual data. In addition to analyzing the performance metrics shown in Figure 5, analyzing these types of visual plots can be used by analysts to assess the quality of actual datasets in practice for model development. Figures 5, 6A, 6B, 7A, 7B, 8A, and 8B illustrate example visual plots. Here, the analysis is based on performance metric trends corresponding to various bias metric values. In other embodiments, the analysis can be based on other variables corresponding to bias metric values, such as model coefficient trends corresponding to various bias metrics selected by the analyst.

[0075] Various embodiments include a system for reducing outlier bias in a measured target variable for a facility. FIG. 9 illustrates an example of such an embodiment. The system illustrated in FIG. 9 includes a computer unit 1012 capable of processing a dataset, such as a dataset containing various performance measures for an industrial facility. The computer unit 1012 includes a processor 1014 and a storage subsystem 1016, on which a computer program implements the dynamic outlier bias removal method disclosed herein. The system 1010 includes an input unit 1018. The input unit 1018 further includes a measurement device 1020 that measures a given target variable and provides a corresponding dataset. The measurement device 1020 can be configured to measure any target variable of interest, such as the number of parts leaving an industrial plant facility per unit time or the volume of refined material produced by a refinery facility per unit time. Alternatively, multiple target variables can be measured simultaneously. In the illustrated embodiment, the measurement device 1020 includes a sensor 1022. Those skilled in the art will appreciate that the scope of the present invention includes various sensors used to measure various physical attributes of materials and / or components produced by or used in industrial facilities. Examples include sensors capable of detecting and quantifying chemicals, such as greenhouse gas emissions. In addition, those skilled in the art will appreciate that measuring a target variable of interest includes any means of collecting, receiving, measuring, storing, and processing data. Target variables, datasets, and data can include all types of data of interest, including, but not limited to, industrial process data, computer system data, financial data, economic data, stock, bond, and futures data, internet search data, security data, human identification data such as voice, cloud data, big data, insurance data, etc. The scope and teachings of the present disclosure and the present invention are not limited to such types of target variables, datasets, or data. Those skilled in the art will also appreciate that sensors and measurement devices can be or include computers, computer systems, and processors. Furthermore, system 1010 includes an output unit 1024 capable of outputting processed data.Output devices include a monitor, printer, or transmission device (not shown).

[0076] In one embodiment, the system 1010 activates the sensor 1022. The sensor 1022 then detects and quantifies a given compound, such as carbon dioxide. The detection and quantification can be performed continuously or in discrete time steps. As each measurement is completed, a data set is generated, stored in the storage subsystem 1016, and input into the computer unit 1012. The data set is processed by a dynamic outlier bias removal computer program stored in the storage subsystem 1016 and is truncated according to various embodiments of the methods disclosed herein. Once the computer program has completed data processing, the processed data is output by the output unit 1024. In embodiments where the output unit 1024 is a monitor or printer, the results are visualized in a diagram. In one embodiment where the output unit 1024 includes a transmission device, the processed data is sent to a central database or control center, where the data is further processed (not shown). Thus, the system according to various disclosed embodiments provides a powerful tool for comparing different facilities within a company or technology area in an automated manner where outlier bias is reduced.

[0077] In a preferred embodiment, the measurement device 1020 includes one or more sensors that detect and quantify chemicals. Due to global warming, greenhouse gas emissions from facilities are becoming an increasingly important target variable. Facilities that emit smaller amounts of greenhouse gases will rank better than those that emit larger amounts, although the latter will have better overall productivity. Greenhouse gases include, for example, carbon dioxide (CO), ozone (O), water vapor (H), hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), chlorofluorocarbons (CFCs), sulfur hexafluoride (SF), methane (CH), nitrous oxide (N), carbon monoxide (CO), nitrogen oxides (NO), and chlorofluorocarbons (CFCs). x) and non-methane volatile organic compounds (NMVOCs). Automated detection and quantification of these compounds can be used to develop industry standards for predetermined allowable emissions of greenhouse gases. However, application of dynamic outlier bias removal results in the removal of outliers caused by abnormal production situations, such as operational errors or even accidents. Thus, use of various embodiments disclosed herein results in the development of accurate and meaningful standards. Once an industry standard is developed, the system is used to compare emissions against the standard.

[0078] Those skilled in the art will further appreciate that the scope of the present invention includes application of various disclosed embodiments to reduce outlier bias of a target variable in a target variable associated with a financial instrument, such as an equity security (e.g., common stock) or a derivative contract (e.g., a forward, futures, option, swap, etc.). For example, in one embodiment, the system 1010 includes an input unit 1018 that receives data related to a financial instrument, such as common stock, and provides a corresponding data set. The target variable may be a stock price. Furthermore, variables related to the target variable may be determined using various well-known methods for valuing financial instruments, such as discounted cash flow analysis. Such related variables may include associated dividends, retained earnings, or cash flows, earnings per share, price-to-earnings ratios or growth rates, etc. Once a database of target values ​​and related variable values ​​is formed, various embodiments of the dynamic outlier bias removal disclosed herein may be applied to the database to obtain an accurate model for valuing the financial instrument.

[0079] The foregoing disclosure and description of the preferred embodiment of the present invention is illustrative and exemplary, and it will be understood by those skilled in the art that various changes can be made in the details of the illustrated systems and methods without departing from the scope of the invention.

Claims

1. 1. A system for reducing outlier bias in a target variable measured for a facility, comprising: a computer unit including a processor and a storage subsystem for processing the data set; an input unit for inputting the data set to be processed, the input unit including a measurement device for measuring a target variable for the facility and providing a data set corresponding thereto; an output unit that outputs the processed dataset; a computer program stored on the storage subsystem that, when executed, causes the processor to: selecting a target variable for the facility; selecting a set of actual values ​​of the target variable; identifying a plurality of variables related to the target variable for the facility; obtaining a dataset for the facility, the dataset including a plurality of values ​​for the plurality of variables; selecting a bias criterion; selecting a set of model coefficients; (1) generating a set of predictions for the dataset; (2) generating an error set for the data set; (3) generating a set of error thresholds based on the error set and the bias criterion; (4) generating a truncated data set based on the error set and the set of error thresholds; (5) generating a set of new model coefficients; (6) repeating steps (1) through (5) using the set of new model coefficients until a truncated performance termination criterion is met; and a computer program including instructions to execute the A system including:

2. The system of claim 1 , wherein the measurement device comprises one or more sensors.

3. The system of claim 2 , wherein the sensor detects and quantifies chemicals to the facility.

4. 1. A system for reducing outlier bias in a target variable measured for a financial instrument, comprising: a computer unit for processing a data set, the computer unit including a processor and a storage subsystem; an output unit that outputs the processed dataset; a computer program stored on the storage subsystem that, when executed, causes the processor to: selecting a target variable for the financial instrument; selecting a set of actual values ​​of the target variable; identifying a plurality of variables for the financial instrument that are related to the target variable; obtaining a data set for the financial instrument, the data set including a plurality of values ​​for the plurality of variables; selecting a bias criterion; selecting a set of model coefficients; (1) generating a set of predictions for the dataset; (2) generating an error set for the data set; (3) generating a set of error thresholds based on the error set and the bias criterion; (4) generating a truncated data set based on the error set and the set of error thresholds; (5) generating a set of new model coefficients; (6) repeating steps (1) through (5) using the set of new model coefficients until a censored performance termination criterion is met; and and a computer program including instructions to execute the A system including:

5. The financial instrument is a common stock; The system of claim 4 , wherein the target variable is the price of the common stock.

6. 6. The system of claim 5, wherein the plurality of variables for the financial instrument are related to the target variable and include at least one of dividends, retained earnings, cash flow, earnings per share, price-to-earnings ratio, and growth rate.

Citation Information

Patent Citations

  • Adapting method for engine control parameter and its system

    JP2004068729A

  • Integrated circuit device abnormality detection apparatus, method and program

    JP2008166644A

  • Network performance prediction system, network performance prediction method, and program

    JP2009253362A

  • Working hour estimation device, method, and program

    JP2010250674A

  • Systems and computer-implemented methods for measuring industrial process performance for industrial process facilities and computer-implemented methods for reducing outlier bias in measured target variables for financial instruments

    JP7244610B2