System and computer-implemented method for measuring industrial process performance for an industrial process facility
A computer-implemented method iteratively reduces outlier bias in industrial data models by using error thresholding and coefficient recalibration, addressing subjective outlier removal in data analysis and improving model accuracy.
Patent Information
- Application Number
- JP2023036170
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-02-20
- Filing Date
- 2023-03-09
- Publication Date
- 2025-10-30
- Estimated Expiration
- 2034-02-20
AI Technical Summary
Existing data analysis methods struggle with subjective and biased removal of outlier data, leading to unfair and unrepresentative benchmarks in industrial data-driven models, particularly in greenhouse gas emissions standards.
A computer-implemented method for objectively reducing outlier bias through iterative processes involving error thresholding, coefficient recalibration, and optimization techniques to ensure fair and accurate data analysis.
The method effectively minimizes outlier bias, ensuring representative data analysis and improved model performance by iteratively refining model coefficients, thereby enhancing the accuracy and reliability of industrial data models.
Smart Images

Figure 0007762682000009 
Figure 0007762682000010 
Figure 0007762682000011
Abstract
Description
[Technical Field]
[0001] The present invention provides a data analysis method in which outlier components are removed (or filtered) from the analysis development. Analysis relates to the mathematical model that uses data in the calculation or development of simple statistics. The filtering of outlier data is a complex operation involving data quality and data sets for the purpose of operating data authentication or for the application of representative standards, statistics, or subsequent analysis; Calculating data suitable for regression analysis, time series analysis, or mathematical model development The purpose is to
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This continuation-in-part patent application is a continuation-in-part patent application filed on August 19, 2011, entitled "Dynamic Outlier Bias R U.S. Provisional Patent Application No. 13 / 213,780, entitled "Induction System and Method" Priority is claimed from the US Provisional Patent Application Publication No. 2006 / 0110994, filed on May 1, 2006, which is incorporated herein by reference in its entirety. [Background technology]
[0003] Removing outlier data in the baseline or in data-driven model development is fundamental It is important to carry out preliminary analytical work to ensure that a representative and fair analysis is developed from the data available. For example, carbon dioxide (CO2), ozone (O3), water vapor (H2O), Hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), chlorofluorocarbons Carbon dioxide (CFC), sulfur hexafluoride (SF6), methane (CH4), nitrous oxide (N2O) , carbon monoxide (CO), nitrogen oxides (NO x ) and non-methane volatile organic compounds (NMVO C) To develop a fair benchmark for greenhouse gas standards for emissions, standards development The collected industrial data used must exhibit certain characteristics. Extremely good or bad performance may bias the benchmark calculated against other industrial zones. The inclusion of such performance results in the benchmark calculation may be considered unfair or It is judged to be unrepresentative. Historically, performance outliers have required subjective input. Current systems and methods are data-driven approaches. This approach allows the task to be completed at a preliminary analysis or preliminary model development stage. rather, it is an integral part of the model development.
[0004] Bias removal is a subjective process where justification is documented in a prescribed format to justify any changes to the data. However, any form of outlier removal will alter the results of the calculation. Such data filtering is a form of data censoring with the possibility of may or may not reduce bias or error in the analysis and is not intended to be a full analytical disclosure. Strict data removal guidelines and documentation to remove outliers should be included in the analysis results. Therefore, there are many applications in the industry that involve data quality operations, data validation, statistical calculations or mathematical models. Objectively remove outlier data bias using dynamic statistical processes useful for analysis, analysis, and analysis of statistical data. There is a need to provide a new system and method for outlier bias removal. Methods can also be used to classify data into representative categories. The data is applied to develop a mathematical model customized for each group. as multiplicative and additive factors in mathematical models, and other inherently nonlinear quantities. The coefficients are defined as value parameters, e.g., f(x,y,z)=a*x+b*y c + In the mathematical model of d*sin(ez)+f, a, b, c, d, e and f are all coefficients. The values of these terms can be fixed or can be part of the development of a mathematical model. do. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2007 / 117233(A1) Pamphlet Summary of the Invention
[0006] A preferred embodiment is a computer-implemented method for reducing outlier bias that includes the following steps: The steps include selecting a bias criterion, providing a data set, and Step 1: Provide a set of model coefficients; Step 2: Select a set of target values (1) generating a set of multiple predictions for the completed data set; (2) generating an error set for the data set; 3) generating a set of error thresholds based on the error set and the bias criterion; (4) calculating a program based on the error set and the set of error thresholds; (5) the processor generating a truncated data set; (6) generating a set of new model coefficients; and (7) assigning the new model coefficients to the set of new model coefficients. The set is used to perform step (1 In a preferred embodiment, the steps (1) to (5) are repeated. A set is generated based on the data set and the set of model coefficients. In a preferred embodiment, the error set is calculated by dividing the set of predicted values by the set of predicted values and the set of predicted values by the set of predicted values. A set of absolute errors and a set of relative errors generated based on the set of reference values. In another embodiment, the error set comprises the set of multiple predictions. and the set of target values. wherein generating the set of new coefficients further comprises: The matching of the set of predicted values with the set of actual values can be achieved using a model. and minimizing the set of errors between the set. In this case, the censored performance termination criteria are based on one standard error and one coefficient of determination. Based on.
[0007] Another embodiment includes a computer-implemented method for reducing outlier bias, the method comprising: The method includes the steps of selecting an error criterion, selecting a data set, and selecting a set of a plurality of actual values; selecting a set of a plurality of model coefficients; selecting an initial set, completing the data set and the plurality of model coefficients; (1) generating a set of model predictions based on the initial set; based on the model predicted values and the set of actual values for the data set (2) generating a set of errors; A set of error thresholds is calculated based on the completed set of numerical errors and the error criteria. (3) generating an outlier-removed dataset. Therefore, the filtering is performed on the completed data set and the data set of multiple error thresholds. (4) a step based on the filtered data set and a plurality of old coefficients; generating a set of new coefficients based on the set of new coefficients, (5) generating said set of coefficients by a computer processor; A plurality of new model coefficients are calculated based on the filtered data set and the set of new model coefficients. generating a set of multiple outlier bias-reduced model predictions, The generation of this set of reduced-bias model predictions is performed by a computer processor. (6) calculating a plurality of values based on the model predicted values and the set of a plurality of actual values; generating a set of model performance values for the plurality of coefficients from the previous iteration, Complete one performance using the set of new coefficients instead of the set Repeating steps (1) to (6) until the criteria are met, and multiple models. Storing the set of predicted values on a computer data medium.
[0008] Another embodiment includes a computer-implemented method for reducing outlier bias, the method comprising: The steps include selecting one target variable for the facility, selecting a set of a plurality of actual values of variables, the set of actual values being associated with the target variable; a step of identifying a plurality of variables for the facility; a step of obtaining a data set for the facility; wherein the data set includes a plurality of values for the plurality of variables; selecting a bias criterion for (1 ) a plurality of predicted values based on the completed data set and the set of model coefficients. (2) generating a set of predicted values and a set of actual values; generating a set of multiple censored model performance values based on the set; (3) the set of predicted values and the set of actual values for the target variable; (4) generating an error set based on the error set and the bias base; (5) generating a set of error thresholds based on the data set; and a processor for generating a truncated data set based on the set of error thresholds. (6) generating a set of the censored data set and a plurality of model coefficients; generating a set of new model coefficients based on the set of (7) generating the project based on the data set and the set of new model coefficients; (8) generating a set of new predicted values by the processor; and a plurality of new censored model performance values based on the set of a plurality of actual values. generating a set of new coefficients, and using the set of new coefficients to generate a truncated parameter Repeat steps (1) to (8) until the performance termination criteria are met. storing said set of new model predictions in a computer data medium. do.
[0009] Another embodiment includes a computer-implemented method for reducing outlier bias, the method comprising: That is, a step of determining one target variable for the facility, The target variable is a metric for an industrial facility related to its production, financial performance, or emissions. a step of identifying a plurality of variables for the facility, The variables are multiple direct variables for the facility that affect the target variable, and each of them is A complex parameter for the facility that is a function of at least one direct facility variable that affects the target variable. A set of transformed variables, a step, an absolute error, and a relative error. selecting an error criterion; obtaining a data set for the facility; a step of: selecting a set of actual values of the target variables; selecting the completed data set and the initial set of model coefficients; generating a set of model predictions based on the model predictions; A complete set of errors is generated based on the complete set and the set of actual values. generating a set of relative errors according to the formula: m =((predicted value m - Actual value m ) / actual value m ) 2 (where "m" is a reference number), and the absolute error is calculated using the formula: error m =(predicted value m - Actual value m ) 2 Steps, multiple model predictions calculated using and a plurality of model performance values based on the set of a plurality of actual values. generating a set of overall model performance values, The first step involves the first standard error and the first coefficient of determination. a set of errors based on the model predicted values and the set of actual values for (2) generating a complete set of errors and a complete set of data; generating a set of error thresholds based on the error criteria for the set; (3) Removing data having an error value equal to or greater than the error threshold value to eliminate outliers generating a value-removed data set, the filtering of which is completed; (4) a linear optimization step based on the set of data and the set of error thresholds; and a nonlinear optimization model for predicting the predicted values of the plurality of predicted values using at least one of the optimization model and the nonlinear optimization model. The outlier removal is performed by minimizing the error between the set of values and the set of multiple actual values. a plurality of outlier bias reductions based on the removed data set and the set of model coefficients; generating a set of new model predictions, (5) the outlier-removed data; generating a set of new coefficients based on the set and the old set of coefficients; wherein generating said set of new coefficients is performed by said computer processor. (6) calculating a new set of forecast model values and a new set of actual values; generating a set of overall model performance values based on the The set of multiple model performance values includes a second standard error and a second coefficient of determination. , using the set of new coefficients instead of the set of coefficients from the previous iteration. Repeat steps (1) to (6) using the same method until one performance termination criterion is met. the performance termination criteria being a standard error termination value and a decision Including coefficient termination values and meeting the performance termination criteria means the standard error termination The value is greater than the difference between the first and second standard errors and the coefficient of determination end value is greater than the first and a second coefficient of determination, and a plurality of new model predictions. Storing the set of values on a computer data medium.
[0010] Another embodiment includes a computer-implemented method for reducing outlier bias, the method comprising: The method includes the steps of selecting an error criterion, selecting a data set, and selecting a set of the plurality of actual values; selecting a set of the plurality of model predicted values; selecting an initial set of model predicted values and a set of actual values; determining a set of errors based on the set; (1) determining the completeness of the set of errors; determining a set of error thresholds based on the generated set and the error criterion; (2) generating an outlier-removed data set, filtering based on the data set and the set of error thresholds; (3) Based on the outlier-removed dataset and multiple previous model predictions, generating a set of value bias reduced model predicted values, The generation of this set of reduced model predictions is a step performed by a computer processor. (4) based on said set of multiple new model predicted values and said set of multiple actual values; determining a set of errors for the plurality of model predictions from the previous iteration; A single performance termination criterion is met using multiple new model predictions instead of a set. Repeat steps (1) to (4) until the desired result is achieved, and then reduce the bias of multiple outliers. storing the set of completed model predictions on a computer data medium.
[0011] Another embodiment includes a computer-implemented method for reducing outlier bias, the method comprising: That is, a step of determining one target variable for a facility, identifying a plurality of variables for which the plurality of variables influence the target variable; There are multiple direct variables for the facility that affect the target variable, and at least one direct variable that affects the target variable. A set of transformed variables for the facility that are all functions of one direct facility variable. a step of selecting an error criterion including an absolute error and a relative error; obtaining a data set containing a plurality of values for the plurality of variables; selecting a set of a plurality of actual values of the target variable; selecting a set, and applying a set of model coefficients to the data set. generating a set of a plurality of model predictions by a set of performance values based on the set and a set of actual values; determining a first standard error and a second standard error of the set of performance values; (1) the set of model predictions and the completed data; generating a set of errors based on the set of actual values for the data set; The relative error is calculated using the formula: Relative error m =((predicted value m - Actual value m ) / actual value m ) 2(where "m" is a reference number), and the absolute error is calculated using the formula: Absolute Error m =(prediction value m - Actual value m ) 2 ) and (2) the completed data set. and determining a plurality of error thresholds based on the completed set of errors and the error criteria. (3) generating a set of values that are equal to or greater than the set of error thresholds; Generate an outlier-removed data set by removing data with erroneous values. the filtering being based on the data set and a plurality of error thresholds. (4) a step based on the set, the outlier-removed data set and a plurality of old coefficients (5) generating a set of new coefficients based on the set of and a nonlinear optimization model for predicting the predicted values of the plurality of predicted values using at least one of the optimization model and the nonlinear optimization model. The outlier removal is performed by minimizing the error between the set of values and the set of multiple actual values. A plurality of outlier-bias-reducing methods are performed based on the set of the removed data set and the plurality of new model coefficients. generating a set of reduced model predictions, wherein generating the model predictions comprises: (6) generating a plurality of outlier bias reduced models; A plurality of updated performance data are generated based on the set of predicted performance data and the set of actual performance data. generating a set of performance values, said set of updated performance values being The iteration includes a second standard error and a second coefficient of determination. a performance termination criterion using multiple new coefficient sets instead of the set repeating steps (1) to (6) until the performance is satisfied; The performance termination criteria include one standard error termination value and one coefficient of determination termination value, and The end-of-month criterion is met if the end-of-month standard error is the difference between the first and second standard errors. and the end value of the coefficient of determination is greater than the difference between the first and second coefficients of determination. and computing the set of multiple outlier bias reduction factors as a computer data set. and storing the data on a data medium.
[0012] Another embodiment is to estimate the feasibility of a data set used in developing a model. The method includes a computer-implemented method for evaluating a plurality of providing a target data set including a number of data values; generating a random target data set based on a set of bias metrics; selecting a bias metric based on the data set and the selected bias metric; generating a target dataset with reduced outlier bias; The processor performs one outlier bias reduction based on the dataset and each selected bias criterion. generating a random dataset, the outlier bias reduced dataset, and Calculating a set of error values for the outlier bias reduced random data set. a step of: calculating a set of correlation coefficients for the data set; The data set and the corresponding line are then compared based on the bias criteria and the corresponding error values and correlation coefficients. generating a plurality of bias reference curves for a random data set; the bias criterion curve for the set and the bias criterion curve for the random data set and comparing the outlier bias-reduced target dataset with the outlier. The bias-reduced random target data set is generated using a dynamic outlier bias removal method The random target data set is expanded from a plurality of values within the range of the plurality of data values. The set of error values may consist of a plurality of randomly selected data values. where the set of correlation coefficients may include a set of standard errors of the Other embodiments further include a set of values for the bias relative to the target data set. Based on a comparison of the reference curve and the bias reference curve for the random target data set. , regarding the feasibility of the target dataset to support the deployed model and vice versa. generating an automated advice based on the correlation coefficient threshold. and / or generated based on analyst-selected parameters, such as error thresholds. Still another embodiment further includes the following steps: providing a real data set containing a plurality of real data values corresponding to model predicted values; generating a random actual data set based on the actual data set; The processor determines an outlier based on the actual data set and each selected bias metric. generating a reduced bias actual data set, the random actual data set, and The processor generates an outlier-bias-reduced random execution result based on each selected bias metric. generating a real dataset, the real dataset having the outlier bias reduced for each selected bias criterion; Based on the random target data set and the outlier bias reduced random actual data generating a random data plot, and determining the deviation for each selected bias criterion; Based on the target dataset with reduced outlier bias and the actual target dataset with reduced outlier bias, generating a realistic data plot based on the random data plot; and comparing the plots with the corresponding real data plots corresponding to each selected bias criterion. be.
[0013] A preferred embodiment includes a server including a processor and a storage subsystem, and a data center. a database including the data and stored by said storage subsystem; and a computer program stored by the system. A computer program contains instructions that, when executed, cause the processor to: This includes selecting a bias criterion and providing a set of model coefficients. and (1) selecting a set of multiple target values for the data set. (2) generating a set of measurements; and (3) generating a set of errors for the data set. (3) determining one of a plurality of error thresholds based on the error set and the bias criterion; (4) generating a set of error thresholds based on the error set and the set of error thresholds; (5) generating a set of new model coefficients based on the censored data set; (6) generating a set of new model coefficients; and Repeat steps (1) to (5) until the performance termination criteria are met. In a preferred embodiment, the set of multiple predictors is In a preferred embodiment, the set of model coefficients is generated. An error set is generated based on the set of multiple predicted values and the set of multiple target values. In another embodiment, the method further comprises a set of absolute errors and a set of relative errors. The error set is the difference between the set of predicted values and the set of target values. In another embodiment, the method for generating the set of new coefficients includes a value calculated as: The step of determining further comprises: determining a plurality of predicted values and a plurality of actual values; This involves minimizing this set of numerical errors using a linear or non-linear optimization model. In a preferred embodiment, the truncation performance can be achieved using The termination criteria are based on one standard error and one coefficient of determination.
[0014] Another embodiment of the present invention comprises a server including a processor and a storage subsystem; a database containing the data set and stored by the storage subsystem; and a computer program stored by the subsystem. The computer program includes instructions that, when executed, cause the processor to: This includes instructions for selecting an error criterion and selecting a set of actual values. selecting an initial set of coefficients; (1) generating a complete set of model predictions from the initial set; based on the set of model predicted values and multiple actual values for the data set created (2) generating a set of errors for the completed data set; A set of error thresholds is calculated based on the completed set of numerical errors and the error criteria. (3) generating an outlier-removed data set; The filtering is based on the completed data set and the set of error thresholds. (4) based on the outlier-removed data set and the set of coefficients; generating a set of multiple outlier bias-reduced model predictions, The generation of the set of reduced value bias model predicted values is performed by a computer processor. (5) based on the outlier-removed data set and the set of multiple old coefficients; generating a set of new coefficients in the (6) performing the outlier bias reduced model; a set of model performance values based on the set of model predicted values and a set of model actual values generating a set of new coefficients in place of the set of coefficients from the previous iteration; Step (1) using the set until one performance termination criterion is met. Repeating (6) and a set of overall multiple outlier bias reduction model predictions The purpose is to store the information on a computer data medium.
[0015] Yet another embodiment includes a server including a processor and a storage subsystem; A database stored by the system that contains a target variable for the facility, a set of actual values of the target variable, a set of variables for the facility that are related to the target variable; a data set for the facility that includes multiple values for the multiple variables; and a computer program stored by the storage subsystem. The computer program, when executed, causes the processor to: It includes instructions to cause the following: selecting a bias criterion, (1) selecting a set of coefficients for the data set and the plurality of model coefficients; (2) generating a set of multiple predicted values based on the set of multiple predicted values; and a set of a plurality of censored model performance values based on the set of a plurality of actual values. (3) generating a set of predictors for the target variable; (4) generating an error set based on the set of actual values of and generating a set of multiple error thresholds based on the bias criterion; (5) A data set and a truncated data set based on the set of error thresholds. (6) generating a set of the censored data set and the model coefficients; (7) generating a set of new model coefficients based on the data set; generating a set of new predicted values based on the set of new model coefficients; (8) generating a plurality of new predicted values based on the set of new predicted values and the set of actual values; generating a set of new censored model performance values for the plurality of new coefficients; Repeat steps (1) to (3) until the performance termination criterion is met. (8) repeating the process and storing the set of new model predictions in the storage subsystem. The idea is to store it in the system.
[0016] Another embodiment includes a server including a processor and a storage subsystem and a single server for the facility. a database containing a data set and stored by said storage subsystem; and a computer program stored by the storage subsystem. The computer program, when executed, causes the processor to: It contains instructions for determining a single target variable, identifying multiple variables, and The plurality of variables includes a plurality of direct variables for the facility that affect the target variable, Each of these variables is a function of at least one direct variable that affects the target variable. a set of transformed variables for each, an absolute error and a relative error selecting a set of actual values of the target variable; and selecting an initial set of coefficients. generating a set of model predictions from the initial set; determining a set of errors based on the set of actual values and the set of actual values; and the relative error is given by the formula: Relative error m =((predicted value m - Actual value m ) / actual value m ) 2 (" m is a reference number), and the absolute error is calculated using the formula: Absolute Error m =(predicted value m - Actual Value m ) 2 and calculating the set of model predictions using a plurality of model predictions and a plurality of actual predictions. determining a set of a plurality of performance values based on said set of actual values; The set of performance values includes a first standard error and a first coefficient of determination; ) generating a set of errors based on the model predicted values and the set of actual values; (2) creating a complete set of errors for the complete data set; and generating a set of multiple error thresholds based on the error criterion; (3) filtering data having error values outside the set of error thresholds; generating an outlier-removed data set by filtering the (4) the analysis is based on the data set and the set of error thresholds; A plurality of predictions are made using at least one of the linear optimization model and one nonlinear optimization model. by minimizing an error between the set of values and the set of actual values; a plurality of new models based on the outlier-removed data set and the set of coefficients; generating a set of predicted values, wherein the generation of the outlier bias reduced model predicted values comprises: (5) the outlier-removed data set; and generating a set of new coefficients based on the set of old coefficients. and generating said set of new coefficients is performed by said computer processor. (6) generating a plurality of new model predictions based on the set of new model predictions and the set of actual values; generating a set of numerical performance values, the plurality of model performance values; The set of coefficients from the previous iteration includes a second standard error and a second coefficient of determination. Complete one performance using the set of new coefficients instead of the set Repeat steps (1) to (6) until the performance criteria are met. The performance termination criteria include one standard error termination value and one coefficient of determination, and The termination criterion is met when the standard error termination value is less than the difference between the first and second standard errors. and the end value of the coefficient of determination is greater than the difference between the first and second coefficients of determination. and storing said set of a plurality of new model predictions on a computer data medium. This is what we should do.
[0017] Another embodiment of the present invention comprises a server including a processor and a storage subsystem; a database containing the data set and stored by the storage subsystem; and a computer program stored by the subsystem. The computer program includes instructions that, when executed, cause the processor to: This includes instructions for selecting an error criterion, selecting a data set, Selecting a set of actual values; selecting an initial set of model predicted values and generating a plurality of model predicted values based on the set of model predicted values and the set of actual values. (1) determining a set of errors, (2) determining the complete set of errors and the errors, (2) determining a set of error thresholds based on a criterion; generating a filtered dataset, the filtering being performed on the dataset and (3) the outlier-removed data set based on said set of error thresholds; and generating a plurality of outlier bias-reduced models based on the completed set of model forecasts. generating a set of model predictions, the set of model predictions being the multiple outlier bias-reduced predictions; (4) the generation of the set is performed by a computer processor; and generating a plurality of error reduction models based on the set of predicted values and a corresponding set of actual values. determining a set of differences, a plurality of outliers in place of the set of model predictions; A performance termination criterion is met while using that set of bias-reduced model predictions. Repeat steps (1) to (4) until the outlier bias reduction factor is satisfied. and storing the set on a computer data medium.
[0018] Another embodiment of the present invention comprises a server including a processor and a storage subsystem; a database containing the data set and stored by the storage subsystem; and a computer program stored by the subsystem. The computer program includes instructions that, when executed, cause the processor to: This includes determining a target variable, identifying multiple variables for the facility, and The plurality of variables include a plurality of variables for the facility that affect the target variable. A function of direct variables and at least one key facility variable, each of which affects the target variable. a set of transformed variables for the facility, and an absolute error and selecting an error criterion including a single relative error; and selecting a set of actual values of the target variable from a data set including the target variable. selecting an initial set of coefficients; and computing the set of model coefficients according to the data. generating a set of multiple model predictions by applying the model to a dataset; a plurality of performance indicators based on said set of model predicted values and said set of a plurality of actual values; determining a set of performance values, said set of performance values being a first target; (1) the set of model predictions and the first coefficient of determination; determining a set of errors based on the set of actual values, is the formula: relative error k =((predicted value k - Actual value k ) / actual value k ) 2 ("k" is a reference number) The absolute error is calculated using the formula: k =(predicted value k - Actual value k ) 2 Use (2) the set of errors for the completed data set and and determining a set of error thresholds based on said error criteria; (3) By removing data with an error value equal to or greater than the difference threshold, an outlier-removed data set is obtained. The filtering involves generating a dataset and multiple errors. (4) based on said set of difference thresholds, said outlier-removed data set and a plurality of (5) generating a set of new coefficients based on the set of old coefficients; A plurality of predicted values are calculated using at least one of a linear optimization model and a nonlinear optimization model. Minimizing an error between the set and the set of actual values and the outliers Multiple outlier bias reduction based on the removed dataset and the set of coefficients (5) generating a set of multiple outlier bias-reduced model predictions; a plurality of updated performances based on the set of values and the plurality of actual values; generating a set of updated performance values, said set of updated performance values being a second Including standard errors and second coefficients of determination, instead of the set of coefficients from the previous iteration Use this set of new coefficients until a performance termination criterion is met. Repeat steps (1) to (5) with the performance termination criteria being one Includes a standard error exit value and a coefficient of determination exit value, and meets the performance exit criteria. The standard error ending value is greater than the difference between the first and second standard errors and the end value of the coefficient of determination is greater than the difference between the first and second coefficients of determination; and and storing the set of multiple outlier bias reduction factors in a computer data medium. be.
[0019] Yet another embodiment is to perform the simulation of a data set used in developing a model. The system includes a processor and a storage subsystem. a server including a target dataset including a plurality of model predictions and a storage subsystem; The database stored by the storage subsystem and the and a computer program that, when executed, The command contains instructions that cause the processor to: generating a target dataset and a set of bias criteria; and generate multiple outlier-bias-reduced datasets based on each selected bias metric. and determining an outlier based on the random target data set and each selected bias metric. generating a bias-reduced random target data set; a plurality of error values for the outlier bias-reduced random target dataset; calculating a set of the outlier bias reduced target dataset and the outlier bias; Calculating a set of correlation coefficients for the reduced random target data set. , the target data based on the corresponding error value and correlation coefficient for each selected bias criterion. generating a set and a plurality of bias reference curves for the random target data set; and the bias reference curve for the target data set and the random target data set. The processor performs dynamic outlier bias removal by comparing the bias reference curve against the The outlier bias reduced target dataset and the outlier bias reduced run are calculated using the method. generating a random target dataset, the random target dataset being a subset of the plurality of data; It may consist of a plurality of randomly selected data values derived from a plurality of values within a range of values; The set of error values may include a set of standard errors. The set includes a set of coefficient of determination values. Additionally, it contains instructions that, when executed, cause the processor to: the bias reference curve for the target data set and the bias reference curve for the random target data set The objective is to generate automated advice based on comparison with the bias reference curve. Advice may be provided based on analyst-selected criteria, such as correlation coefficient thresholds and / or error thresholds. In yet another embodiment, the system The database of the system further comprises a plurality of actual data values corresponding to the model predictions. The program further includes a single actual data set, and when executed, It contains instructions that cause the processor to: generating a random actual data set; generating an outlier bias-reduced actual data set based on the bias criterion value; One outlier bias reduced based on a random real data set and each selected bias metric generating a random real data set, and determining the outlier bias for each selected bias criterion; Based on the reduced random target data set and the outlier bias reduced random actual data generating a random data plot based on the selected bias criterion; The outlier bias reduced target dataset and the outlier bias reduced actual target dataset generating a realistic data plot based on the random data plot; and the actual data plots corresponding to each selected bias criterion.
[0020] Another embodiment is a system for reducing outlier bias in the target variable measured for a facility. The present invention also includes a system including a computer unit for processing a data set. The computer unit includes a processor and a storage subsystem, the data being processed, an input unit for inputting a set of measurements of a given target variable and a corresponding data set; an input unit including a measuring device for providing a processed data set; and an output unit for outputting a processed data set. , including computer programs stored by said storage subsystem. The computer program contains instructions that, when executed, cause the processor to perform the following steps: The steps include selecting the target variable for the facility, identifying a plurality of variables for the facility; obtaining a dataset comprising a plurality of values for the plurality of variables; a step of selecting a bias criterion; a step of selecting a set of model coefficients; (1) generating a set of predictions for the dataset; (2) generating an error set for the data set; (3) generating a set of error thresholds based on the set and the bias criterion; 4) A truncated data set based on the error set and the set of error thresholds. (5) generating a set of new model coefficients. and (6) using the set of new model coefficients to perform a censored performance test. This is the step of repeating steps (1) to (5) until the process termination criteria are met.
[0021] Additionally, other embodiments may involve the purchase of equity securities (e.g., common stock) or derivative contracts (e.g., Target variables measured against financial instruments such as stocks, forwards, futures, options, swaps, etc. The system includes a system for reducing outlier bias in a data set. a computer unit including a processor and a storage subsystem; unit, an input unit that receives the dataset to be processed, and an input unit including a storage device for storing data for an output unit that outputs the processed data set; This includes computer programs that, when executed, instructions to cause a processor to perform the steps of: selecting the target variable, the target variable (e.g., dividend, revenue, cash flow) identifying a plurality of variables for the instrument that are related to the financial instrument's performance, such as the obtaining a data set corresponding to the plurality of variables, the data set comprising: a step including a plurality of values for the model coefficients; a step of selecting a bias criterion; a step of selecting a set; (1) selecting a set of a plurality of predictors for the data set; (2) generating an error set for the data set. (3) determining a set of error thresholds based on the error set and the bias criterion; (4) generating a plurality of error thresholds based on the error set and the set of error thresholds. (5) generating a set of the new model coefficients; (6) generating a set of new model coefficients using the set of new model coefficients. Repeat steps (1) to (5) until the termination criteria are met. It's Tep. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a flowchart illustrating one embodiment of a method for identifying and removing data outliers. [Figure 2] 1 is a flowchart illustrating one embodiment of a method for identifying and removing data outliers for data quality operations. [Figure 3] 1 is a flow chart illustrating one embodiment of a method for identifying and removing data outliers for data validation. [Figure 4] 1 is an exemplary node for implementing the method of the present invention; [Figure 5] 1 is an exemplary graph for quantitative evaluation of a data set. [Figure 6]6A and 6B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the randomized and realistic data sets relative to the entire data set, respectively. [Figure 7] 7A and 7B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the unselected data set and the realistic data set, respectively, after removing 30% of the data as outliers. [Figure 8] 8A and 8B are graphs for quantitative evaluation of the data sets of FIG. 5, illustrating the unselected data set and the realistic data set, respectively, after removing 50% of the data as outliers. [Figure 9] 1 illustrates an example system used to reduce outlier bias in a target variable measured for a facility. DETAILED DESCRIPTION OF THE INVENTION
[0023] The following disclosure relates to systems and methods for accessing and managing structured content. Many different embodiments or examples are provided that implement different features of the components, processes, and Specific examples of processes and implementations are described to help clarify the invention. It is merely an example and is not intended to limit the invention beyond what is set forth in the claims. Well-known elements are not intended to obscure the preferred embodiments of the present invention in unnecessary detail. For the most part, the preferred embodiments of the present invention are presented without detailed description. Details not necessary to obtain a complete understanding of the embodiments are omitted, as such details are within the skill of those skilled in the art. It is omitted whenever possible.
[0024] A mathematical description of one embodiment of dynamic outlier bias reduction is as follows.
number
number
[0025]
number
[0026] Another mathematical description of one embodiment of dynamic outlier bias reduction is as follows:
number
number
[0027]
number
[0028] After each iteration, new model coefficients are calculated from the current censored data set, The removed data from the previous analysis plus the current censored data are recombined. This combination encompasses all data values in the complete data set. The current model coefficients are then applied to the completed data to calculate the completed set of forecasts. The absolute and relative errors are calculated for the complete set of predictions. A new bias-based percentile threshold is calculated. If the absolute or relative error is greater than the threshold, A new truncated data set is created by removing all large data values, The nonlinear optimization model is then applied to the newly censored data set to generate a new model. This process ensures that all data values are included in the model data set. The possibility of including the model coefficients as the best fit to the data can be examined at each iteration. When the iteration converges to a value that matches the iteration, some data values that were excluded in the previous iteration are included in the subsequent iteration. It may also be included in the
[0029] In one embodiment, variability in greenhouse gas emissions leads to bias in model predictions. Errors in environmental conditions and calculation procedures may result in overestimation or underestimation of emission results. These non-industrial influences may cause results for specific facilities to be biased in model predictions. Unless these differences are eliminated, the results will be fundamentally different from similar facilities. Biases also exist due to unique operating conditions.
[0030] If the analyst is certain that the facility's calculations are incorrect or have unique extenuating characteristics, If reliable, bias can be manually removed by simply removing the facility's data from the calculation. However, facility performance data from many different companies, regions and countries can be collected. When measuring performance, precise a priori knowledge of the data details is not realistic. No analyst-based data removal procedures were documented for the model results, and no data were It has the potential to add unsupported bias.
[0031] In one embodiment, to determine statistical outliers to be removed from the model coefficient calculations: Dynamic outlier bias reduction is applied to the procedure using the data and a predetermined overall error criterion. This uses a global error criterion provided by the data, e.g., using a percentile function. Dynamic outlier bias reduction is a data-driven process that uses Its use in this embodiment is illustrative but not limited to reducing bias in model predictions. Dynamic outlier bias reduction can also be used to reduce the number of outliers from any statistical data set, for example. It is used to remove outliers. This is useful, for example, in arithmetic means, linear regressions, and trend line calculations. This includes, but is not limited to, use in calculations. Outlier facilities will still be ranked lower in the calculations. Although the outliers are ranked, they are not included in the filters applied to calculate model coefficients or statistical results. Not used in filtered datasets.
[0032] A commonly used standard procedure for removing outliers is to multiply the standard deviation (σ) of a data set by We can simply define all data that fall outside the 2σ interval from the mean as outliers, for example. This procedure generally has statistical assumptions that cannot be tested in practice. A description of the dynamic outlier bias reduction method applied in one embodiment is summarized in FIG. and uses both relative and absolute error. For example, for facility "m": Relative Error m =((predicted value m - Actual value m ) / actual value m ) 2 (1) Absolute Error m =(predicted value m - Actual value m ) 2 (2) This becomes:
[0033] In step 110, the analyst selects an error threshold that defines outliers to be removed from the calculation. Specify value criteria, e.g., relative and absolute error using percentile operations as error functions. The 80th percentile value for the relative error can be set. Calculation of data values below the 80th percentile and data values at the 80th percentile for absolute errors The remaining values are either removed or considered outliers. In this example, for a data value to be avoided from being removed, the data value must be relative and Both the relative and absolute errors must be less than the 80th percentile. In other embodiments, the percentile thresholds for both the pairwise errors can be varied independently. In this case, only one percentile threshold is used.
[0034] In step 120, the model standard error and the coefficient of determination (r 2 ) percent change criteria While the values of these statistics vary from model to model, the performance of the previous iterative procedure is The cents change can be preset, for example, at 5 percent. The value can be used to terminate the iteration procedure. Other termination criteria are a simple number of iterations It could be.
[0035] In step 130, an optimization calculation is performed to generate model coefficients and predicted values for each facility. will be carried out.
[0036] In step 140, the relative weights for all facilities are calculated using equations (1) and (2). Both the mean and absolute error are calculated.
[0037] In step 150, the error function with the threshold criteria specified in step 110 is calculated. is applied to the data calculated in step 140 to determine the outlier threshold.
[0038] In step 160, the data is calculated as a relative error, an absolute error, or to include only facilities where both errors are less than the error threshold calculated in step 150. It is filtered.
[0039] In step 170, an optimization calculation is performed using the outlier-removed data set. do.
[0040] In step 180, the standard error and r 2 The percent change in If the percent change is greater than the criteria, step 140 is performed. If not, the iterative procedure continues at step 190. The resulting model calculated from this dynamic outlier bias reduction criterion is then finalized. The model results are based on the current iteration, regardless of the status of previously removed or accepted data. This rule applies to all facilities, not just those listed above.
[0041] In another embodiment, the process begins with the selection of predetermined iteration parameters. Specifically, (1) absolute error and relative error, one or both of which are used in an iterative process; (2) coefficient of determination (r 2 (also known as) improvement value, and (3) standard This is the error improvement value.
[0042] The process involves the creation of an original data set, a set of actual data, and a set of At least one coefficient or one factor used to calculate the predicted value in It starts with a coefficient or set of coefficients that are applied to the original data set to produce a set of predicted values. The set of coefficients includes, but is not limited to, scalar, exponential, parametric, and periodic functions. The set of predicted data is then compared to the set of actual data. Standard errors and coefficients of determination are calculated based on the difference from the actual data. The data points are associated with a metric to remove data outliers based on the percentile and relative error. The absolute and relative errors calculated are used. No ordering of the data is required. Absolute and / or relative Any data outside the range associated with the percentile value for the error is removed from the original dataset. Using absolute and relative errors to filter data This method is illustrative and for illustrative purposes only. This is because it can be done for only or for other functions.
[0043] Data associated with absolute and relative errors that fall within a user-selected percentile range are considered outliers. A filtered dataset, where each iteration of the process produces its own filtered data set. This first outlier-removed data set is used to compare the actual values. At least one coefficient is determined by optimizing the error. The coefficients are then used to generate a prediction based on the first outlier-removed data set. The outlier bias reduced coefficient is the mechanism by which knowledge is transferred from one iteration to the next. It functions as:
[0044] After the first outlier-removed data set is created, the standard error and coefficient of determination are calculated, and The standard error and coefficient of determination are compared with those of the original data set. If both of the differences are less than their respective improvement values, the process stops. If at least one is not met, the process continues with another iteration. The use of numbers to check the iterative process is illustrative and exemplary only. The checks can be based on standard errors only or coefficients of determination only, different statistical checks, or other (number of replicates) This can be done using performance exit criteria such as
[0045] If the first iteration fails to meet the improvement criteria, a new set of predictions is determined. The second iteration begins by applying the first outlier bias-reduced data coefficient to the raw data. In this case, the original data is processed again and the coefficients of the first outlier-removed data set are used. The absolute and relative errors for the data points and the original data set during use The standard error and coefficient of determination values are established. The data are then filtered to remove second outliers. A removed data set is formed, and coefficients based on the second outlier-removed data set are calculated. is determined.
[0046] However, the second outlier-removed data set is not necessarily the same as the first outlier-removed data set. It is not a subset of the original data set, but a second set of model coefficients with reduced outlier bias. , the second standard error and the second coefficient of determination. Once these values are determined, , the second standard error is compared to the first standard error, and the second coefficient of determination is compared to the first coefficient of determination. can be.
[0047] If the improvement (standard error and coefficient of determination) exceeds the difference between these parameters, the process is Otherwise, another iteration is performed by processing the original data again. At this point, the original data set must be processed and a new set of predictions must be generated. The second outlier bias reduced coefficient is used. User selected hundred for absolute and relative error. Determine the set of third outlier-biased coefficients by filtering based on quantile values A third outlier-removed data set is created that is optimized to improve the error correction. Continue until success or some other termination criterion (such as a convergence criterion or a specific number of iterations) is met. .
[0048] The output of this process is a set of coefficients or model parameters, where the coefficients or A model parameter is a mathematical value (or set of values) that can be used to calculate, for example, data, a linear equation, or a Model predictions for comparing the slope and intercept values of an equation, exponents, or polynomial coefficients. The output of Dynamic Outlier Bias Reduction is not its own output value, but rather This is a factor that modifies the data to determine the force value.
[0049] In another embodiment illustrated in FIG. 2, dynamic outlier bias reduction is performed based on the data for a particular use. Data quality assessment to assess the consistency and accuracy of data to ensure it is appropriate for the For data quality operations, this method does not involve an iterative procedure. Other data quality methods can also be used in conjunction with Dynamic Outlier Bias Reduction during the process. The method is applied to the calculation of the arithmetic mean of a given data set. For example, consecutive data values fall within the same range. Any widely spaced values constitute poor quality data. In this case, the error term is The error values are constructed from continuous values and dynamic outlier bias reduction is applied to these error values.
[0050] In step 210, the initial data is listed in any order.
[0051] Step 220 configures the function or operation to be performed on the data set. In the example form, the functions and operations are such that each line corresponds to the average of all data above that line. It is an ascending ranking of data followed by successive arithmetic mean calculations.
[0052] Step 230 uses the successive values from the result of step 220 to generate a relative and calculate the absolute error.
[0053] Step 240 allows the analyst to input a desired outlier removal error criterion (%). The quality criterion value is calculated from the error calculation in step 230 based on the data in step 220. is the result value of
[0054] Step 250 shows the data quality outlier filtered data set. If the absolute error exceeds a specified error criterion given in step 240, the specified value is removed. can be.
[0055] Step 260 involves calculating the arithmetic mean of the completed data set and the outlier-removed data set. The analyst shall use the identified This is the final step to determine whether the outlier-removed data components are actually of poor quality. Dynamic outlier bias reduction systems and methods allow analysts to directly remove data. The optimal implementation guidelines encourage analysts to review and provide results on the appropriateness of the implementation. It will check the fruit.
[0056] In another embodiment illustrated in FIG. 3, dynamic outlier bias reduction is performed when the data is Data that tests the reasonable accuracy of a data set to determine whether it is appropriate for This method does not involve an iterative procedure for data authentication operations. In the example, dynamic outlier bias reduction is applied to the calculation of the Pearson correlation coefficient between two data sets. The Pearson correlation coefficient is the value of a data point relative to other data points in a data set. Validating a dataset against this statistic means that the results are highly sensitive to It is important to ensure that the majority of data is representative of what is suggested, without the influence of extreme values. The data authentication process in this example ensures that consecutive data values are within a specified range. That is, values that are too far apart (e.g., a particular Any values outside the specified range indicate poor quality data. This is achieved by constructing error terms for successive values of the function. The outlier-removed dataset is then treated as certified data by applying outlier bias reduction. become.
[0057] In step 310, the pairs of data are listed in any order.
[0058] Step 320 calculates the relative and absolute errors for each ordered pair in the data set. Calculate.
[0059] Step 330 allows the analyst to input desired data validation criteria. In the example, both the relative and absolute error thresholds of 90% are selected. The quality metric items are the resulting absolute and relative values for the data presented in step 320. The error percentile value.
[0060] Step 340 shows the outlier removal process, where both relative and absolute The error value of exceeds the value corresponding to the user-selected percentile value entered in step 330. Using criteria, potentially invalid data is removed from the data set. error criterion can be used, so multiple criteria can be applied, as shown in this example. If so, any combination of error values can be applied to determine the rules for outlier removal. can.
[0061] Step 350 calculates the statistical results of the authenticated data and raw data values. is the Pearson correlation coefficient. These results were then examined by the analyst for operational validity. can be done.
[0062] In another embodiment, dynamic outlier bias reduction is used to validate the entire data set. Standard error improvement, coefficient of determination improvement, and absolute and relative error thresholds are selected. , then the dataset is filtered according to an error criterion. Even if the data is of high quality, it may still have error values outside the absolute and relative error thresholds. Therefore, it is important to determine whether any removal of data is necessary. It is important to note that after the first iteration, the outlier-removed data set shows significant improvements in standard error and coefficient of determination. If the original dataset passes the refinement criteria, it is certified. The dataset is too small to be considered significant (e.g., the selected improvement This is because it produces standard errors and coefficients of determination (less than the standard error).
[0063] In another embodiment, the data outlier removal iterations affect the calculation. Dynamic outlier bias reduction is used to provide insight into whether the graph or data A data table is provided, allowing the user to track the progress of the data outlier removal calculation as each iteration is performed. This step-by-step approach allows the analyst to add value and knowledge to the results. We can observe unique properties of the calculation that can add to the Therefore, we propose a dynamic outlier bias reduction approach to calculating representativeness factors for multidimensional datasets. The effect of reduction is shown.
[0064] As an example, a linear regression calculation was performed on a poor quality dataset of 87 records. The regression equation is of the form y=mx+b. Table 1 shows the iteration process for five iterations. It is worth noting that the results of the process are shown below. Convergence is achieved in iterations. The changes in the regression coefficients can be observed. The outlier bias reduction method reduced the calculated data set based on 79 records. A relatively low coefficient of determination (r 2 =39%) is r 2 for statistics and for calculated regression coefficients A lower (<95%) criterion should be tested to examine the additional effect of outlier removal. This shows that... [Table 1]
[0065] Table 2 shows the results of applying dynamic outlier bias reduction using 80% relative and absolute error criteria. It is noteworthy that the outlier error criterion is 15 percentage points (95% to 80%). % change is an additional 35% reduction in the allowable data (which includes 79 to 51 records). r with small 2 This resulted in a 35 percentage point increase (from 39% to 74%) in Analysts should communicate the outlier-removed results to a wide audience and also organize the data of the analysis results. The outliers in Tables 1 and 2 were analyzed to provide insight into the effect of data variability. Graphical illustration of the change in regression line can be used along with the filtered data and numerical results. . [Table 2]
[0066] As illustrated in FIG. 4, one embodiment of a system used to perform the method includes a computer. The hardware includes a computer system with sufficient system memory to perform the required numerical calculations. The processor 410 includes a memory 420. The processor 410 is configured to perform the method. Executes computer programs stored in system memory 420. Display 440 A video and storage controller 430 is used to enable the operation of the system. , including various data storage devices for data entry, e.g., floppy disks Unit 450, Internal / External Disk Drive 460, Internal CD / DVD 470, Tape unit 480, and other types of electronic storage media 490. These storage media are used to identify data sets and outliers. The removal criteria are entered into the system, the outlier-removed data set is stored, and the calculation factors are calculated. The calculations are stored as statistical data, as well as system-generated trend lines and trend line repeat graphs. Applying to software packages or for example Microsoft Excel This can be done from data entered in spreadsheet format using Cell®. The calculations are performed using customized software designed for enterprise-specific system implementation. using a software program or a database and spreadsheet program such as Excel This is done using commercially available software compatible with the system. The system also features dynamic outlier bias. Other databases to provide data for use with the reduction system and method calculations. The output device may interface with a proprietary or public external storage medium 300. The facility provides system-generated graphs and reports, such as calculation worksheets, over an intranet or internet. Management staff, printer 520, input devices 450, 460, 470 via the internet , 480, 490, and a proprietary storage database. These output devices may be long-distance communication devices 510 that transmit to the network 530. It is illustrative and exemplary only.
[0067] As illustrated in Figures 5, 6A, 6B, 7A, 7B, 8A and 8B, in one embodiment Dynamic outlier bias reduction is used to quantitatively and qualitatively assess the quality of a dataset. This means that the errors and correlations of the data values in the data set are within a reasonable range. The error and correlation are compared to a benchmark data set consisting of randomly distributed data values. In one embodiment, the error is set to be the standard error of the data set. Correlation is the coefficient of determination (r 2 ) In another embodiment, the correlation may be expressed as a function of the coefficient of correlation, commonly referred to as Kendall's tau (τ). The correlation coefficient can be specified as the Kendall rank correlation coefficient. In this case, the correlation is the Spearman rank correlation coefficient or the Spearman rho coefficient. As mentioned above, dynamic outlier bias reduction is used to reduce the number of samples identified as outliers. It is used to systematically remove data values that are not relevant to the underlying model or process. The table is not described. Outliers are usually associated with a relatively small number of data values. In reality, however, data sets are insidiously contaminated with spurious values or random noise. The graphical representations in Figures 5, 6A, 6B, 7A, 7B, 8A and 8B show the underlying model. Dynamic outlier bias reduction system to identify situations where a rule is not supported by the data and how the method can be applied. The calculated relative and / or absolute error between the model and the actual data value is, for example, 80%. This is done by removing data values that are greater than a bias criterion based on a percentile. This means that the relative or absolute error percentile value is the 80th percentile (80th of the data values). % have an error less than this value) , the data value is removed.
[0068] As illustrated in Figure 5, realistic model development data was deployed within a real data set. In practice, the analyst may choose a random value. Such understanding is difficult because researchers typically have no prior knowledge of the contamination of their datasets. The dynamic outlier bias reduction system and method are used to obtain repeated results from several model calculations. Figure 5 shows an example model for both data sets. The standard error, that is, the measure of the amount of error that cannot be explained by the model, is shown below. , the coefficient of determination (%) that indicates how much of the data variability is explained by the model or r 2The percentile values next to each point represent the bias measure. For example, 90% means that the data values for the relative or absolute error value are greater than the 90th percentile. Indicates that the data value with the highest error is to be removed as an outlier from the model. 0% corresponds to removing it every iteration.
[0069] As illustrated in Figure 5, bias is observed for both the random and realistic dataset models. Increasing the criteria reduces the error, i.e., the standard error and the coefficient of determination are both However, the standard error for the random data set is The standard error is two to three times larger than in realistic model data sets. The coefficient of determination requirement for e.g. In Figure 5, the bias criterion of 70% for a random data set is and an 80% r with an approximate 85% bias criterion for realistic data. 2 achieved However, the corresponding standard errors for a random data set are This means that the model dataset analysis is performed using different bias criteria. The calculations were systematically performed on a representative pseudodata set and are shown in Figure 5. By plotting the results in this way, the analyst can determine the acceptable bias for the data set. criteria (i.e., acceptable percentage of data values removed), and therefore the overall data Furthermore, such systematic model dataset analysis can be used to assess the quality of the dataset. , the implementation of the dataset used in the model development based on a configurable set of parameters. It can be used to automatically give advice on the validity of data. In one embodiment where the model is developed using dynamic outlier bias removal for the and representative pseudodata sets for model data sets calculated under different bias criteria. The error and correlation coefficient values for the data set in supporting the deployed model are feasibility of the model and, essentially, the feasibility of the deployed model in supporting the dataset. This can be used to automatically give advice on
[0070] As illustrated in Figure 5, these model performance values for several cases By observing the behavior of the data, it is possible to determine whether the data values are representative of the process being modeled. For example, referring to Figure 5, 100% The standard error for a realistic data set under bias criteria (i.e., without bias reduction) is approximately Similarly, the error at a 65% bias criterion (i.e., the highest error is 35% of the removed data values) This corresponds to the standard error for a random data set. This supports the conclusion that
[0071] In addition to the quantitative analysis described above, facilitated by the exemplary graph of FIG. 5, Dynamic Outlier Bias Reduction are available as equally, if not more powerful, subjective procedures to aid in the assessment of the quality of a dataset. This allows for model predictions to be made for both outliers and included outcomes. , by plotting the actual target values given by the data.
[0072] 6A and 6B show the 100% points of both the realistic and random curves in FIG. The large scatter in Figure 6A indicates that the plot is This is consistent with the model's inability to fit intentional randomness. , consistent with and general to a set of actual data, where model predictions and actual values are The predicted values are clustered around the line where they are equal to the actual target value (hereinafter referred to as the actual = predicted line).
[0073] 7A and 7B illustrate the results from the 70% point in FIG. 5 (i.e., data 30% were removed as outliers). In Figures 7A and 7B, the outlier bias reduction was The model between Figures 7A and 7B is shown to remove the points furthest from the predicted line. The large variation in model accuracy indicates the process by which this dataset is modeled. This is what we are doing.
[0074] 8A and 8B show the results from the 50% point in FIG. 5 (i.e., the 50% of the data). % of the data were removed as outliers). In this case, about half of the data were identified as outliers. Even with this much variability removed from the dataset, the model still performs as shown in Figure 8. In A, we do not strictly describe the random data set. Actual = Predicted The general variability is in Figures 6A and 7A when considering the filtered data in each case. Figure 8B shows that when 50% of the variability is removed, the model is able to The performance shown in Figure 5 shows that we were able to generate prediction results that closely matched the data. In addition to performance-based analysis, analysis of these types of visual plots allows the analyst to It can be used to assess the quality of real datasets in implementations of 5, 6A, 6B, 7A, 7B, 8A and 8B illustrate visualization plots. is based on performance metric trends corresponding to various bias metric values. The analysis involves the analysis of bias criteria, such as model coefficient trends corresponding to various bias criteria selected by the analyst. The value may be based on other variables.
[0075] Various embodiments provide a system for reducing outlier bias in a target variable measured for a facility. An example of such an embodiment is shown in Figure 9. The system illustrated in Figure 9 is used in an industrial facility. Processing a dataset such as one that contains various performance measures for The computer unit 1012 is capable of processing the 2 is a program for implementing the dynamic outlier bias removal method disclosed herein. The system 1010 includes an input unit 1014 and a storage subsystem 1016. The input unit 1018 further includes a counter 1018 for measuring a given target variable and The measurement device 1020 may include a measurement device that provides a corresponding data set. It can be configured to measure a target variable, e.g., The number of parts leaving an industrial plant or refineries produced by a refinery per unit of time. Alternatively, several target variables can be measured simultaneously. In this embodiment, the measurement device 1020 includes a sensor 1022. Those skilled in the art will recognize that this is within the scope of the present invention. The various physical attributes of the substance and / or the properties produced by or present in the industrial facility. It can be seen that this includes various sensors used to measure the components used in the can detect and quantify chemicals such as greenhouse gas emissions. Additionally, those skilled in the art will appreciate that measuring the target variable of interest involves collecting, receiving, and It can be seen that this includes any means of acquiring, measuring, storing and processing the target variable, data set and data include industrial process data, computer system data, financial data, economic data, data, stock, bond and futures data, internet search data, security data, Human identification data such as voice, cloud data, big data, insurance data, and other data of interest This disclosure and the present invention may include all types of data, including but not limited to: The scope and implications of are not limited to the type of target variable, data set, or data. If the sensors and measuring devices are connected to computers, computer systems and processors, It is also understood that the system 1010 may be or include the processed The output unit 1024 can output data. The device includes a printer or transmitter (not shown).
[0076] In one embodiment, the system 1010 activates the sensor 1022. 2 then performs the detection and quantification of a given compound, such as carbon dioxide. Measurements can be performed continuously or in discrete time steps. A data set is generated and stored in the storage subsystem 1016 and The data set is input to the storage subsystem 1012. Outlier bias removal is processed by a computer program and various methods disclosed herein are Once the computer program has completed its data processing, it is terminated according to the embodiment. The processed data is then output by the output unit 1024. In embodiments where 24 is a monitor or printer, the results are visualized in a diagram. In one embodiment where unit 1024 includes a transmitting device, the processed data is stored in a central database. The data is then sent to a data center or control center where it is further processed (not shown). ). Thus, the systems according to various disclosed embodiments provide automated methods for reducing outlier bias. This provides a powerful tool for comparing different facilities within a company or technology area in a systematic way.
[0077] In a preferred embodiment, the measurement device 1020 comprises one or more sensors for detecting and quantifying chemicals. Due to global warming, greenhouse gas emissions from facilities are becoming increasingly important. Facilities that emit small amounts of greenhouse gases are more likely to produce greenhouse gases than facilities that emit large amounts. However, the latter has better overall productivity. Gases include, for example, carbon dioxide (CO2), ozone (O3), water vapor (H2O), and hydrochloric acid. Hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), chlorofluorocarbons Carbon dioxide (CFC), sulfur hexafluoride (SF6), methane (CH4), nitrous oxide (N2O), Carbon monoxide (CO), nitrogen oxides (NO x ) and non-methane volatile organic compounds (NMVOCs) ) Automated detection and quantification of these compounds is essential for achieving the desired greenhouse gas emissions allowance. However, the dynamic outlier bias By applying the removal, abnormal situations in production, such as operational errors or even accidents, can be prevented. That is, the use of the various embodiments disclosed herein results in the removal of outliers. Once an industry standard is developed, the system is used to compare emissions with the standard.
[0078] Those skilled in the art will further appreciate that the scope of the present invention does not include the use of equity securities (e.g., common stock) or derivatives. related to financial instruments such as financial contracts (e.g., forwards, futures, options, swaps, etc.) The approach of various disclosed embodiments for reducing outlier bias in a target variable For example, in one embodiment, system 1010 includes an input unit 1018 for receiving data relating to financial instruments such as common stock. , giving the corresponding data set. The target variable can be the stock price. Furthermore, the target The variables related to the variables may be used in conjunction with various well-known methods of valuing financial instruments, e.g., discounted cash Such relevant variables can be determined using methods such as flow analysis. dividends, retained earnings, or cash flow, earnings per share, price-to-earnings ratio or Once the database of target values and related variable values is created, Various embodiments of the dynamic outlier bias removal disclosed herein can be applied to the database to Accurate models for valuing financial instruments can be obtained.
[0079] The foregoing disclosure and description of the preferred embodiments of the present invention are illustrative and exemplary only and will be understood by those skilled in the art. Various changes may be made in the details of the exemplary systems and methods without departing from the scope of the invention. It is understood that variations may be made.
Claims
1. 1. A system comprising: a computer unit for processing at least one industrial process data set; an input unit in the computer unit for inputting the at least one industrial process data set, the input unit including a measurement device for measuring at least one target variable for the industrial process facility and providing a data set corresponding thereto; and an output unit for outputting a processed industrial process data set. a computer program stored in the computer unit, which when executed causes the computer unit to: selecting the at least one target variable for the industrial process facility; selecting a set of a plurality of actual values of the at least one target variable; identifying a plurality of variables associated with the at least one target variable for the industrial process facility; acquiring the at least one industrial process data set for the industrial process facility; obtaining a single bias metric; obtaining a set of model coefficients; (1) applying the set of model coefficients to the at least one industrial process data set to generate a set of predicted values; (2) generating an error set for the at least one industrial process data set using the set of predicted values and the set of actual values; (3) generating a set of error thresholds based on the error set and the bias criterion; (4) modifying the set of model coefficients by filtering the at least one industrial process data set based on the error set and the set of error thresholds to generate a set of new model coefficients; (5) repeating steps (1) through (4) using the set of new model coefficients until a performance termination criterion is met; and a computer program including instructions to execute the Including, The system, wherein the set of errors generated using the set of predicted values and the set of actual values includes a set of absolute errors and a set of relative errors.
2. The system of claim 1 , wherein the computer unit for processing the at least one industrial process data set includes a measurement device for measuring industrial process performance for the industrial process facility.
3. The system of claim 1 , wherein the computer unit includes a processor and a storage subsystem.
4. The system of claim 2 , wherein the measurement device comprises one or more sensors.
5. The system of claim 4 , wherein the one or more sensors detect and quantify chemicals for the industrial process facility.
6. The system of claim 1 , wherein the at least one target variable is related to at least one physical attribute of a substance or ingredient used in or produced by the industrial process facility.
7. The system of claim 1 , wherein the at least one target variable comprises an emission rate of at least one selected from a plurality of greenhouse gases within the industrial process facility.
8. The system of claim 1 , wherein the at least one industrial process data set includes a plurality of values for the plurality of variables.
9. 1. A computer-implemented method for measuring industrial process performance for an industrial process facility, comprising: at least one computer unit processing at least one industrial process data set; at least one computer unit inputting the at least one industrial process data set to be processed into at least one input unit, the at least one input unit being a measurement device measuring at least one target variable of a plurality of target variables for the industrial process facility and providing a data set corresponding thereto; at least one output unit outputs the processed industrial process data set; At least one computer unit stores at least one computer program in a storage subsystem; Including, The at least one computer program, when executed, causes the at least one computer unit to: selecting the at least one target variable for the industrial process facility; selecting a set of a plurality of actual values of the at least one target variable; identifying a plurality of variables associated with the at least one target variable for the industrial process facility; acquiring the at least one industrial process data set for the industrial process facility, the at least one industrial process data set including a plurality of values for the plurality of variables; obtaining a single bias metric; obtaining a set of model coefficients; (1) applying the set of model coefficients to the at least one industrial process data set to generate a set of predicted values; (2) generating an error set for the at least one industrial process data set using the set of predicted values and the set of actual values; (3) generating a set of error thresholds based on the error set and the bias criterion; (4) modifying the set of model coefficients by filtering the at least one industrial process data set based on the error set and the set of error thresholds to generate a set of new model coefficients; (5) repeating steps (1) through (4) using the set of new model coefficients until a performance termination criterion is met; and instructions to execute The computer-implemented method, wherein the set of errors generated using the set of predicted values and the set of actual values includes a set of absolute errors and a set of relative errors.
10. The computer-implemented method of claim 9 , wherein the at least one computer unit includes a processor and a storage subsystem.
11. The computer-implemented method of claim 9 , wherein the measurement device includes one or more sensors.
12. The computer-implemented method of claim 11 , wherein the one or more sensors detect and quantify chemicals for the industrial process facility.
13. 10. The computer-implemented method of claim 9, wherein the at least one target variable is related to at least one physical attribute of a substance or component used in or produced by the industrial process facility.
14. 10. The computer-implemented method of claim 9, wherein the at least one target variable comprises emissions of at least one selected from a plurality of greenhouse gases in the industrial process facility.
15. 10. The computer-implemented method of claim 9, wherein the at least one industrial process data set comprises a plurality of values for the plurality of variables.
Citation Information
Patent Citations
Adapting method for engine control parameter and its system
JP2004068729A
Integrated circuit device abnormality detection apparatus, method and program
JP2008166644A
Network performance prediction system, network performance prediction method, and program
JP2009253362A
Working hour estimation device, method, and program
JP2010250674A
JPP7244610B