A method for optimizing the quality of real-time production data in power enterprises
By optimizing the real-time data quality of power companies' production through data grading and correcting self-learning neural network models, the data quality problems caused by equipment aging have been solved, achieving efficient monitoring of power companies' production and operation and a comprehensive improvement in data quality.
Patent Information
- Application Number
- CN202310917437.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-07-25
AI Technical Summary
Existing technologies cannot effectively monitor and optimize the quality of real-time production data of power companies, especially when equipment is aging and system reliability is declining, leading to data deviations, distortions, and omissions, which affect the accuracy and reliability of production operations.
By employing data grading, constructing a basic data quality model, and a data optimization model based on a modified self-learning neural network, we comprehensively monitor and optimize real-time data quality from four aspects: completeness, standardization, timeliness, and accuracy through step-by-step monitoring and optimization. We also combine neural networks for data optimization and manual judgment correction.
It enables highly accurate judgment and improvement of important real-time production data, enhances data quality, and ensures the effectiveness and security of enterprise production and operation.
Smart Images

Figure CN117033356B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of online data quality monitoring, and specifically to a method for optimizing the real-time production data quality of power enterprises. Background Technology
[0002] Currently, real-time data from power company production and operation processes is primarily acquired automatically through various sensors at the production site. This data is then further integrated into various enterprise information systems and big data platforms via data acquisition instruments, on-site automated control systems (such as DCS), and information reporting modules. Over time, these sensors and automation / information systems may experience aging or malfunctions, insufficient accuracy, or decreased reliability. This can lead to data deviations, distortions, instability, and even data omissions or significant drift, severely impacting the effectiveness and accuracy of production and operation process control. Therefore, continuous automatic online monitoring, analysis, and optimization of real-time production data are necessary to enable enterprise operations management personnel to promptly address data anomalies, continuously improve data quality, and ultimately enhance the accuracy and reliability of production and operation process control.
[0003] Existing methods for real-time data quality monitoring typically involve automatically determining the data quality of collected real-time data from power companies and issuing alarms to remind relevant personnel to correct abnormal data quality. However, these methods only consider data timeliness (due to untimely updates) and limited accuracy (based on normal value ranges) when assessing data quality. They are only applicable to data quality checks under stable operating conditions and are not very accurate when assessing data quality under complex and variable operating conditions.
[0004] The method designed in this invention combines basic data quality model assessment with neural network-based modeling and data optimization. It comprehensively monitors and optimizes real-time data quality in four aspects: integrity, standardization, timeliness, and accuracy, achieving focused and comprehensive coverage in the data quality improvement process. It can perform basic quality analysis and assessment on general data, as well as perform higher-level calculations and optimizations on core, important, and relevant data, enabling more accurate and detailed judgment and improvement of data quality. Summary of the Invention
[0005] The purpose of this invention is to design a method for optimizing the quality of real-time production data in power enterprises. This method includes basic preparation (including data classification and grading and basic data quality model design), data acquisition, data screening and cleaning (including historical data screening and cleaning, basic data quality judgment of target data, and data labeling), data correlation analysis, construction / verification of a modified input neural network (MINN) data optimization model based on historical data, data optimization based on the MINN model, result verification, and data correction. This method collaboratively improves the quality of real-time production data in power enterprises from three directions: basic data quality model judgment, machine learning model-based data optimization, and manual judgment correction. This comprehensively enhances the reliability and credibility of core real-time production data, achieving a comprehensive improvement in data quality.
[0006] The present invention discloses a method for optimizing the quality of real-time production data of power enterprises, which executes the following steps S1-S8 to complete the verification and quality optimization of power enterprise production data:
[0007] Step S1: Classify and grade the various types of production data of the power company. Classification includes dividing the various types of production data into automatically collected data and manually entered data. Grading includes grading the importance of the various types of production data according to their scope of influence, number of direct references, and importance coefficient.
[0008] Step S2: Construct a basic data quality model for evaluating the quality of production data. The basic data quality model defines the integrity score, standardization score, timeliness score, and accuracy score for various types of production data, and calculates the total data quality score for each type of production data based on the four scores.
[0009] Step S3: Based on the importance classification of various types of production data, collect historical data and real-time data of automatically collected data and manually entered data of preset levels respectively, and store the historical data in the temporary historical data table of the real-time database, and store the real-time data in the real-time database.
[0010] Step S4: Clean and filter historical data, determine the quality of real-time data based on the basic data quality model, determine whether real-time data is abnormal or normal based on the calculation results of the basic data quality model, and record abnormal data in the basic data quality abnormal data details table.
[0011] Time-stamp alignment is performed on historical data that has been cleaned and filtered, as well as real-time data that has been determined to be normal.
[0012] Step S5: For the historical data and real-time data obtained in step S4, respectively, use principal component analysis to perform correlation analysis and remove production data that has no correlation with other production data.
[0013] Step S6: Construct a data optimization model based on a self-learning neural network. Use the historical data obtained in step S4 as input training samples, modify the training samples, and use the target value and adjustment amount of the training samples as outputs to iteratively train the data optimization model to obtain a well-trained data optimization model.
[0014] Step S7: Input the trained data tuning model with the real-time data obtained in step S4, tune the real-time data based on the data tuning model, use the output of the data tuning model as the tuning data, and determine the quality of each tuning data.
[0015] Step S8: After completing the calculation and quality judgment of a batch of optimization data, calculate the relative error percentage between the optimization data and the original collected data of all production data, sort and record them in descending order, and complete the verification and quality optimization of the power company's production data.
[0016] Beneficial effects: In response to the problem of declining real-time data quality in many power companies due to aging data acquisition and transmission equipment and systems, various power companies currently use simple rule-based judgment methods to check and correct data quality. For example, they use simple threshold range rules to determine data accuracy or certain special specifications to determine the validity of specific data. However, these methods have limited applicability and low accuracy, and cannot meet the needs of efficiently checking and optimizing massive amounts of real-time data.
[0017] This invention designs a method for optimizing the quality of real-time production data in power enterprises. Through data grading, constructing a basic data quality rule model, and a data optimization model based on a modified self-learning neural network, it comprehensively monitors and optimizes real-time data quality in four aspects: integrity, standardization, timeliness, and accuracy. This enables data quality verification and online correction of important real-time production data. Based on the results, it alerts operation and management personnel to locate and analyze problematic data and their possible causes, and handles data anomalies using appropriate methods. This achieves a comprehensive improvement in data quality, ensuring the effectiveness of normal production and operation management for the enterprise. Attached Figure Description
[0018] Figure 1 This is a flowchart of a method for optimizing the quality of real-time production data in power enterprises according to an embodiment of the present invention;
[0019] Figure 2 This is a schematic diagram of metadata analysis for indicator X provided in an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of the MINN network structure provided according to an embodiment of the present invention. Detailed Implementation
[0021] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0022] Currently, most power companies have implemented real-time data collection and online monitoring of production data, and have launched production operation management systems to help improve their operational management. However, over time, due to the aging of on-site equipment and the decline in system reliability, the real-time data collected from various enterprise information systems and big data platforms suffers from a decline in data quality, including incompleteness, standardization, timeliness, and accuracy. This can range from affecting the accuracy of daily management to potentially causing malfunctions in on-site control systems or even triggering various production safety accidents leading to shutdowns and losses, significantly impacting the company's normal operations. At present, most power companies monitor and manage the data quality of collected real-time data primarily through simple threshold range rules for data verification and validation. This approach has limited applicability, significant limitations, and low accuracy, failing to meet the needs for efficient data quality optimization.
[0023] The purpose of this invention is to design a method for optimizing the quality of real-time production data in power enterprises. By classifying data, constructing a basic data quality rule model, and a data optimization model based on a modified self-learning neural network, the method enables data quality verification and online correction of important real-time production data. Based on the results, the method alerts operation and management personnel to locate and analyze problematic data and their possible causes, and handles data anomalies through appropriate methods, thereby achieving a comprehensive improvement in data quality and ensuring the effectiveness of normal production and operation management of the enterprise.
[0024] This invention provides a method for optimizing the quality of real-time production data in power enterprises, referring to... Figure 1 Perform the following steps S1-S8 to complete the verification and quality optimization of power company production data:
[0025] Step S1: Classify and grade the various types of production data of the power company. Classification includes dividing the various types of production data into automatically collected data and manually entered data. Grading includes grading the importance of the various types of production data according to their scope of influence, number of direct references, and importance coefficient.
[0026] First, it's necessary to define the scope of data quality optimization. Select a portion of important real-time production data from the enterprise's massive data / metric library to conduct data quality optimization. This can be done by classifying and grading data to identify high-priority data objects. The specific process is as follows:
[0027] (1) Data Classification
[0028] This invention categorizes the target data for data quality improvement into two main types: automatically collected data and manually entered data. The definitions are as follows:
[0029] Automatic data acquisition: This includes data collected at different time frequencies such as seconds, minutes, and hours. It is mainly obtained automatically through various sensors, such as various voltage, current, power, real-time power generation, pressure, temperature, flow rate, and real-time calculation efficiency data.
[0030] Manually entered data: Data that cannot be obtained through automatic collection is supplemented by manual entry. This includes data that is entered periodically (such as fuel calorific value, liquid level that cannot be collected, etc.) and data that is entered initially (such as design values of various performance parameters of equipment or system, etc.).
[0031] (1) Data classification
[0032] The impact of data on various business areas of power companies (such as finance, operations, human resources, etc.) is based on the scope N. all , Number of direct citations N y The importance coefficient α is used to classify the two types of data mentioned above. The level index S and data level of each data item in the two types of data are calculated respectively (assuming that there are three levels in total).
[0033] The level index S, used to characterize the importance level of various types of production data, is defined as follows:
[0034] S=k1*N all +k2*N y +k3*α
[0035] In the formula, N all N represents the scope of influence of production data. y The value represents the number of direct citations, α represents the importance coefficient, and k1, k2, and k3 are the influence factors of the production data's scope of influence, number of direct citations, and importance coefficient, respectively, with a custom value range of 0 to 1.
[0036] For both automatically collected and manually entered production data, the data for each category is sorted in descending order of level index S to obtain the automatically collected data sequence {S1}. i {i = 1, 2, ..., m} and manually entered data {S2} iLet m = 1, 2, ..., n, where m is the number of types of production data in the automatically collected data sequence and n is the number of types of production data in the manually entered data sequence; each of these is divided into three equal parts to obtain each subsequence {S1}. j1}、{S1 j2}、{S1 j3} and {S2 j1}、{S2 j2}、{S2 j3 The importance levels of the data are Level 1, Level 2, and Level 3, with Level 1 being the highest and Level 3 the lowest. After completing the data classification and grading, priority online data quality optimization can be performed on the higher-level data (Level 1) as needed.
[0037] The above N all N y The calculation process for α is explained below:
[0038] Based on the classification, natural attributes, and management attributes of production data indicators, an enterprise indicator dictionary is developed. Then, based on this dictionary, the influence range N of all indicators is calculated using a metadata analysis module or manually. all and the number of direct citations N y Secondly, based on the actual application situation of each business domain of the enterprise, a coefficient α is assigned to the business importance of all data.
[0039] (1) Design of indicator dictionary
[0040] The reference design of the enterprise indicator dictionary definition table is shown in the table below. Using this table, and in conjunction with the company's indicator system, a detailed indicator dictionary table for each business area can be obtained.
[0041]
[0042]
[0043] (2) Metadata Analysis
[0044] Impact analysis and correlation analysis were performed on each data point to calculate the number N of directly influential data points at all downstream levels. zj and the number of data points N that are indirectly affected jj Further calculate the impact range N of the production data. all and the number of direct citations N y The specific formula is as follows:
[0045] N all =N zj +N jj
[0046] N y =N jj
[0047] Direct impact refers to all impact data at the first level downstream of the current production data, while indirect impact refers to all impact data at the second level and all other levels downstream of the current production data. (See diagram below.) Figure 2 , Figure 2 In the middle, the N of index X zj =3, N jj =7, therefore N all =10, N y =3.
[0048] (3) Importance coefficient
[0049] Each business domain assigns an importance coefficient α to all indicators within its scope according to their degree of business importance. Generally, important financial and business indicators that affect or reflect the company's performance are assigned higher values, and vice versa. The company's performance indicator model can generally be used as a reference for assigning the importance coefficient α.
[0050] Step S2: Construct a basic data quality model for evaluating the quality of production data. The basic data quality model defines the integrity score, standardization score, timeliness score, and accuracy score for various types of production data, and calculates the total data quality score for each type of production data based on the four scores.
[0051] In step S2, basic data quality judgment rules are preset, and the total data quality score F for each type of production data is calculated. 总 As shown in the following formula:
[0052] F 总 =μ1*F w +μ2*F g +μ3*F j +μ4*F z =μ1*F w +μ2*∑(F gx *k x )+μ3*F j +μ4*F z
[0053] In the formula, F w For completeness score, F g For normative scoring, F j For timeliness, F z For accuracy, F gx For each sub-category of normative type, k xThe weights for each subdivided normative type are given, with μ1 to μ4 representing the score coefficients for the four dimensions, respectively. μ1 + μ2 + μ3 + μ4 = 1, and μ1 to μ4 are typically set to 0.25, but can be customized. The scores for the four dimensions are calculated as follows:
[0054] F = 100 - Δ * N w
[0055] In the formula, Δ represents the deduction for violating the preset basic data quality judgment rules, and N represents the deduction for violating the basic data quality judgment rules. w The number of times the basic data quality assessment rules have been violated.
[0056] The basic data quality assessment rules are as follows. If a data does not meet the assessment rules, points will be deducted accordingly:
[0057] 1) Data integrity rule: Whether null values are allowed.
[0058] 2) Data normative rules can be further refined into the following subcategories:
[0059] Data types: numeric, character, date, etc.
[0060] Unit: Standard unit of measurement for data
[0061] Can it be negative: Yes / No
[0062] Can it be 0: Yes / No
[0063] Decimal places: Non-negative integers
[0064] other
[0065] 3) Data timeliness rules: data update time.
[0066] 4) Data accuracy rules: value range rules, which set upper and lower limits for data.
[0067] Step S3: Based on the importance classification of various types of production data, collect historical data and real-time data of automatically collected data and manually entered data of preset levels respectively, and store the historical data in the temporary historical data table of the real-time database, and store the real-time data in the real-time database.
[0068] Select high-importance production data from both automatically collected and manually entered data, and initiate a data collection task. Data collection includes historical data collection and real-time data collection.
[0069] (1) Historical data collection
[0070] Historical data can be collected from a time-series database using a general OPC interface or other interfaces. The time range can be selected to be within the past year, and the data collection interval can be set to per second, per minute, or other parameters as needed. The historical data collection cycle can be set to be executed weekly or monthly. After each execution, the historical data collected this time will be stored in a temporary historical data table in the real-time database, overwriting the historical data collected last time.
[0071] The collected historical data is mainly used for subsequent step S6: the construction / verification of the data optimization model of the self-learning neural network based on historical data, and provides a model basis for step S7: data optimization based on the self-learning neural network model.
[0072] (2) Real-time data acquisition
[0073] Real-time data is collected from the data source automation system using a general OPC interface or other interfaces and stored in a real-time database. The collected real-time data is used for further data quality assessment and data optimization. Generally, it can be stored for three months, which is longer than the data quality assessment and data optimization cycle. After the storage period expires, the data will be automatically deleted from the real-time database.
[0074] Step S4: Clean and filter historical data, determine the quality of real-time data based on the basic data quality model, determine whether real-time data is abnormal or normal based on the calculation results of the basic data quality model, and record abnormal data in the basic data quality abnormal data details table.
[0075] (1) Historical data filtering and cleaning
[0076] In step S3, historical data is collected in batches at preset collection cycles. In step S4, the historical data in each batch is cleaned and filtered according to the following rules. All of the following rules must be met. If any one rule is not met, all historical data collected in that batch will be discarded:
[0077] 1) The production system is operating in a steady or quasi-steady state;
[0078] Steady state: The rate of change of all core operating parameters of the system is less than their respective reference rate of change (e.g., the rate of change of main steam temperature is less than 0.4% / second) and lasts for a certain period of time (e.g., 1 minute or more);
[0079] Quasi-steady state: The rate of change of all core operating parameters of the system is within their respective reference rate of change range (e.g., the rate of change of main steam temperature is less than 1% / second and greater than or equal to 0.4% / second), and lasts for a certain period of time (e.g., more than 1 minute);
[0080] 2) All historical data do not violate the basic data quality assessment rules;
[0081] 3) All the historical data correlation logics are reasonable and without anomalies;
[0082] Qualified historical data that meets all the above conditions will be further transferred and stored in the historical data table of the real-time database.
[0083] (2) Basic data quality determination and data marking of target real-time data
[0084] Verify the real-time data using the basic data quality determination rules. Real-time data that does not meet the basic data quality rules is determined as abnormal data, recorded in the detailed list of basic data quality abnormal data, and relevant information is sent to relevant personnel for warning notifications via email or text messages, etc. Data that meets all the basic data quality rules is normal data. The detailed list of basic data quality abnormal data includes information such as data name, data type, data level, data collection time, each dimension of the data (such as business scope, date, type, etc.), total data quality score, and sub-item scores of each data quality element.
[0085] (3) Time scale alignment
[0086] For the historical data that has been cleaned and screened, as well as the real-time data determined as normal data, perform time scale alignment operations. According to the timestamps of each production data, perform time scale alignment at integer minute moments to obtain the data values of all production data at a fixed interval time (such as 1 min). The time scale alignment calculation method uses the equal-proportion linear estimation method, described as follows:
[0087] Let the data values of a certain measurement point at non-standard times t1 and t2 be x1 and x2, then for the standard time t0 (t1 < t0 < t2), the calculated standard data x0 is:
[0088]
[0089] Step S5: Respectively for the historical data and real-time data obtained in step S4, use the principal component analysis method to perform correlation analysis, and eliminate the production data that has no correlation with other production data;
[0090] In step S5, respectively for the historical data and real-time data obtained in step S4, perform standardization processing, and use the PCA principal component analysis method to calculate the covariance between each production data of the historical data and real-time data respectively, and eliminate the production data whose covariance with other production data is all 0.
[0091] Correlation analysis is performed on the historical data that has been filtered and cleaned in step S4 to provide a training dataset for building a modified self-learning neural network data tuning model. The target data identified as normal in the above steps undergoes further correlation analysis, and then the created modified self-learning neural network data tuning model is used for data tuning.
[0092] There are many methods for data correlation analysis. In this invention, Principal Component Analysis (PCA) is used to perform correlation analysis on the target data. PCA transforms most of the highly correlated variable data into new, independent variable data (i.e., principal components).
[0093] 1) Standardization Processing
[0094] Mapping the original collected data to the interval [-1, 1], its standardized transformation is as follows:
[0095] X = 2(X' - X') min ) / (X' max -X' min )-1
[0096] Where X' represents the collected and measured value of the original variable data, X' min and X' max These are the upper and lower limits of the original variable data X'.
[0097] 2) PCA Calculation
[0098] Suppose that the m values of the original n target data variables are represented as the following matrix X after standardization:
[0099]
[0100] Its covariance matrix is Y:
[0101]
[0102] Where cov(Xi,Xj) represents the covariance of the i-th and j-th columns of data in X.
[0103] Eigenvalue decomposition of Y yields Y = U∧U T
[0104] Where: ∧ is the eigenvalue λ of Y. i Diagonal array
[0105] The cumulative contribution rate R of the first h variables h for:
[0106] R h = (λ1+λ2+...+λ) h) / (λ1+λ2+...+λ n )
[0107] Based on the cumulative contribution rate threshold R0 (which can be taken as 0.7 to 0.9) <R h The number of principal components, h, can be determined. Furthermore, by using the covariance matrix for Y, data variables whose covariance with all other variables is close to 0 are selected. These variables are uncorrelated with other data and can be excluded from subsequent data tuning, model training, and tuning steps, resulting in the final number of data variables, n'.
[0108] Step S6: Construct a data optimization model based on a modified self-learning neural network (MINN). Use the historical data obtained in step S4 as input training samples, modify the training samples, and use the target value and adjustment amount of the training samples as outputs to iteratively train the data optimization model to obtain a well-trained data optimization model.
[0109] The prototype of MINN is Input Neural Network (INN), which has good applicability in dimensionality reduction and data error detection for nonlinear systems. By locally modifying it to form the MINN described in this invention, the stability and convergence of the network can be improved to a certain extent, thereby improving the accuracy of model application.
[0110] Following the steps outlined above, a MINN model is trained using well-filtered and cleaned historical data. The MINN model consists of three layers: an input layer, intermediate hidden layers, and an output layer. The input layer corresponds to the principal components of the collected data variables; therefore, its number of nodes is N. in Less than the output layer N out The input and output layers use linear activation functions, while the intermediate layers use the non-linear activation function Logistic. See the schematic diagram of the MINN network structure below. Figure 3 .
[0111] (1) Determine the network structure
[0112] This includes the number of input nodes, the number of intermediate hidden layer nodes, and the number of output nodes.
[0113] Number of input data nodes N in =h, which is the number of principal components of all measurement point data;
[0114] Output number of nodes N out =n', which is the final number of data variables collected;
[0115] Number of intermediate hidden layer nodes Rounding;
[0116] (2) Initialize the network parameters of MINN
[0117] That is, the input value matrix and the network weight matrix, and the initial values of each data item are randomly selected in the range of [-1, 1].
[0118] (3) Network training
[0119] Suppose there are N training history data points after filtering, cleaning, and saving. X For each batch, the original training objective function for INN is:
[0120]
[0121] Where X' ij and X ij These are the network output value of the j-th generation in the i-th batch and the measurement value of the training sample, respectively.
[0122] The training objective of the above formula is to minimize the sum of the differences between the output target value and the actual value of all batches and measurement point data. However, since the training process does not restrict the network input values, the range of input values may be abnormally large, affecting the stability of the network. This invention makes certain modifications by adding a constraint term for the input parameters to the objective function, so that the training values of the network input parameters fall within a reasonable range, thereby improving the stability of the network to a certain extent. In step S6, during the iterative training of the data optimization model, the training objective function is as follows:
[0123]
[0124] Among them, A ij’ Let μ be the input value for the j'th network node in the i-th batch, μ be the input parameter weight coefficient, and p1, p2, and p3 be three different constraint parameters for the input value of the data optimization model. Their values are related to the model structure and the number of training samples. Different data can be used to train the model multiple times to compare the convergence results. The most suitable value is selected based on the actual model training effect to ensure that the model can converge to the given range well and quickly.
[0125] The network is iteratively trained using the conventional backpropagation (BP) algorithm (i.e., gradient descent). The estimated network output and adjusted network input values are calculated for each set of sample data. The adjusted network weights are then calculated and updated. This calculation process is repeated multiple times until the objective function converges to a given range. Finally, the final network weights are determined, and the data-optimized model is established.
[0126] Its model can be trained and updated regularly based on the latest historical data collected to continuously improve and optimize it, and version control of the model can be achieved.
[0127] Step S7: Input the trained data tuning model with the real-time data obtained in step S4, tune the real-time data based on the data tuning model, use the output of the data tuning model as the tuning data, and determine the quality of each tuning data.
[0128] Unlike conventional backpropagation (BP) networks, the MINN model's computation process does not involve obtaining the output from the input in a single step. Instead, it requires multiple iterative calculations, similar to network training, to minimize the validation objective function and confirm reasonable network input values. Compared to network training, this reduces the computation of network weights. In step S7, during the optimization of real-time data based on the data optimization model, the optimization objective function is as follows:
[0129]
[0130] The data collected in step S4, marked as normal, is then used to select the target data for optimization based on a self-learning neural network. This data is then input into the trained MINN data optimization model for calculation, yielding the final optimized data. The calculation method can be set to either real-time calculation or scheduled batch calculation. If the number of data points is small (e.g., less than 10), real-time calculation can be used. If the number of data points is large (e.g., more than 10), considering system efficiency, scheduled batch calculation is recommended.
[0131] Let the set of measured values of a certain batch of real-time data be {X}. i The set of output values obtained by the data optimization model for the set of numbers i = 1, 2, ..., n is {X}. i If for any i, |X, ..., n}, then |X, ..., n} i '-X i |≤Ω i Ω i If the residual threshold corresponding to the i-th data variable in the model training result is specified under the given confidence limit, then the optimization data obtained in this calculation is qualified. Otherwise, the output target data items in the optimization objective function that do not meet the condition need to be corrected and recalculated until all measurement results meet the condition. At this time, the model data value can be regarded as the final model optimization data value for each data.
[0132] Finally, the obtained model tuning values are inversely standardized to recover the true tuning values of the data with dimensions, and the final results are included in the tuning calculation data table of the target data.
[0133] Step S8: After completing the calculation and quality judgment of a batch of optimization data, calculate the relative error percentage between the optimization data and the original collected data of all production data, sort and record them in descending order, and complete the verification and quality optimization of the power company's production data.
[0134] After completing a MINN model data optimization calculation, the relative error percentage between the optimized values and the original collected values for all data in all batches is further calculated, and the data is sorted and recorded in descending order. Alerts are pushed to relevant operations and management personnel for the top 10%–20% of data with the largest changes during this MINN model optimization. Based on manual confirmation, the corrected values are either accepted or ignored. Ignored data has invalid optimization results and is not recorded or statistically analyzed as abnormal data.
[0135] The system comprehensively records valid data showing significant changes in MINN data optimization calculation results and valid data deemed abnormal based on basic data quality criteria, forming an abnormal data information detail table. The definitions of each field are as follows:
[0136] field name Data types illustrate id string Data Number name string Data Name data_type string Data types level string Data Level model_cal string Participation in data optimization: Yes / No time datetime Data collection time col_type string Data acquisition method: Automatic / Manual ori_value float Collected data values cal_value float Model tuning values deviation float Optimization calculation relative deviation: % source_name string Data source name remark string Remark
[0137] Locate and confirm the data sources of any abnormal data, and remind staff to conduct precise point-to-point checks. Eliminate potential data quality risks by upgrading system hardware and software, communication links, or replacing sensors, and continuously improve data quality.
[0138] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A method for optimizing the quality of real-time production data in power enterprises, characterized in that, Perform the following steps S1-S8 to complete the verification and quality optimization of power company production data: Step S1: Classify and grade the various types of production data of the power company. Classification includes dividing the various types of production data into automatically collected data and manually entered data. Grading includes grading the importance of the various types of production data according to their scope of influence, number of direct references, and importance coefficient. Step S2: Construct a basic data quality model for evaluating the quality of production data. The basic data quality model defines the integrity score, standardization score, timeliness score, and accuracy score for various types of production data, and calculates the total data quality score for each type of production data based on the four scores. Step S3: Based on the importance classification of various types of production data, collect historical data and real-time data of automatically collected data and manually entered data of preset levels respectively, and store the historical data in the temporary historical data table of the real-time database, and store the real-time data in the real-time database. Step S4: Clean and filter historical data, determine the quality of real-time data based on the basic data quality model, determine whether real-time data is abnormal or normal based on the calculation results of the basic data quality model, and record abnormal data in the basic data quality abnormal data details table. Time-stamp alignment is performed on historical data that has been cleaned and filtered, as well as real-time data that has been determined to be normal. Step S5: For the historical data and real-time data obtained in step S4, respectively, use principal component analysis to perform correlation analysis and remove production data that has no correlation with other production data. Step S6: Construct a data optimization model based on a self-learning neural network. Use the historical data obtained in step S4 as input training samples, modify the training samples, and use the target value and adjustment amount of the training samples as outputs to iteratively train the data optimization model to obtain a well-trained data optimization model. Step S7: Input the trained data tuning model with the real-time data obtained in step S4, tune the real-time data based on the data tuning model, use the output of the data tuning model as the tuning data, and determine the quality of each tuning data. Step S8: After completing the calculation and quality judgment of a batch of optimization data, calculate the relative error percentage between the optimization data and the original collected data of all production data, sort and record them in descending order, and complete the verification and quality optimization of the power company's production data.
2. The method for optimizing the quality of real-time production data in power enterprises according to claim 1, characterized in that, The method for classifying the importance of various types of production data in step S1 is as follows: The level index S, used to characterize the importance level of various types of production data, is defined as follows: S=k1*N all +k2*N y +k3*α In the formula, N all N represents the scope of influence of production data. y α represents the number of direct citations, k1 represents the importance coefficient, and k2 and k3 represent the influence factors of the production data's scope of influence, number of direct citations, and importance coefficient, respectively. Based on the indicator classification, natural attributes, and management attributes of production data, the direct and indirect influence relationships between production data are determined, and the scope of influence N of production data is calculated. all and the number of direct citations N y The specific formula is as follows: N all =N zj +N jj N y =N jj In the formula, N zj N represents the number of data points that directly affect production data. jj This indicates the number of data points that indirectly affect production data; For both automatically collected and manually entered production data, the data for each category is sorted in descending order of level index S to obtain the automatically collected data sequence {S1}. i {i = 1, 2, ..., m} and manually entered data {S2} i Let m = 1, 2, ..., n, where m is the number of types of production data in the automatically collected data sequence and n is the number of types of production data in the manually entered data sequence; each of these is divided into three equal parts to obtain each subsequence {S1}. j1 }、{S1 j2 }、{S1 j3 } and {S2 j1 }、{S2 j2 }、{S2 j3 The importance levels of} are Level 1, Level 2, and Level 3, respectively, with Level 1 being the highest and Level 3 the lowest.
3. The method for optimizing the quality of real-time production data in power enterprises according to claim 1, characterized in that, In step S2, basic data quality judgment rules are preset, and the total data quality score F for each type of production data is calculated. 总 As shown in the following formula: F 总 =μ1*F w +μ2*F g +μ3*F j +μ4*F z =μ1*F w +μ2*∑(F gx *k x )+μ3*F j +μ4*F z In the formula, F w For completeness score, F g For normative scoring, F j For timeliness, F z For accuracy, F gx For each sub-category of normative type, k x The weights for each subdivided normative type are given, and μ1 to μ4 are the score coefficients for the four dimensions, respectively. The scores for the four dimensions are calculated as follows: F=100-Δ*N w In the formula, Δ represents the deduction for violating the preset basic data quality judgment rules, and N represents the deduction for violating the basic data quality judgment rules. w The number of times the basic data quality assessment rules have been violated.
4. The method for optimizing the quality of real-time production data in power enterprises according to claim 3, characterized in that, In step S3, historical data is collected in batches at preset collection cycles. In step S4, the historical data in each batch is cleaned and filtered according to the following rules. All of the following rules must be met. If any one rule is not met, all historical data collected in that batch will be discarded: The production system is operating in a steady or quasi-steady state; all historical data does not violate the basic data quality judgment rules; all historical data are logically correlated and without anomalies. Qualified historical data that meets all the above conditions will be further transferred to the historical data table in the real-time database. Real-time data is checked using basic data quality assessment rules. Real-time data that does not meet the basic data quality rules is judged as abnormal data, and its total data quality score F is assigned. 总 The scores of the four dimensions are recorded in the basic data quality anomaly data details table. Real-time data that meets all basic data quality rules is considered normal data. For historical data that has been cleaned and filtered, as well as real-time data that has been determined to be normal, a time-stamp alignment operation is performed. Based on the timestamp of each production data, the time stamps are aligned according to integer minutes to obtain the data values of all production data at fixed intervals.
5. The method for optimizing the quality of real-time production data in power enterprises according to claim 1, characterized in that, In step S5, the historical data and real-time data obtained in step S4 are standardized, and the PCA principal component analysis method is used to calculate the covariance between each production data in the historical data and real-time data, and production data whose covariance with each other is 0 are eliminated.
6. The method for optimizing the quality of real-time production data in power enterprises according to claim 1, characterized in that, In step S6, during the iterative training of the data-optimized model, the training objective function is as follows: Among them, A ij’ For the j'th network node in the i-th batch, μ is the input value, p1, p2, and p3 are three different constraint parameters of the data tuning model input value; N x N represents the number of batches of historical training data that have been filtered, cleaned, and saved. out Output the number of nodes in the network structure, N in Input the number of data nodes for the network structure; X′ ij and X ij These are the network output value of the j-th generation in the i-th batch and the measurement value of the training sample, respectively.
7. The method for optimizing the quality of real-time production data in power enterprises according to claim 6, characterized in that, In step S7, during the optimization of real-time data based on the data optimization model, the optimization objective function is as follows: Let the set of measured values of a certain batch of real-time data be {X}. i The set of output values obtained by the data optimization model for the set of numbers i = 1, 2, ..., n is {X}. i If for any i, |X, ..., n}, then |X, ..., n} i '-X i |≤Ω i Ω i If the residual threshold corresponding to the i-th data variable in the model training result is specified under the given confidence limit, then the optimization data obtained in this calculation is qualified. Otherwise, the output target data items in the optimization objective function that do not meet the conditions need to be corrected and recalculated until all measurement calculation results meet the conditions. At this time, the model data value can be regarded as the final model optimization data value of each data.
Citation Information
Patent Citations
User security level identification method and system integrated with boosting tree constructed model, electronic equipment and medium
CN114880635A
Enterprise credit scoring method and device
CN115330526A