A data cleaning method for the intelligent monitoring panel system of a coal-fired power plant

The processing of coal-fired power plant data through CCA correlation analysis, normal distribution and K-means detection combined with improved wavelet threshold method is solved, data abnormalities and missing problems are improved, data quality is improved, and subsequent analysis is ensured.

CN119066055BActive Publication Date: 2025-07-22JIANGXI DATANG INT XINYU NO 2 POWER GENERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411115889.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-07-22
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

In coal-fired power plants, data abnormalities, missing and noise problems seriously affect the judgment of the operating status of power equipment. A single data processing method is inefficient and has a large workload.

Method used

The CCA correlation analysis method is used for data correlation analysis, and the normal distribution and K-means method are combined to detect missing and outliers, and denoising is performed through the improved wavelet threshold method to comprehensively improve data quality.

Benefits of technology

Effectively balance noise interference with useful signals, improve data accuracy and reliability, and ensure the accuracy and reliability of subsequent analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119066055B_ABST
    Figure CN119066055B_ABST
Patent Text Reader

Abstract

The present application discloses a data cleaning method for an intelligent monitoring system of a coal-fired power plant. The method first uses CCA correlation analysis method to perform data correlation analysis, then uses the missing value detection method of normal distribution to detect and process data missing values, and the outlier detection method of K-means to detect and process outliers; finally, the processed data is denoised by an improved wavelet threshold method, so as to efficiently improve the data quality of the intelligent monitoring system of the coal-fired power plant and ensure the accuracy and reliability of subsequent data analysis (such as training / inference of prediction models).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data cleaning methods, and in particular to a data cleaning method for an intelligent monitoring panel system of a coal-fired power plant. Background Art

[0002] With the deep integration of industrial Internet of Things technology in coal-fired power plants, the associated production operation data has shown a blowout growth. The power data, which is complex in structure, diverse in types and huge in volume, presents the characteristics of multi-source heterogeneity. It is particularly important to extract valuable information from the vast amount of power production data. However, during the process of power production data collection and transmission, it is inevitably affected by the external environment, parameter equipment, and operating conditions, resulting in problems such as data anomalies, missing values, and noise, which seriously affect the judgment of power plant operation and maintenance personnel on the operating status of relevant power equipment.

[0003] The purpose of data cleaning is to accurately identify errors in data, including the detection of abnormal data and the correction of incorrect data. Currently, the data in the monitoring panel system dataset of coal-fired units is generally detected and processed by using a data missing value detection method or a data outlier detection method. The inventor realizes that a single method is not ideal in terms of data continuity, accurate fitting, and balancing the contradiction between noise interference and useful signals, and the workload is relatively large. Summary of the Invention

[0004] The present application provides a data cleaning method for an intelligent monitoring panel system of a coal-fired power plant, aiming to comprehensively improve the effect of detecting and correcting data.

[0005] The solution provided by the present application is as follows:

[0006] A data cleaning method for an intelligent monitoring panel system of a coal-fired power plant, comprising:

[0007] Obtain a monitoring panel system dataset of a coal-fired unit, where the dataset includes data in the electrical system, steam-water system, and boiler system related to the operating status of the coal-fired unit;

[0008] Perform CCA data correlation analysis on the data in the monitoring panel system dataset of the coal-fired unit to obtain the magnitude of the correlation between the state monitoring parameters of the coal-fired unit, quantitatively measure the degree of correlation between the state monitoring parameters, and provide a reference for the next step to select necessary data for missing value detection and outlier detection;

[0009] With reference to the degree of correlation between the state monitoring parameters, select necessary data from the data in the monitoring panel system dataset of the coal-fired unit for missing value detection based on normal distribution and outlier detection based on K-means, and correspondingly process missing data and abnormal data;

[0010] The processed data is denoised using the improved wavelet threshold method; the improved wavelet threshold method is implemented through the following formula:

[0011]

[0012] where ω j,k is the j-th layer and k-th wavelet coefficient, E j and E noise are the energy distributions of the j-th layer signal and noise respectively, N is the number of samples, is the mean of the j-th layer wavelet coefficients, α is the protection factor, and σ is the standard deviation.

[0013] Furthermore, the CCA data correlation analysis is implemented through the following formula:

[0014] Two groups of original variables X and Y:

[0015]

[0016] Composite variables R and S:

[0017]

[0018] The correlation coefficients of composite variables R and S are as follows:

[0019]

[0020] where U and V are transformation coefficient matrices, cov(R,S) represents the covariance of composite variables R and S, and Var(R) and Var(S) represent the variances of composite variables R and S respectively.

[0021] Furthermore, the missing value detection based on the normal distribution is implemented through the following formula:

[0022] Mean and standard deviation Maximum likelihood estimation:

[0023]

[0024] Traverse the sample set:

[0025]

[0026] Furthermore, the outlier detection based on K-means is implemented through the following formula:

[0027] 1) Arbitrarily select B elements from the dataset A as the initial clustering centers;

[0028] 2) Calculate the distances between the remaining elements and the B clustering centers, and assign them to the nearest clustering center;

[0029] 3) Recalculate the arithmetic means of the B clusters and update to generate new cluster centers;

[0030] 4) Repeat steps 2) and 3) until the clustering result no longer changes and the clustering ends.

[0031] Furthermore, the processing of missing data includes interpolation or completely removing the data points containing missing values from the dataset.

[0032] Furthermore, the processing of abnormal data includes deletion, replacement, marking, and imputation.

[0033] This application also provides a computer device, including a memory, a processor, and a computer program stored on the memory. It is characterized in that the processor executes the computer program to implement the steps of the above data cleaning method for the intelligent monitoring system of a coal-fired power plant.

[0034] This application also provides a computer program product, including a computer program / instructions. It is characterized in that when the computer program / instructions are executed by a processor, the steps of the above data cleaning method for the intelligent monitoring system of a coal-fired power plant are implemented.

[0035] Compared with the prior art, this application has at least the following beneficial effects:

[0036] This application comprehensively considers the advantages of the data correlation analysis method, the data missing value detection method, and the data outlier detection method. First, the CCA correlation analysis method is used for data correlation analysis to determine whether there is a strong correlation between various parameters, and necessary data is selected accordingly (because for some strongly correlated parameters, only their representatives need to be selected). The missing value detection method based on the normal distribution is used for the detection and processing of data missing values, and the outlier detection method of K-means is used for the detection and processing of outliers. Finally, the improved wavelet threshold method is used for denoising the processed data, which can fully balance the contradiction between noise interference and useful signals, make up for the limitation of the inconsistent local features of signals at different scales in the traditional wavelet threshold method, and achieve good denoising effects under different decomposition scales, thus ultimately efficiently improving the data quality of the intelligent monitoring system of a coal-fired power plant and ensuring the accuracy and reliability of subsequent data analysis (such as the training / inference of prediction models). BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a schematic diagram of a data cleaning method for an intelligent monitoring system of a coal-fired power plant provided by an embodiment of this application;

[0038] Figure 2 It is a comparison diagram before and after denoising using the improved wavelet threshold method in an embodiment of this application. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0040] In the description of the present application: Unless otherwise specified, expressions such as "including", "comprising", "having", etc. also mean "not limited to" (certain units, components, devices, steps, etc.).

[0041] As Figure 1 shown, this embodiment provides a method for cleaning data of an intelligent monitoring panel system of a coal-fired power plant. The method mainly includes: obtaining a dataset of the monitoring panel system of a coal-fired unit, CCA data correlation analysis, missing value detection and processing based on normal distribution, outlier detection and processing based on K-means, and denoising using an improved wavelet threshold method. Specifically:

[0042] Obtain a dataset of the monitoring panel system of a coal-fired unit, where the dataset includes data in the electrical system, steam-water system, and boiler system related to the operating state of the coal-fired unit;

[0043] Perform CCA data correlation analysis on the data in the dataset of the monitoring panel system of a coal-fired unit to obtain the magnitude of the correlation between the state monitoring parameters of the coal-fired unit, quantitatively measure the degree of correlation between the state monitoring parameters, and provide a reference for the next step to select necessary data for missing value detection and outlier detection; in addition, the results of CCA data correlation analysis can also be used as parameters in subsequent model training and inference;

[0044] With reference to the degree of correlation between the state monitoring parameters, select necessary data from the data in the dataset of the monitoring panel system of a coal-fired unit for missing value detection based on normal distribution and outlier detection based on K-means, and correspondingly process the missing data (such as interpolation, completely removing the data points containing missing values from the dataset, etc.), and process the abnormal data (such as deletion, replacement, marking, imputation, etc.);

[0045] Use the improved wavelet threshold method to perform denoising processing on the processed data.

[0046] Among them, there is no strict order between the two processes of missing value detection and handling based on normal distribution and outlier detection and handling based on K-means. Either data processing method can be carried out first. However, CCA data correlation analysis needs to be placed at the first step of data processing because CCA data correlation analysis studies the relationships between variables. According to the CCA data correlation analysis, the association rule measurement results among various monitoring parameters are obtained to determine whether there is a strong correlation between parameters. For some strongly correlated parameters, only representatives need to be selected. Therefore, only the selected parameters need to undergo a series of subsequent data processing, reducing the workload of data detection and processing.

[0047] CCA data correlation analysis is specifically implemented through the following formula:

[0048] Two groups of original variables X and Y:

[0049]

[0050] Composite variables R and S:

[0051]

[0052] The correlation coefficients of composite variables R and S are as follows:

[0053]

[0054] In the formula, U and V are transformation coefficient matrices, cov(R,S) represents the covariance of composite variables R and S, and Var(R) and Var(S) represent the variances of composite variables R and S respectively.

[0055] Missing value detection based on normal distribution is used to detect missing values in the data set of the coal-fired unit monitoring and control system, and to make up for the negative impact of missing values on data analysis. Missing value detection based on normal distribution is specifically implemented through the following formula:

[0056] Mean and standard deviation Maximum likelihood estimation:

[0057]

[0058] Traverse the sample set:

[0059]

[0060] Outlier detection based on K-means is used to detect outliers in the data set of the coal-fired unit monitoring and control system, which is used to find the data that is different from most of the data in the data set. Outlier detection based on K-means is specifically implemented through the following formula:

[0061] 1) Arbitrarily select B elements from dataset A as the initial clustering centers;

[0062] 2) Calculate the distances between the remaining elements and the B clustering centers, and assign them to the nearest clustering centers;

[0063] 3) Recalculate the arithmetic means of the B clusters and update to generate new clustering centers;

[0064] 4) Repeat steps 2) and 3) until the clustering results no longer change, and the clustering ends.

[0065] Denoise the data in the monitoring system dataset of coal-fired power units by improving the wavelet threshold denoising method to improve the accuracy and reliability of the data. The improved wavelet threshold denoising is specifically implemented through the following formula:

[0066]

[0067] In the formula, ω j,k is the k-th wavelet coefficient of the j-th layer, E j and E noise are the energy distributions of the signal and noise of the j-th layer respectively, N is the number of samples, is the mean of the wavelet coefficients of the j-th layer, α is the protection factor, and σ is the standard deviation.

[0068] The traditional wavelet threshold method uses a fixed threshold method for calculation, which has limitations such as unreasonable decomposition and inconsistent local characteristics of the signal at different scales. The improved wavelet threshold method proposed in this embodiment is a dynamically corrected wavelet threshold method, that is, it dynamically adjusts the threshold of the wavelet coefficients according to different scales, which can fully balance the contradiction between noise interference and useful signals, and makes up for the limitation of the inconsistent local characteristics of the signal at different scales of the traditional wavelet threshold method. Under different decomposition scale conditions, good denoising effects are achieved.

[0069] Case analysis example:

[0070] Take the primary fan of Unit 1 of a certain power plant as an example for case analysis. Select 17 measuring points related to the operation status of the primary fan of Unit 1 from the SIS of the plant, specifically including the primary fan X-axis vibration measuring point (1), the primary fan Y-axis vibration measuring point (1), the primary fan front bearing temperature measuring points (3), the primary fan middle bearing temperature measuring points (3), the primary fan rear bearing temperature measuring points (3), the motor stator coil temperature measuring points (3), and the motor bearing temperature measuring points (3).

[0071] Taking the X-axis vibration measurement points (1) and (3) of the primary air fan as an example for correlation analysis. Table 1 is the correlation coefficient table. It can be seen from Table 1 that the correlation between the X-axis vibration measurement points of the primary air fan and the bearing temperature measurement points at the rear of the primary air fan is relatively low, while the correlation between the bearing temperature measurement points at the rear of the primary air fan is relatively strong.

[0072] Table 1 Correlation Coefficient Table

[0073]

[0074] Table 2 and Table 3 are the initial clustering centers and the final clustering centers using K-means clustering.

[0075] Table 2 Initial Clustering Center

[0076]

[0077] Table 3 Final Clustering Center

[0078]

[0079]

[0080] Figure 2 It is a comparison chart before and after denoising by the improved wavelet threshold method. In the figure, the black solid curve represents the signal before denoising, and the red dashed curve represents the signal after denoising.

[0081] In one embodiment, a computer device is further provided, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above-mentioned data cleaning method for the intelligent monitoring system of coal-fired power plants.

[0082] In one embodiment, a computer program product is further provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned data cleaning method for the intelligent monitoring system of coal-fired power plants are implemented.

[0083] The technical features of the above embodiments can be combined arbitrarily. For the sake of brief description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

Claims

1. A data cleaning method for an intelligent monitoring panel system of a coal-fired power plant, characterized in that, Including: Obtain the data set of the coal-fired unit monitoring system, which includes data in the electrical system, steam-water system, and boiler system related to the operating status of the coal-fired unit; First, perform CCA data correlation analysis on the data in the coal-fired unit monitoring system data set to obtain the magnitude of the correlation between the status monitoring parameters of the coal-fired unit, quantitatively measure the correlation degree between the status monitoring parameters, and provide a reference for the next step to select necessary data for missing value detection and outlier detection; Refer to the correlation degree between the status monitoring parameters, select necessary data from the data in the coal-fired unit monitoring system data set for missing value detection based on normal distribution and outlier detection based on K-means, and correspondingly process missing data and abnormal data; Finally, use the improved wavelet threshold method to denoise the processed data; the improved wavelet threshold method is implemented through the following formula: where ω j,k is the k-th wavelet coefficient of the j-th layer, E j and E noise are the energy distributions of the signal and noise of the j-th layer respectively, N is the number of samples, is the mean of the wavelet coefficients of the j-th layer, α is the protection factor, and σ is the standard deviation.

2. The data cleaning method of the intelligent monitoring panel system for coal-fired power plants according to claim 1, wherein The CCA data correlation analysis is implemented through the following formula: Two groups of original variables X and Y: Composite variables R and S: The correlation coefficients of the composite variables R and S are as follows: In the formula, U and V are transformation coefficient matrices, cov(R,S) represents the covariance of the composite variables R and S, and Var(R) and Var(S) represent the variances of the composite variables R and S respectively.

3. The data cleaning method of the intelligent monitoring system for coal-fired power plants according to claim 1, wherein The missing value detection based on normal distribution is implemented through the following formula: Mean and standard deviation Maximum likelihood estimation: Traverse the sample set:

4. The data cleaning method of the intelligent monitoring system for coal-fired power plants according to claim 1, characterized in that The outlier detection based on K-means is implemented through the following steps: 1) Arbitrarily select B elements from the data set A as the initial clustering centers; 2) Calculate the distances between the remaining elements and the B clustering centers and assign them to the nearest clustering centers; 3) Recalculate the arithmetic means of the B clusters and update to generate new clustering centers; 4) Repeat steps 2) and 3) until the clustering results no longer change and the clustering ends.

5. The data cleaning method of the intelligent monitoring panel system for coal-fired power plants according to claim 1, wherein The processing of missing data includes interpolation or completely removing the data points containing missing values from the data set.

6. The data cleaning method of the intelligent monitoring system for coal-fired power plants according to claim 1, characterized in that The processing of abnormal data includes deletion, replacement, marking, and imputation.

7. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the data cleaning method of the intelligent monitoring system for coal-fired power plants described in claim 1.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the data cleaning method of the intelligent monitoring system for coal-fired power plants described in claim 1 are implemented.

Citation Information

Patent Citations

  • Power equipment state evaluation method based on big data analysis

    CN111768082A

  • Turbine operation state parameter optimization method based on correlation analysis and FCM clustering

    CN116956090A