Power plant data management method and management platform based on industrial internet

By calculating the correlation number and iteratively adjusting the neighborhood radius, the problem of noise data detection error in power plant data management is solved, and more efficient data management is achieved.

CN120296444AInactive Publication Date: 2025-07-11QINGDAO YONGTAIYUAN THERMAL POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510448293.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing power plant data management, there are errors in the noise data detection process, which leads to the inability of clustering algorithms to accurately detect noise data, affecting data quality and management efficiency.

Method used

By analyzing the correlation of different types of data, calculating the correlation coefficient, determining the initial neighborhood radius, and iteratively adjusting the optimal neighborhood radius, combining with the clustering algorithm to eliminate noise data.

Benefits of technology

It improves the clustering effect of data, accurately identify and eliminate noise data, and improves the quality and management efficiency of power plant data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296444A_ABST
    Figure CN120296444A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data management, in particular to a power plant data management method and management platform based on the industrial internet, and the method comprises the steps: obtaining the monitoring data of each data type of a power plant at each moment; calculating a correlation coefficient of any two data types; obtaining each association set; determining a measurement distance of any two moments in each association set, and obtaining an initial neighborhood radius of each association set; determining the adjustment degree of each association set, and evaluating the neighborhood radius of each association set; determining the adjusted neighborhood radius of each association set; and obtaining the optimal neighborhood radius of each association set, clustering all moments in the association sets in combination with a clustering algorithm, and eliminating noise data in all data types in the power plant. The detection precision of the noise data in the power plant data is improved, and the quality of the power plant data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of big data management, and in particular to a power plant data management method and management platform based on the industrial Internet. Background Art

[0002] Introducing a big data management platform in a power plant can collect, sort, analyze, and apply a large amount of different types of data information, so as to deeply mine the potential value of the power plant from the data, assist power plant managers in formulating corresponding countermeasures, effectively control the production and operation links of the power plant, and improve the comprehensive benefits of the power plant.

[0003] When managing data, the quality of the data is poor and there is a lot of noise. When detecting the noise in the data through a clustering algorithm, the performance of the clustering algorithm highly depends on the neighborhood radius parameter. If the neighborhood radius is too small, the high-density area will be overly divided into multiple small clusters, and a large number of sparse points will be misjudged as noise. If the neighborhood radius is too large, clusters with different densities will be merged into an oversized cluster, and noise points will be wrongly included in the cluster, resulting in errors in the process of detecting noise points in power plant data and unable to accurately detect noise data. Summary of the Invention

[0004] To solve the above technical problems, a power plant data management method and management platform based on the industrial Internet are provided to solve the existing problems.

[0005] The solution of this application to solve the technical problem is to provide a power plant data management method and management platform based on the industrial Internet, including the following steps:

[0006] In the first aspect, an embodiment of this application provides a power plant data management method based on the industrial Internet. The method includes the following steps:

[0007] Obtain the monitoring data of each data type in the power plant at each moment;

[0008] Analyze the frequency characteristics of the simultaneous occurrence of the monitoring data of any two data types, as well as the correlation between the monitoring data of the any two data types, and calculate the correlation coefficient of the any two data types;

[0009] Classify all data types based on the correlation coefficient to obtain each association set;

[0010] Determine the metric distance between any two moments in each association set through the difference situation of the monitoring data of each data type in each association set at any two moments, and the average level of the correlation coefficient between each data type and the remaining data types, and obtain the initial neighborhood radius of each association set;

[0011] Analyze the discreteness of the metric distances between each moment in each associated set and the remaining moments within its initial neighborhood radius, as well as the difference between the maximum metric distance and the initial neighborhood radius. Combine the quantity situation of the remaining moments to determine the adjustment degree of each associated set, and evaluate the adjustment of the neighborhood radius of each associated set;

[0012] Analyze the discreteness of the metric distances between each moment in each associated set and the remaining moments within its initial neighborhood radius, as well as the difference between the minimum metric distance and the initial neighborhood radius, to determine the adjusted neighborhood radius of each associated set; Iterate the neighborhood radius of each associated set to obtain the optimal neighborhood radius of each associated set. Combine the clustering algorithm to cluster the monitoring data of all data types within the associated sets at different moments, and eliminate the noise data in all data types in the power plant.

[0013] Preferably, calculating the correlation coefficient between any two data types includes:

[0014] Form data point pairs from the monitoring data at the same moment between any two data types; Count the frequency of occurrence of each data point pair in any two data types;

[0015] Calculate the reciprocal of the mean of the differences between the frequencies of all any two data point pairs in any two data types, which is denoted as the co-occurrence correlation degree;

[0016] Calculate the correlation degree of the monitoring data at all moments between any two data types;

[0017] The correlation coefficient is the product of the absolute value of the correlation degree and the co-occurrence correlation degree.

[0018] Preferably, the further acquisition process of each associated set is as follows:

[0019] Cluster the correlation coefficients between any one data type and all the remaining data types to obtain multiple clusters;

[0020] Calculate the average value of all the correlation coefficients within each cluster, and denote the cluster corresponding to the maximum average value as the relevant cluster;

[0021] Form a relevant class set from any one data type and all the corresponding data types in its relevant cluster;

[0022] Denote the intersection between the relevant class set of any one data type and the relevant class sets of all the data types in its relevant cluster as the associated set.

[0023] Preferably, determining the metric distance between any two moments in each associated set includes:

[0024] Form a feature vector from the monitoring data of all data types at each moment in the association set;

[0025] Calculate the mean of the correlation coefficients of each data type in the association set with all other data types, denoted as the association weight;

[0026] Calculate the difference between the elements corresponding to the same data type in the feature vectors of any two moments in the association set, denoted as the relative difference amount;

[0027] Based on the association weight, perform a weighted sum of the relative difference amounts corresponding to all data types at any two moments in the association set, as the metric distance between any two moments in the association set.

[0028] Preferably, the initial neighborhood radius is the mean of the metric distances between all any two moments in the association set.

[0029] Preferably, the determination of the adjustment degree of each association set includes:

[0030] Take a preset multiple of the number of all data types in the association set as the minimum number of points of the association set;

[0031] Obtain the maximum and minimum values of the metric distances between any moment in the association set and all other moments within its initial neighborhood radius, denoted as the neighborhood maximum distance and the neighborhood minimum distance respectively;

[0032] Calculate the dispersion degree of the metric distances between any moment and all other moments within its initial neighborhood radius;

[0033] Calculate the difference between the neighborhood maximum distance and the initial neighborhood radius, denoted as the first difference;

[0034] Count the number of all moments within the initial neighborhood radius of any moment, and denote the difference between the number and the minimum number of points as the second difference; calculate the ratio of the first difference to the second difference; denote the product of the ratio and the dispersion degree as the adjustment factor;

[0035] The adjustment degree is the mean of the adjustment factors of all moments in each association set.

[0036] Preferably, the evaluation of adjusting the neighborhood radius of each association set includes:

[0037] Calculate the mean of the maximum adjustment factor and the minimum adjustment factor in each association set, denoted as the evaluation threshold;

[0038] If the adjustment degree is greater than the evaluation threshold, adjust the neighborhood radius of each association set, otherwise, do not adjust the neighborhood radius of each association set.

[0039] Preferably, determining the adjusted neighborhood radius of each association set includes:

[0040] Denote the normalization result of the sum of the degrees of dispersion at all times in each association set as the first adjustment value;

[0041] Denote the difference between the neighborhood minimum distance corresponding to each time in the association set and the initial neighborhood radius as the third difference, calculate the cumulative sum of the products of the third differences at all times in the association set and the neighborhood minimum distance, and denote the normalization result after negative mapping of the cumulative sum as the second adjustment value;

[0042] The adjusted neighborhood radius R of the h-th association set ′ h is calculated as follows: wherein, R h is the initial neighborhood radius of the h-th association set, Z h,1 is the first adjustment value of the h-th association set, and Z h,2 is the second adjustment value of the h-th association set.

[0043] Preferably, removing the noise data in all data types in the power plant includes:

[0044] Cluster all times in the association set based on the optimal neighborhood radius and the minimum number of points, and denote the times corresponding to the marked noise points as noise times; remove the monitoring data at all noise times in each data type in each association set.

[0045] In a second aspect, an embodiment of the present application further provides a power plant data management platform based on the industrial Internet, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned power plant data management method based on the industrial Internet.

[0046] The present application has at least the following beneficial effects:

[0047] This application analyzes the association of different data types, calculates the correlation coefficient between any two data types, and its beneficial effects are as follows: considering the co-occurrence frequency of the monitoring data of the two data types and the correlation between the two data types to reflect the association of the two data types, so as to classify the data types subsequently and obtain each association set. The beneficial effect is that the data types with strong correlation are classified into one category, improving the efficiency and accuracy of the management of different data types; secondly, by calculating the difference of the monitoring data of each data type between different times in each association set, calculating the metric distance between any two times in each association set, and obtaining the initial neighborhood radius of each association set. The beneficial effect is that it considers the differences between all types of data collected at different times to evaluate the degree of difference between the data at different times, so as to identify the data at abnormal times subsequently, making it possible to more accurately eliminate noise data; further, calculating the adjustment degree of each association set, and its beneficial effect is that it considers the discrete situation of the metric distances between different times within the initial neighborhood radius of each time and the proximity of the maximum metric distance to the initial neighborhood radius, so as to evaluate whether the initial neighborhood radius should be adjusted; for those that need to adjust the neighborhood radius, judging whether to increase or decrease the neighborhood radius based on the discrete situation of the metric distances between different times within the initial neighborhood radius of each time and the difference between the minimum metric distance and the initial neighborhood radius, so as to obtain the adjusted neighborhood radius of each association set. Iterating the neighborhood radius of each association set to obtain the optimal neighborhood radius of each association set, and combining the clustering algorithm to cluster the monitoring data of all types of data within the association set at different times, eliminating the noise data in each type of data in the power plant. The beneficial effect is that it can select a suitable optimal neighborhood radius for different association sets, enabling the DBSCAN algorithm to perform clustering based on the optimal neighborhood radius of different association sets, improving the clustering effect of the data, being able to more accurately detect the noise data in the power plant data, reducing the detection error of the noise data, enhancing the quality of the power plant data, and further improving the management efficiency and accuracy of the power plant data. Brief Description of the Drawings

[0048] The following further elaborates in detail a method for managing power plant data based on the industrial Internet according to this application with reference to the drawings.

[0049] Figure 1 It is a flowchart of the steps of a method for managing power plant data based on the industrial Internet provided by an embodiment of this application;

[0050] Figure 2 It is a flowchart of the steps of a method for obtaining the correlation coefficient between any two data types provided by an embodiment of this application;

[0051] Figure 3It is a flowchart of the steps for obtaining the initial neighborhood radius of each association set provided by the embodiments of the present application. Detailed implementation manners

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details a power plant data management method and management platform based on industrial Internet in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0054] Please refer to Figure 1 , which shows a flowchart of the steps of a power plant data management method based on industrial Internet provided by an embodiment of the present application. The method includes the following steps:

[0055] Step 1: Obtain the monitoring data of each data type of the power plant at each moment.

[0056] Obtain the monitoring data of each data type of the power plant at different moments;

[0057] In this embodiment, taking a thermal power plant as an example, data of multiple data types such as boiler fuel quantity, feed water quantity, desuperheating water quantity, primary air volume, oxygen content, boiler flue gas temperature, furnace pressure, and steam drum water level in the thermal power plant are collected through intelligent sensors. Among them, the time interval for data collection is 1 s. As other implementation manners, the implementer can set it according to the actual situation.

[0058] Therefore, the boiler fuel quantity, feed water quantity, desuperheating water quantity, primary air volume, oxygen content, boiler flue gas temperature, furnace pressure, and steam drum water level are recorded as each monitoring parameter, and all the collected data are normalized to remove the dimension of the data.

[0059] In this embodiment, the maximum-minimum normalization method is used for normalization. Among them, the maximum-minimum normalization method is a well-known technology and will not be elaborated here. As other implementation manners, the implementer can adopt other methods of existing technologies, such as the Z-score normalization method, etc. This embodiment does not make special restrictions on this.

[0060] So far, the monitoring data of each data type of the power plant at different moments are obtained.

[0061] Step 2: Analyze the frequency characteristics of the simultaneous occurrence of the monitoring data of any two data types, and the correlation of the monitoring data between the any two data types, and calculate the correlation coefficient of the any two data types.

[0062] When managing a large amount of data in a thermal power plant, it is necessary to remove noise from the data and manage the data classification. Therefore, when using the DBSCAN algorithm to process the data of the thermal power plant, it is necessary to set the initial neighborhood radius and the minimum number of points. The traditional algorithm generally sets them manually, while the performance of the DBSCAN algorithm highly depends on the initial neighborhood radius parameter. If the initial neighborhood radius is too small, the high-density area will be overly segmented into multiple small clusters, and a large number of sparse points will be misjudged as noise. If the initial neighborhood radius is too large, clusters with different densities will be merged into an oversized cluster, and noise points will be wrongly included in the cluster, resulting in a low recognition accuracy for noise points.

[0063] Secondly, due to the large variety and huge amount of data in the thermal power plant, first, analyze the correlation of the monitoring data changes between different types of data, and calculate the correlation coefficient to reflect the correlation between different types of data.

[0064] Furthermore, the step flowchart of the method for obtaining the correlation coefficient of any two data types provided by the embodiments of the present application is as Figure 2 shown.

[0065] First, calculate the co-occurrence correlation degree through the occurrence frequencies of two monitoring data at the same moment between any two data types. Specifically:

[0066] Form data point pairs from the monitoring data at the same moment between any two data types;

[0067] Count the occurrence frequency of each data point pair in the any two data types;

[0068] Calculate the reciprocal of the mean of the differences between the frequencies of all any two data point pairs in the any two data types as the co-occurrence correlation degree of the any two data types;

[0069] In this embodiment, calculate the mean of the absolute values of the differences between the frequencies of all any two data point pairs in the any two data types.

[0070] It should be noted that for the sake of easy understanding, assume that for the two data types of boiler fuel quantity and boiler flue gas temperature, the monitoring data of the boiler fuel quantity at each moment is represented by the sequence [1, 3, 1, 3, 2, 3], and the monitoring data of the boiler flue gas temperature at each moment is represented by the sequence [2, 1, 2, 1, 3, 1]. Then the data point pairs are (1, 2), (3, 1), (1, 2), (3, 1), (2, 3), (3, 1). Therefore, the occurrence frequency of the data point pair (1, 2) is The occurrence frequency of the data point pair (3, 1) is The occurrence frequency of the data point pair (2, 3) is 6 1 .

[0071] It should be noted that when calculating the reciprocal, to avoid the denominator being 0, a preset value greater than 0 is added to the denominator. In this embodiment, the preset value greater than 0 is taken as 0.01. Secondly, the greater the co-occurrence correlation degree, it indicates that when the monitoring data of two data types change, the frequency distribution of the data point pairs appears relatively concentrated, that is, some data point pairs appear more frequently, reflecting a strong correlation between the two data types.

[0072] Secondly, to reflect the characteristic relationship that the change of the monitoring data of one data type causes the change of the monitoring data of another data type, by analyzing the relevant changes of the monitoring data between any two data types and combining the co-occurrence correlation degree, the correlation coefficient is calculated as follows:

[0073] Calculate the degree of correlation of the monitoring data at all times between any two data types;

[0074] In this embodiment, the degree of correlation is measured by calculating the Pearson correlation coefficient of the monitoring data at all times between any two data types. Among them, the calculation method of the Pearson correlation coefficient is a well-known technology. As other implementation manners, implementers can adopt other methods of the prior art, for example, the Spearman correlation coefficient, etc. This embodiment does not make special restrictions on this, that is, calculate the Pearson correlation coefficient of the monitoring data at all times between the boiler fuel quantity and the boiler flue gas temperature; or calculate the Pearson correlation coefficient of the monitoring data at all times between the primary air volume and the oxygen content.

[0075] Take the product of the co-occurrence correlation degree and the change correlation degree as the correlation coefficient of any two data types;

[0076] It should be noted that the greater the absolute value of the degree of correlation, it indicates that the monitoring data of the two data types have a higher correlation, and the greater the obtained change correlation degree, it indicates that the changes of the monitoring data between the two data types are correlated. When the monitoring data of a certain data type fluctuates, it will cause the monitoring data of another data type to also fluctuate accordingly. The greater the correlation coefficient, it indicates that the two data types have a stronger correlation.

[0077] Thus, the correlation coefficient of any two data types is obtained.

[0078] Step 3, based on the correlation coefficient, classify all data types to obtain each association set; determine the measurement distance between any two moments in each association set through the difference situation of the monitoring data of each data type in each association set at any two moments and the average level of the correlation coefficients of each data type with the remaining data types, and obtain the initial neighborhood radius of each association set.

[0079] Furthermore, based on the correlation coefficient, all types of data are clustered, specifically as follows:

[0080] Cluster the correlation coefficients of any one type of data with all the other types of data to obtain multiple clusters;

[0081] In this embodiment, the K-means clustering algorithm is used for clustering to obtain two clustering clusters. Among them, the K-means clustering algorithm is a well-known technology and will not be elaborated here. That is, through the K-means clustering algorithm, the boiler fuel quantity, feed water quantity, desuperheating water quantity, primary air volume, oxygen content, boiler flue gas temperature, furnace pressure, and steam drum water level in the thermal power plant are clustered.

[0082] Calculate the average value of all the correlation coefficients within each cluster, and mark the cluster corresponding to the largest average value as the relevant cluster;

[0083] Form a relevant class set with any one type of data and all the corresponding types of data in its relevant cluster;

[0084] Mark the intersection between the relevant class set of any one type of data and the relevant class sets of all the types of data in its relevant cluster as the associated set;

[0085] It should be noted that for the convenience of understanding, assume that the data types of the boiler fuel quantity, feed water quantity, desuperheating water quantity, primary air volume, oxygen content, boiler flue gas temperature, furnace pressure, and steam drum water level are respectively denoted as A, B, C, D, E, F, G, V. Taking type A data as an example, calculate the correlation coefficients of A with B, C, D, E, F, G, V respectively, denoted as r A,B 、r A,C 、r A,D 、r A,E 、r A,F 、r A,G 、r A,V , and cluster r A,B 、r A,C 、r A,D 、r A,E 、r A,F into two categories. Assume that r A,B 、r A,D 、r A,E 、r A,G 、r A,V are in one category, r A,C 、r A,F are in one category, and the average value of r A,C 、r A,F is the largest. Therefore, the relevant cluster is {r A,C 、r A,F} Therefore, the set of related classes for data type A is {A, C, F}; correspondingly, assuming that the set of related classes for data type C is {C, A, E, G}, and the set of related classes for data type F is {F, C, A, V}, then taking the intersection of the sets of related classes for the three data types A, C, and F, the associated set is {C, A}.

[0086] Through the above process, all types of data can be divided into multiple classes, that is, multiple associated sets are obtained. The monitoring data of all types of data at each moment within each associated set can be regarded as the data distribution in a high-dimensional space. In traditional algorithms, the Euclidean distance is used as the distance metric index. However, the Euclidean distance is relatively intuitive and effective in a low-dimensional space, but performs poorly in a high-dimensional space. Therefore, by analyzing the differences in the monitoring data between different types of data within each associated set, the distance between the monitoring data at different moments is measured in the high-dimensional space to determine the initial neighborhood radius in the DBSCAN algorithm. The flowchart of the steps for obtaining the initial neighborhood radius of each associated set provided in the embodiments of the present application is as Figure 3 shown, and specifically includes:

[0087] Form the monitoring data of all types of data at each moment within each associated set into a feature vector;

[0088] Calculate the mean value of the correlation coefficients between each type of data and all other types of data within each associated set, and denote it as the association weight;

[0089] Calculate the difference between the corresponding elements of the same type of data in the feature vectors at any two moments within each associated set, and denote it as the relative difference amount;

[0090] In this embodiment, calculate the absolute value of the difference between the corresponding elements of the same type of data in the feature vectors at any two moments within the associated set, and denote it as the relative difference amount.

[0091] Based on the association weight, perform a weighted sum of the relative difference amounts corresponding to all types of data at any two moments within each associated set, and use it as the measurement distance between any two moments within each associated set;

[0092] In this embodiment, taking the hth associated set as an example, the calculation formula for its measurement distance is:

[0093]

[0094] where is the measurement distance between the tth moment and the xth moment in the hth associated set, ω h,m is the association weight of the mth type of data in the hth associated set, is the element corresponding to the m-th data type in the feature vector at the t-th moment in the h-th associated set. is the element corresponding to the m-th data type in the feature vector at the x-th moment in the h-th associated set, M h is the number of all data types in the h-th associated set.

[0095] It should be noted that the larger the obtained relative difference amount, the greater the difference in the monitored data of this data type in the associated set. The larger the associated weight, the greater the influence of this data type on the remaining data types. By introducing the associated weight, the data difference between two moments is increased, making the data difference between different moments more sensitive, which is beneficial to the detection and identification of noise data.

[0096] Take the mean of the metric distances between all any two moments in each associated set as the initial neighborhood radius of each associated set.

[0097] Thus, the initial neighborhood radius of each associated set is obtained.

[0098] Step 4: Analyze the dispersion of the metric distances between each moment in each associated set and the remaining moments within its initial neighborhood radius, as well as the difference between the maximum metric distance and the initial neighborhood radius. Combine the quantity situation of the remaining moments to determine the adjustment degree of each associated set and evaluate the adjustment of the neighborhood radius of each associated set.

[0099] Take a preset multiple of the number of all data types in each associated set as the minimum number of points for each associated set;

[0100] In this embodiment, the preset multiple is 2. Taking the h-th associated set as an example, the number of all data types in the h-th associated set is M h , and take 2×M h as the minimum number of points of the DBSCAN algorithm corresponding to each associated set.

[0101] Furthermore, based on the initial neighborhood radius and the minimum number of points, determine the adjustment degree, specifically:

[0102] Obtain the maximum and minimum values of the metric distances between any moment in each associated set and the remaining all moments within its initial neighborhood radius, and record them as the neighborhood maximum distance and neighborhood minimum distance corresponding to any moment respectively;

[0103] Calculate the dispersion degree of the metric distances between any moment in each associated set and the remaining all moments within its initial neighborhood radius;

[0104] In this embodiment, the degree of dispersion is measured by calculating the standard deviation of the metric distances between any moment in each associated set and all the other moments within its initial neighborhood radius. As other implementation manners, implementers can adopt other methods in the prior art, such as variance, coefficient of variation, etc. This embodiment does not make special limitations thereon.

[0105] Calculate the difference between the maximum neighborhood distance and the initial neighborhood radius, and denote it as the first difference;

[0106] In this embodiment, calculate the absolute value of the difference between the maximum neighborhood distance and the initial neighborhood radius, and denote it as the first difference.

[0107] Count the number of all moments within the initial neighborhood radius of any moment, and denote the difference between the number and the minimum number of points as the second difference;

[0108] In this embodiment, denote the absolute value of the difference between the number and the minimum number of points as the second difference.

[0109] Calculate the ratio of the first difference to the second difference, and denote the product of the ratio and the degree of dispersion as the adjustment factor of any moment in each associated set;

[0110] In this embodiment, taking the t-th moment in the h-th associated set as an example, the calculation formula of its adjustment factor is:

[0111]

[0112] where, Q h,t is the adjustment factor of the t-th moment in the h-th associated set, σ h,t is the degree of dispersion of the t-th moment in the h-th associated set, is the maximum neighborhood distance corresponding to the t-th moment in the h-th associated set, R h is the initial neighborhood radius of the h-th associated set, N h,t is the number of all moments within the initial neighborhood radius of the t-th moment in the h-th associated set, P h is the minimum number of points of the h-th associated set, and ε is a preset value greater than 0 to avoid the denominator being 0. In this embodiment, the value of ε is 0.1. As other implementation manners, implementers can set it by themselves according to the actual situation.

[0113] Calculate the mean value of the adjustment factors of all moments in each associated set as the adjustment degree of each associated set;

[0114] It should be noted that the greater the degree of dispersion, the more chaotic the data distribution at different times within the initial neighborhood radius at the corresponding time, and the less concentrated the distribution of the monitoring data at different times; the greater the first difference, the smaller the data distribution range within the initial neighborhood radius at the corresponding time, and the initial neighborhood radius should be adjusted; secondly, the greater the second difference, the greater the deviation between the actual quantity and the expected quantity at all times within the initial neighborhood radius. After adjusting the initial neighborhood radius, the change in the quantity of data has a small impact, and the more the initial neighborhood radius should be adjusted, the greater the obtained adjustment factor. Correspondingly, the greater the degree of adjustment, the more the initial neighborhood radius of the associated set should be adjusted.

[0115] Furthermore, based on the degree of adjustment, the initial neighborhood radius of each associated set is evaluated, specifically as follows:

[0116] Calculate the mean of the maximum adjustment factor and the minimum adjustment factor in each associated set, denoted as the evaluation threshold;

[0117] If the degree of adjustment is greater than the evaluation threshold, the neighborhood radius of each associated set is adjusted; otherwise, the neighborhood radius of each associated set is not adjusted;

[0118] It should be noted that if the degree of adjustment is greater than the evaluation threshold, it indicates that the moments with large adjustment factors in each associated set account for a relatively large proportion. Therefore, the neighborhood radius needs to be adjusted.

[0119] Thus, the degree of adjustment of each associated set is obtained.

[0120] Step 5: Analyze the dispersion of the metric distances between each moment in each associated set and the other moments within its initial neighborhood radius, as well as the difference between the minimum metric distance and the initial neighborhood radius, to determine the adjusted neighborhood radius of each associated set; iterate the neighborhood radii of each associated set to obtain the optimal neighborhood radius of each associated set, and combine the clustering algorithm to cluster the monitoring data of all data types within the associated sets at different times, and eliminate the noise data in all data types in the power plant.

[0121] Furthermore, the process of adjusting the neighborhood radius of each associated set is as follows:

[0122] Denote the normalized result of the sum of the degrees of dispersion of all moments in each associated set as the first adjustment value;

[0123] In this embodiment, the sigmoid function is used for normalization processing. The sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, the implementer can use other methods according to the existing technology, such as the tanh function, etc. This embodiment does not make special restrictions on this.

[0124] Denote the difference between the neighborhood minimum distance corresponding to each moment in the associated set and the initial neighborhood radius as the third difference, calculate the cumulative sum of the product of the third difference and the neighborhood minimum distance for all moments in the associated set, and denote the normalized result after negative mapping of the cumulative sum as the second adjustment value;

[0125] In this embodiment, denote the absolute value of the difference between the neighborhood minimum distance corresponding to each moment in the associated set and the initial neighborhood radius as the third difference.

[0126] In this embodiment, the process of negative mapping is as follows: perform negative mapping through an exponential function. Denote the cumulative sum as β, then the result of exp(-β) is used as the second adjustment value. Since the value range of exp(-β) is in (0, 1], there is no need to perform normalization processing, and directly use the result of negative mapping as the second adjustment value; As other embodiments, after the implementer uses other methods for negative mapping, the sigmoid function can be used for normalization processing. Among them, the sigmoid function is a well-known technology and will not be elaborated here.

[0127] Secondly, in this embodiment, the calculation formula for the second adjustment value of each associated set is:

[0128]

[0129] where Z h,2 is the second adjustment value of the h-th associated set, is the neighborhood minimum distance corresponding to the t-th moment in the h-th associated set; R h is the initial neighborhood radius of the h-th associated set, T is the number of all moments, and exp() is the exponential function with the natural constant as the base.

[0130] The calculation method for the adjusted neighborhood radius of each associated set is:

[0131]

[0132] where R ′ h is the adjusted neighborhood radius of the h-th associated set, R h is the initial neighborhood radius of the h-th associated set, Z h,1 is the first adjustment value of the h-th associated set, and Z h,2 is the second adjustment value of the h-th associated set.

[0133] It should be noted that the larger the first adjustment value, the more dispersed the distance distribution at different times within the neighborhood radius of the h-th associated set, which indicates that the monitoring data corresponding to different times is more likely not to belong to the same category, and thus the neighborhood radius should be decreased; the larger the second adjustment value, the closer the minimum distance of the neighborhood corresponding to different times in the h-th associated set is to the initial neighborhood radius, which reflects that the neighborhood radius is too small and the neighborhood radius should be increased.

[0134] Denote the adjusted neighborhood radius R ′ h as the neighborhood radius of the second iteration. According to the above process, calculate the adjustment degree of the second iteration, evaluate the neighborhood radius of the second iteration. If the neighborhood radius of the second iteration needs to be adjusted, calculate the first adjustment value and the second adjustment value of the second iteration to obtain the adjusted neighborhood radius at the second iteration, and iterate the neighborhood radius in sequence until the neighborhood radius no longer needs to be adjusted finally, then stop the iteration to obtain the optimal neighborhood radius of each associated set.

[0135] Combine all data types at each time in the associated set to form a feature vector. Based on the optimal neighborhood radius and the minimum number of points, cluster the feature vectors corresponding to all times in the associated set through the DBSCAN algorithm (Density-Based Spatial Clustering of Applications with Noise), and denote the times corresponding to the marked noise points as noise times.

[0136] It should be noted that the DBSCAN algorithm is a well-known technology and will not be elaborated here. Secondly, the DBSCAN algorithm will classify the feature vectors corresponding to different times in the associated set into three categories. Among them, the times that are neither core points nor border points are marked as noise points, that is, if there are at least the minimum number of points within the neighborhood radius of a certain time, then the time is a core point. If the number of points within the neighborhood of a certain time is less than the minimum number of points but belongs to the neighborhood of a certain core point, then this time is used as a border point. If the number of points within the neighborhood of a certain time is less than the minimum number of points and does not belong to the neighborhood of a certain core point, then it is marked as a noise point.

[0137] Delete the monitoring data of all noise times in each data type in the associated set, so as to clean the monitoring data of all data types in the thermal power plant at different times.

[0138] Based on the same inventive concept as the above method, an embodiment of the present application further provides a power plant data management platform based on industrial Internet, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above methods of a power plant data management method based on industrial Internet are implemented.

[0139] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0141] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation to the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all belong to the protection scope of the technical solution of the present application.

Claims

1. A power plant data management method based on the industrial Internet, characterized in that, The method includes the following steps: Obtain the monitoring data of each data type in the power plant at each moment; Analyze the frequency characteristics of the simultaneous occurrence of the monitoring data of any two data types, as well as the correlation of the monitoring data between any two data types, and calculate the correlation coefficient between any two data types; Classify all data types based on the correlation coefficient to obtain each association set; Determine the metric distance between any two moments in each association set through the difference of the monitoring data of each data type in any two moments within each association set, and the average level of the correlation coefficient between each data type and the remaining data types, to obtain the initial neighborhood radius of each association set; Analyze the dispersion of the metric distance between each moment in each association set and the remaining moments within its initial neighborhood radius, as well as the difference between the maximum metric distance and the initial neighborhood radius, and combine the quantity of the remaining moments to determine the adjustment degree of each association set, and evaluate the adjustment of the neighborhood radius of each association set; Analyze the dispersion of the metric distance between each moment in each association set and the remaining moments within its initial neighborhood radius, as well as the difference between the minimum metric distance and the initial neighborhood radius, to determine the adjusted neighborhood radius of each association set; iterate the neighborhood radius of each association set to obtain the optimal neighborhood radius of each association set, and combine the clustering algorithm to cluster the monitoring data of all data types in the association set at different moments, and eliminate the noise data in all data types in the power plant.

2. The method for power plant data management based on industrial Internet according to claim 1, wherein The calculation of the correlation coefficient between any two data types includes: Form data point pairs from the monitoring data at the same moment between any two data types; count the frequency of each data point pair in any two data types; Calculate the reciprocal of the mean of the differences between the frequencies of all any two data point pairs in any two data types, denoted as the co-occurrence correlation degree; Calculate the correlation degree of the monitoring data between any two data types at all moments; The correlation coefficient is the product of the absolute value of the correlation degree and the co-occurrence correlation degree.

3. A power plant data management method based on industrial Internet as claimed in claim 1, characterized in that, The further acquisition process of each association set is as follows: Cluster the correlation coefficients between any one data type and all the remaining data types to obtain multiple clusters; Calculate the average value of all the correlation coefficients within each cluster, and denote the cluster corresponding to the maximum average value as the relevant cluster; Form a relevant class set from any one data type and all the corresponding data types in its relevant cluster; Denote the intersection between the relevant class set of any one data type and the relevant class sets of all the data types in its relevant cluster as the association set.

4. The method for managing power plant data based on industrial Internet according to claim 1, wherein The determination of the metric distance between any two moments in each association set includes: Form a feature vector from the monitoring data of all data types at each moment in the association set; Calculate the mean value of the correlation coefficients between each data type in the association set and all the remaining data types, denoted as the association weight; Calculate the difference between the elements corresponding to the same data type within the feature vectors of any two moments in the association set, denoted as the relative difference amount; Based on the correlation weights, perform a weighted sum of the relative difference amounts corresponding to all types of data for any two moments in the correlation set, and use it as the metric distance between any two moments in the correlation set.

5. The method for managing power plant data based on industrial Internet according to claim 1, characterized in that, The initial neighborhood radius is the mean of the metric distances between any two moments in the correlation set.

6. The method for power plant data management based on industrial Internet according to claim 1, characterized in that, The determination of the adjustment degree of each correlation set includes: Use a preset multiple of the number of all types of data in the correlation set as the minimum number of points in the correlation set; Obtain the maximum and minimum values of the metric distances between any moment in the correlation set and all other moments within its initial neighborhood radius, and denote them as the neighborhood maximum distance and the neighborhood minimum distance respectively; Calculate the dispersion degree of the metric distances between any moment and all other moments within its initial neighborhood radius; Calculate the difference between the neighborhood maximum distance and the initial neighborhood radius, and denote it as the first difference; Count the number of all moments within the initial neighborhood radius of any moment, and denote the difference between the number and the minimum number of points as the second difference; calculate the ratio of the first difference to the second difference; denote the product of the ratio and the dispersion degree as the adjustment factor; The adjustment degree is the mean of the adjustment factors of all moments in each correlation set.

7. The method for power plant data management based on industrial Internet according to claim 6, wherein The evaluation of adjusting the neighborhood radius of each correlation set includes: Calculate the mean of the maximum adjustment factor and the minimum adjustment factor in each correlation set, and denote it as the evaluation threshold; If the adjustment degree is greater than the evaluation threshold, adjust the neighborhood radius of each correlation set; otherwise, do not adjust the neighborhood radius of each correlation set.

8. The method for power plant data management based on industrial Internet according to claim 6, wherein The determination of the adjusted neighborhood radius of each correlation set includes: Denote the normalized result of the sum of the dispersion degrees of all moments in each correlation set as the first adjustment value; Denote the difference between the neighborhood minimum distance corresponding to each moment in the correlation set and the initial neighborhood radius as the third difference, calculate the cumulative sum of the products of the third differences of all moments in the correlation set and the neighborhood minimum distance, and denote the normalized result after negative mapping of the cumulative sum as the second adjustment value; The adjusted neighborhood radius R of the h-th associated set ′ h is calculated as follows: where R h is the initial neighborhood radius of the h-th associated set, and Z h,1 is the first adjustment value of the h-th associated set, and Z h,2 is the second adjustment value of the h-th associated set.

9. The method for managing power plant data based on industrial Internet according to claim 6, characterized in that, The elimination of noise data in all types of data in the power plant includes: Based on the optimal neighborhood radius and the minimum number of points, cluster all moments in the correlation set, and denote the moments marked as noise points as noise moments; eliminate the monitoring data of all noise moments in each type of data in each correlation set.

10. A power plant data management platform based on the industrial Internet, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for managing power plant data based on industrial Internet according to any one of claims 1-9.