A thermal process alarm data filtering method and system based on AHC-GP hybrid model

By combining the nearest neighbor propagation and agglomerative hierarchical clustering algorithms with the Gaussian process model, the problem of rampant alarms in the thermal process of thermal power units was solved, accurate filtering of thermal process alarm data and positioning of key alarm data were achieved, and the missed detection rate and false positive rate were reduced.

CN114372515BActive Publication Date: 2025-09-09HANGZHOU E ENERGY ELECTRIC POWER TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111578578.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-09-09
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

During the operation of thermal power units, a large amount of nonlinear, strongly coupled, high-dimensional and time-varying thermal process alarm data is generated, resulting in an alarm flood. Operators find it difficult to accurately locate important alarm data and eliminate them in a timely manner, and existing models are unable to effectively filter out redundant alarms.

Method used

The nearest neighbor propagation algorithm is used to determine the optimal number of clusters, and the agglomerative hierarchical clustering algorithm is combined to cluster the data set. The Gaussian process model is used for data classification, and a data filtering model is constructed based on the posterior alarm probability estimate to achieve accurate filtering of thermal process alarm data.

Benefits of technology

Accurately locate key alarm data, eliminate redundant alarms, reduce missed detection and false positive rates, have good data filtering accuracy, and reduce alarm flooding caused by abnormal propagation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114372515B_ABST
    Figure CN114372515B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for filtering thermal process alarm data based on the AHC-GP hybrid model, which belongs to the field of data processing technology. Existing data processing solutions cannot effectively eliminate redundant alarms and cannot suppress alarm overflow in time, exacerbating the "alarm flooding" problem caused by abnormal propagation. The present invention provides a method for filtering thermal process alarm data based on the AHC-GP hybrid model, which can pre-process the data set, and use the nearest neighbor propagation algorithm to determine the optimal number of clusters, and then use the agglomerative hierarchical clustering algorithm to cluster the data set to distinguish different working conditions; secondly, the Gaussian process model is used to classify the data, and the posterior alarm probability estimate is combined to construct a data filtering model to achieve accurate filtering of thermal process alarm data. Furthermore, the present invention can accurately locate the key alarm data in the data set and eliminate redundant alarms; the missed detection rate and the false positive rate are both low, and it has good data filtering accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a thermal process alarm data filtering method and system based on an AHC-GP hybrid model, belonging to the technical field of data processing. Background Art

[0002] With the continuous adjustment of my country's energy structure, thermal power generation faces a series of new difficulties and challenges, such as low-emission transformation, the integration of renewable energy generation, and operation with a shift from coal-based types. This leads to frequent fluctuations in unit operation, which inevitably generates a large amount of alarm data. At the power plant's centralized control center, system operators often need to process a large amount of real-time alarm information and make important decisions about the operation of the power generation system. These alarms may be related to equipment failures, improper operation of protective devices, and other issues. Due to the complex control systems of thermal power units, the large number of production equipment and parameters, the close connection between variables, and the inefficiency of existing alarm systems, the problem of "alarm flooding" is very common.

[0003] Furthermore, simple models are difficult to fit due to the nonlinear, strongly coupled, high-dimensional, and time-varying characteristics of thermal process alarm data. Furthermore, datasets often contain multiple complex operating conditions, which complicates data processing. This makes it difficult for operators to accurately locate important alarm data and promptly resolve them, hindering the ability to suppress alarm overflows, exacerbating the "alarm flood" problem caused by the propagation of anomalies. Summary of the Invention

[0004] In response to the defects of the prior art, the purpose of the present invention is to provide a method and system for filtering thermal process alarm data based on the AHC-GP hybrid model, which can preprocess the data set, determine the optimal number of clusters using the nearest neighbor propagation algorithm, and then cluster the data set using the agglomerative hierarchical clustering algorithm to distinguish different working conditions; secondly, use the Gaussian process model to classify the data, and combine the posterior alarm probability estimate to achieve accurate filtering of thermal process alarm data, with good alarm data filtering performance, can accurately locate the key alarm data in the data set, and eliminate redundant alarms; the missed detection rate and false positive rate are both low, and the method and system have good data filtering accuracy.

[0005] To achieve the above object, a technical solution of the present invention is:

[0006] A thermal process alarm data filtering method based on AHC-GP hybrid model,

[0007] The following steps are involved:

[0008] The first step is to obtain the thermal process alarm data set;

[0009] In the second step, the thermal process alarm dataset in the first step is preprocessed, and the nearest neighbor propagation algorithm is used to determine the optimal number of clusters. Then, the agglomerative hierarchical clustering algorithm is used to cluster the thermal process alarm dataset to form a clustered dataset to distinguish different working conditions.

[0010] The third step is to build a Gaussian process model to classify the clustered data set in the second step to obtain a classified data set;

[0011] Step 4: For the classified data set in step 3, a data filtering model is constructed by combining the posterior alarm probability estimate;

[0012] The fifth step is to use the data filtering model in the fourth step to achieve accurate filtering of thermal process alarm data.

[0013] After continuous exploration and experimentation, the present invention provides a thermal process alarm data filtering method based on the AHC-GP hybrid model. The method can preprocess the data set, use the nearest neighbor propagation algorithm to determine the optimal number of clusters, and then use the agglomerative hierarchical clustering algorithm (AHC) to cluster the data set to distinguish different working conditions. Secondly, the Gaussian process model (GP) is used for data classification, and combined with the posterior alarm probability estimate value, a data filtering model is constructed to achieve accurate filtering of thermal process alarm data.

[0014] Furthermore, the present invention utilizes the nearest neighbor propagation algorithm and the agglomerative hierarchical clustering algorithm, and by constructing a Gaussian process model and a data filtering model, it can achieve accurate filtering of alarm data, accurately locate key alarm data in the data set, and eliminate redundant alarms; the missed detection rate and the false positive rate are both low, and it has good data filtering accuracy.

[0015] Furthermore, the alarm data filtering method of the present invention enables operators to accurately locate important alarm data and eliminate alarms in a timely manner, thereby suppressing alarm overflow and reducing the "alarm flooding" problem caused by abnormal propagation.

[0016] As preferred technical measures:

[0017] In the second step, the method for the neighbor propagation algorithm to determine the optimal number of clusters is as follows:

[0018] The optimal number of clusters is detected on the thermal process alarm data set to determine the optimal number of clusters.

[0019] As preferred technical measures:

[0020] In the second step, the agglomerative hierarchical clustering algorithm is a bottom-up hierarchical clustering method. Its method for clustering the thermal process alarm dataset is as follows:

[0021] First, a time label constraint is added to the thermal process alarm dataset to make it continuous in the time dimension, which is consistent with the actual operation characteristics of thermal process data and enhances the accuracy of model clustering.

[0022] Then, the thermal process alarm dataset is divided into initial clusters, and the clusters are merged according to the inter-cluster metric distance until the clustering requirements are met;

[0023] The entire thermal process alarm data set is used as the cluster center, and there is no need to determine the cluster center in advance, so that the clustering result will not be affected by the local optimal solution.

[0024] As preferred technical measures:

[0025] In the third step, the specific method of using the Gaussian process model to classify the cluster data set is as follows:

[0026] First, based on the characteristics of the alarm data of the thermal process, category labels are constructed. The category labels are divided into key alarm data labels and redundant alarm data labels.

[0027] Then, a prediction function is established to effectively find the category label y of any input cluster data set, and divide the cluster data set into key alarm data and redundant alarm data;

[0028] Where y∈{-1, 1}.

[0029] The Gaussian process model, a new research hotspot in machine learning, possesses powerful learning capabilities. Combining the advantages of Bayesian inference learning and kernel machine learning, the model's prior knowledge is intuitive, exhibits excellent generalization, and can fit arbitrary data. Its covariance function significantly reduces the computational complexity of model parameters and can also output probability information, resulting in excellent modeling results for complex, small-sample data.

[0030] As preferred technical measures:

[0031] The data points of the classification dataset are x i (i=1,2,…n), the corresponding category label variable is t=(t1,…,t n ) T .

[0032] For a data point x in a specific classification dataset N+1 , whose category label is t N+1 , and its posterior probability distribution is p(t N+1 |t); and introduce vector α N+1 The Gaussian process prior on the N+1 );

[0033] Vector α N+1 It is a Gaussian process. Based on the existing research experience and the characteristics of the data set, the Gaussian process model uses the variational inference method to perform Gaussian approximation and uses the logistic sigmoid function for transformation. After transformation, the vector α N+1 is a non-Gaussian random process, vector α N+1 The value becomes (0, 1);

[0034] Thus, t is determined accordingly. N+1 A non-Gaussian process on the training data t N , solve the data point x in the classification data set N+1 The posterior alarm probability estimate of .

[0035] As preferred technical measures:

[0036] The α N+1 The Gaussian process prior on is calculated as follows:

[0037]

[0038] Among them, C N+1 is the covariance matrix;

[0039] The posterior alarm probability estimate p(t N+1 =1|t N ) is calculated as follows:

[0040] p(t N+1 |t N )=∫p(t N+1 =1|α N+1 )p(α N+1 |t N )dα N+1 (twenty one)

[0041] p(t N+1 =1|α N+1 )=σ(α N+1 ) (twenty two).

[0042] As preferred technical measures:

[0043] In the fourth step, the data filtering model is constructed as follows:

[0044] First, the posterior alarm probability estimate of each data point in the classification data set is calculated;

[0045] Then, according to the size of the posterior alarm probability estimate, the classification data set is filtered to obtain the key alarm data points;

[0046] The calculation formula of the posterior alarm probability estimate is as follows:

[0047] p(y * ={-1, 1}|X, y, x * ) (twenty three)

[0048] Among them, X is the classification data set, y, y * is the category label, x * is the unlabeled test set sample, and p is the alarm probability estimate.

[0049] The standard deviation of the posterior alarm probability estimate σ ★ and mean μ * The calculation formula is as follows:

[0050]

[0051]

[0052] k * =k(X, x * ) (17)

[0053] k ** =k(x ★ , x ★ ) (18)

[0054] Among them: K is the kernel matrix of the classification data set, k * is the vector value of the new classification data set in the kernel function matrix, I represents an n×n identity matrix, σ n is the standard deviation of the nth data point.

[0055] As preferred technical measures:

[0056] At the same time, the accuracy, false positive rate and missed detection rate are used to evaluate the filtering effect of the data filtering model.

[0057] The calculation formula of the accuracy is as follows:

[0058] ACR=correct number / actual alarm data number;

[0059] The calculation formula of the misjudgment rate is as follows:

[0060] MDR = number of false positives / number of true alarm data;

[0061] The calculation formula of the missed detection rate is:

[0062] MJR = number of missed detections / number of actual alarm data.

[0063] To achieve the above object, the second technical solution of the present invention is:

[0064] A thermal process alarm data filtering method based on AHC-GP hybrid model,

[0065] The following steps are involved:

[0066] Step 1: For the thermal process alarm data set, the nearest neighbor propagation clustering algorithm is used to determine the optimal number of clusters n;

[0067] Step 2: Based on the optimal number of clusters n in step 1, repeat steps 3 to 6 to cluster the thermal process alarm dataset;

[0068] Step 3: Calculate the category and proximity value ESS of each data point in the thermal process alarm data set;

[0069]

[0070] Among them, x i It is the data point of thermal process alarm data set;

[0071] Step 4: Enumerate all binomial clusters and merge the data according to the ESS in step 3, and then calculate the total ESS after merging;

[0072] Step 5: Select the two clusters with the smallest total proximity value ESS in step 4 and merge them;

[0073]

[0074] Step 6: Loop step 3 to step 5 until the category meets the requirements;

[0075] Step 7: For each cluster n∈N selected in step 6, repeat the following steps 8 to 9;

[0076] Step 8: Select kernel function RBF: The prior mean is set to 0 and the Gaussian process classification algorithm is used for training;

[0077] Step 9: For the data trained in step 8, use the variational inference method to obtain the Gaussian approximation, use the logistic sigmoid function to perform discrete transformation, and calculate the posterior alarm probability estimate:

[0078] P(t N+1 =1|X N+1 , t N )=∫P[t N+1 =1|f(x N+1 )]

[0079] P[(xN+1 )|X N+1 , t N ]df(x N+1 )

[0080] Among them, X N+1 is a classification data set, x N+1 is the data point of the classification dataset, t N is the corresponding category label;

[0081] The function f(x) is a Gaussian process, which is transformed by the logistic sigmoid function. After the transformation, the function becomes a non-Gaussian random process, and the function value becomes (0, 1);

[0082] Step 10: Based on the size of the posterior alarm probability estimate in Step 9, the thermal process alarm data set is filtered to obtain key alarm data points.

[0083] To achieve the above object, the third technical solution of the present invention is:

[0084] A thermal process alarm data filtering system based on AHC-GP hybrid model,

[0085] It includes:

[0086] one or more processors;

[0087] a storage device for storing one or more programs;

[0088] When the one or more programs are executed by the one or more processors, the one or more processors implement the thermal process alarm data filtering method based on the AHC-GP hybrid model as described above.

[0089] Compared with the prior art, the present invention has the following beneficial effects:

[0090] After continuous exploration and experimentation, the present invention provides a thermal process alarm data filtering method based on the AHC-GP hybrid model. The method can preprocess the data set, use the nearest neighbor propagation algorithm to determine the optimal number of clusters, and then use the agglomerative hierarchical clustering algorithm to cluster the data set to distinguish different working conditions. Secondly, the Gaussian process model is used for data classification, and the posterior alarm probability estimate is combined to construct a data filtering model to achieve accurate filtering of thermal process alarm data.

[0091] Furthermore, the present invention utilizes the nearest neighbor propagation algorithm and the agglomerative hierarchical clustering algorithm, and by constructing a Gaussian process model and a data filtering model, it can achieve accurate filtering of alarm data, accurately locate key alarm data in the data set, and eliminate redundant alarms; the missed detection rate and the false positive rate are both low, and it has good data filtering accuracy.

[0092] Furthermore, the alarm data filtering method of the present invention enables operators to accurately locate important alarm data and eliminate alarms in a timely manner, thereby suppressing alarm overflow and reducing the "alarm flooding" problem caused by abnormal propagation. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] Figure 1 This is a data filtering flow chart based on the AHC-GP model of the present invention;

[0094] Figure 2 This is a simulation data set diagram of the present invention;

[0095] Figure 3 This is a comparison chart of the estimated values ​​of the alarm probability under various working conditions of the present invention. DETAILED DESCRIPTION

[0096] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0097] On the contrary, the present invention covers any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention as defined by the claims. Furthermore, to facilitate a better understanding of the present invention, certain specific details are described in detail below in the detailed description of the present invention. Those skilled in the art will be able to fully understand the present invention without these details.

[0098] An embodiment of the present invention:

[0099] A thermal process alarm data filtering method based on AHC-GP hybrid model,

[0100] The following steps are involved:

[0101] The first step is to obtain the thermal process alarm data set;

[0102] In the second step, the thermal process alarm dataset in the first step is preprocessed, and the nearest neighbor propagation algorithm is used to determine the optimal number of clusters. Then, the agglomerative hierarchical clustering algorithm is used to cluster the thermal process alarm dataset to form a clustered dataset to distinguish different working conditions.

[0103] The third step is to build a Gaussian process model to classify the clustered data set in the second step to obtain a classified data set;

[0104] Step 4: For the classified data set in step 3, a data filtering model is constructed by combining the posterior alarm probability estimate;

[0105] The fifth step is to use the data filtering model in the fourth step to achieve accurate filtering of thermal process alarm data.

[0106] After continuous exploration and experimentation, the present invention provides a thermal process alarm data filtering method based on the AHC-GP hybrid model. The method can preprocess the data set, use the nearest neighbor propagation algorithm to determine the optimal number of clusters, and then use the agglomerative hierarchical clustering algorithm to cluster the data set to distinguish different working conditions. Secondly, the Gaussian process model is used for data classification, and the posterior alarm probability estimate is combined to construct a data filtering model to achieve accurate filtering of thermal process alarm data.

[0107] Furthermore, the present invention utilizes the nearest neighbor propagation algorithm and the agglomerative hierarchical clustering algorithm, and by constructing a Gaussian process model and a data filtering model, it can achieve accurate filtering of alarm data, accurately locate key alarm data in the data set, and eliminate redundant alarms; the missed detection rate and the false positive rate are both low, and it has good data filtering accuracy.

[0108] Furthermore, the alarm data filtering method of the present invention enables operators to accurately locate important alarm data and eliminate alarms in a timely manner, thereby suppressing alarm overflow and reducing the "alarm flooding" problem caused by abnormal propagation.

[0109] A specific embodiment of the Gaussian process model kernel function of the present invention:

[0110] According to the data under different working conditions, RBF is selected as the kernel function of the Gaussian process model.

[0111] The Gaussian kernel function is a local kernel function that is good at extracting local features. It is a monotonic function of the Euclidean distance between two vectors, also known as the radial basis function (RBF). Its basic form is the following formula (14):

[0112] K(x i , x j )=exp(-||x i -x j || 2 / σ 2 ) (14)

[0113] Among them, σ is the core parameter, which represents the core radius and controls the radial range of action.

[0114] The Gaussian kernel function can well identify data clusters with high local density in the data. At the same time, the Gaussian kernel has strong interpolation ability. The initial input data set can be linearly separated in the high-dimensional feature space after mapping, which has a good clustering effect.

[0115] A specific embodiment of the Gaussian process prior and posterior of the present invention:

[0116] Based on a parameter-free Bayesian framework, the Gaussian process treats the underlying function f as an unknown random variable for model training. Because it considers all model parameters, there is no need to fix any parameter form. The Gaussian process modeling approach assumes that the function f(x, θ) contains some parameters and then mathematically calculates the optimal model parameters θ. The prior for the Gaussian process is typically set to a mean of 0, and the covariance function typically uses a kernel function to calculate the correlation between data samples.

[0117] For the function p(y, f) that follows a Gaussian distribution and has Gaussian noise of ε, all marginalizations involved can be calculated as similar to the posterior probability p(y * |X, y, x * ), and the calculation formulas for its mean and standard deviation are as follows (15) and (16):

[0118]

[0119]

[0120] k * =k(X, x * ) (17)

[0121] k ** =k(x ★ , x ★ ) (18)

[0122] Where: K is the kernel matrix of the sample data set, k * is the vector value of the new sample data in the kernel function matrix, y is the vector of all binary classification labels, and I represents an n×n identity matrix.

[0123] A specific embodiment of the Gaussian process classification of the present invention:

[0124] Gaussian process classification is improved upon Gaussian process regression. For a set of input training data, the posterior probability of the variable is modeled, resulting in a probability value in the interval (0, 1), thereby accurately predicting the category y. If y takes the value of (0, 1) or (-1, 1), the classification is binary. If y takes multiple integer values, the classification is multi-classified. The thermal process alarm data filtering studied in this paper can be considered a special binary classification problem.

[0125] Assume that the function f(x) is a Gaussian process and uses the logistic sigmoid function to transform it. After the transformation, the function becomes a non-Gaussian random process and the function value becomes (0, 1). Assume that the training set is x i (i=1,2,…n), the corresponding data marker variable is t=(t1,…,t n ) THere we take a data sample point x N+1 , the categorical variable is t N+1 , the posterior probability distribution is p(t N+1 |t), then introduce vector α N+1 The Gaussian process prior on the N+1 ). Thus, t is defined accordingly. N+1 A non-Gaussian process on N As a condition, its prediction distribution is solved. N+1 The form of the Gaussian process prior on is shown in formula (19), C N+1 is the covariance matrix:

[0126]

[0127] For the binary classification problem, since the relationship between the two satisfies the following formula, when calculating the predicted probability, only one of them needs to be calculated.

[0128] p(t N+1 =0|t N )+p(t N+1 =1|t N )=1 (20)

[0129] Next, we solve p(t N+1 =1|t N ) as an example, the solution formulas for its prediction distribution are shown in (21) and (22):

[0130] p(t N+1 |t N )=∫p(t N+1 =1|α N+1 )p(α N+1 |t N )dα N+1 (twenty one)

[0131] p(t N+1 =1|α N+1 )=σ(α N+1 ) (twenty two)

[0132] The model's subsequent parameters require Gaussian approximation. Common methods for obtaining Gaussian approximation include variational inference and expectation propagation. Based on existing research experience and the characteristics of the dataset, the model uses variational inference for Gaussian approximation.

[0133] A specific embodiment of the evaluation index of the present invention is:

[0134] The alarm data filtering of thermal processes can be regarded as a special binary classification problem, namely, critical alarm data and redundant alarm data. The principle of Gaussian process binary classification is to establish a prediction function model and effectively find the category label y for any input data, where y∈{-1,1}. The present invention defines the concept of posterior alarm probability estimation by combining the Gaussian process related theory in Definitions 2 and 3 with a given training data set, and performs data filtering analysis based on the size of the probability estimate. The specific formula is abbreviated as (23):

[0135] p(y * ={-1, 1}|X, y, x * ) (twenty three)

[0136] Among them, X is the training set sample, y, y * is the category label, x * is the unlabeled test set sample, and p is the alarm probability estimate.

[0137] At the same time, the accuracy, false positive rate and missed detection rate are used to describe the detection effect of the model algorithm. The specific evaluation indicators and definition formulas are shown in Table 1 below.

[0138] Table 1 Description of evaluation indicators

[0139]

[0140] A preferred embodiment of the present invention:

[0141] First, the present invention preprocesses the dataset, divides the dataset into a training dataset and a test dataset, and uses the nearest neighbor propagation algorithm to determine the optimal number of clusters. The nearest neighbor propagation algorithm (AP) has a good effect on the rapid clustering of high-dimensional, multi-category data and can also be used to determine the optimal number of clusters for the model. Secondly, after the number of clusters is determined, a time dimension constraint is added to the dataset, and an agglomerative hierarchical clustering algorithm is used for clustering to distinguish data under different working conditions in the dataset and increase the accuracy of subsequent data filtering.

[0142] Finally, a Gaussian process model was used for classification. To address the imbalance between critical and redundant alarm data, a posterior alarm probability estimation concept was defined. Based on the estimated alarm probabilities for each category at each point, critical alarm data was located and filtered, eliminating redundant alarms and reducing the occurrence of "alarm flooding." Based on the characteristics of the dataset and previous research results, the Gaussian kernel function (RBF) was selected as the kernel function in the Gaussian process model, and the variational inference method was used to obtain the Gaussian approximation.

[0143] like Figure 1 As shown, the specific algorithm flow of the present invention is as follows:

[0144] Input: training dataset x = {x1 ,...,x N}, training dataset class label t, test set, kernel function, hyperparameters.

[0145] Output: accuracy, false positive rate, missed positive rate, posterior alarm probability estimation, alarm data location.

[0146] Training starts:

[0147] Step 1: Use the neighbor propagation clustering algorithm to determine the optimal number of clusters n.

[0148] Step 2: Use the optimal number of clusters n and repeat steps 3 to 6 to cluster the data set.

[0149] Step 3: Calculate the ESS value for each category and the total.

[0150]

[0151] Step 4: Enumerate all binomial clusters and calculate the total ESS value after merging.

[0152] Step 5: Select the two clusters with the smallest total ESS value to merge.

[0153]

[0154] Step 6: Repeat steps 3 to 5 until the category meets the requirements.

[0155] Step 7: For each cluster n∈N, repeat steps eight to nine.

[0156] Step 8: Select RBF kernel function, The prior mean is set to 0 and the Gaussian process classification algorithm is used for training.

[0157] Step 9: Use variational inference to obtain Gaussian approximation, use logistic sigmoid function for discrete transformation, and calculate the posterior alarm probability estimate:

[0158] P(t N+1 =1|X N+1 , t N )=∫P[t N+1 =1|f(x N+1 )]

[0159] P[(x N+1 )|X N+1 , t N ]df(x N+1 )

[0160] Step 10: Input the test set. Return the accuracy, false positive rate, missed positive rate, posterior alarm probability, and alarm data location.

[0161] The training is over.

[0162] An application embodiment of the present invention:

[0163] The experiment of this invention uses the thermal process data set of a 1000MW unit in a power plant to verify the effectiveness of the proposed algorithm, and analyzes the experimental results. The sampling interval of the historical data used is set to 5s, the total sampling time is 25000s, and a total of 5000 groups. The main steam temperature and main steam pressure data are used to form the original data set. The total data volume of the training data set is 3000 groups, and the total data volume of the test data set is 2000 groups. The simulation data set image is as follows Figure 2 shown.

[0164] The proposed AHC-GP model first performs data preprocessing, cleaning and segmenting the dataset. The optimal number of clusters is tested using the Affinity Propagation (AP) algorithm, and the results show that the optimal number of clusters is 4.

[0165] Considering that thermal process data is time series data with strong temporal correlation, a time dimension constraint was added to the dataset to enhance the model clustering accuracy. Clustering was then performed using an agglomerative hierarchical clustering algorithm to distinguish between various operating conditions and prevent data from being mixed and affecting the alarm data filtering results. If clustering were performed directly using an agglomerative hierarchical clustering algorithm, the model would only consider the magnitude of each point value and not the engineering characteristics of thermal process data, namely, continuity in the temporal dimension. The data clustering results would repeatedly jump between categories and not reflect the actual thermal process operation conditions on site. By adding a time label constraint to the thermal process data, the data no longer repeatedly jumps between operating conditions over time. The temporal continuity of the process data is clearly demonstrated, which aligns with the operational characteristics of actual thermal process data and greatly enhances the accuracy of the model clustering.

[0166] The Gaussian process classification model is established using RBF (Gaussian kernel function) as the kernel function used in the present invention. The training data sets of each category are trained, the model parameters are adjusted to achieve the best detection effect, and then the test data sets are used for testing.

[0167] The estimated value of the posterior alarm probability under each working condition is shown in the figure below: Figure 3 As shown. Figure 3 As can be seen in the images, when critical alarm data is present, the probability estimate for that point is larger, indicating that it is important alarm data. The probability estimates for redundant alarm data points are all smaller. The model can filter critical alarm data points based on the size of the posterior alarm probability estimate for each point.

[0168] The following statistical analysis is conducted on the accuracy of the model under each working condition. The false detection rate and missed detection rate of the model under each working condition are shown in Table 2 below.

[0169] Table 2 Test results of various working conditions

[0170]

[0171] In order to verify the effectiveness of the method proposed in this invention, existing mature classification algorithms such as support vector machine classification (SVC) and K-nearest neighbor (KNN) were selected to compare the accuracy, false detection rate and missed detection rate. The specific algorithm performance is shown in Table 3 below.

[0172] Table 3 Algorithm results comparison table

[0173]

[0174] Comparison results show that the AHC-GP model of our invention outperforms other traditional, mature algorithms in terms of accuracy, missed detection rate, and false detection rate. Using the GP model alone can lead to a large number of data misjudgments. Using an agglomerative hierarchical clustering algorithm to differentiate operating conditions significantly reduces the misjudgment rate in data filtering, resulting in a high overall performance in filtering thermal process alarm data.

[0175] Therefore, the experimental results show that the AHC-GP hybrid model proposed in this invention has good alarm data filtering performance, can accurately identify the location of key alarm data in the data set, eliminate redundant alarms, and reduce the occurrence of the "alarm flooding" problem.

[0176] Furthermore, comparative tests comparing the proposed method with mature algorithms such as Gaussian process models, support vector machines, and K-nearest neighbor algorithms showed that the proposed method had low missed detection and false positive rates, and had high data filtering accuracy.

[0177] An embodiment of a device applying the method of the present invention:

[0178] A thermal process alarm data filtering system based on an AHC-GP hybrid model, comprising:

[0179] one or more processors;

[0180] a storage device for storing one or more programs;

[0181] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned thermal process alarm data filtering method based on the AHC-GP hybrid model.

[0182] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0183] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A thermal process alarm data filtering method based on the AHC-GP hybrid model, It is characterized by: The following steps are involved: The first step is to obtain the thermal process alarm data set; In the second step, the thermal process alarm dataset in the first step is preprocessed, and the nearest neighbor propagation algorithm is used to determine the optimal number of clusters. Then, the agglomerative hierarchical clustering algorithm is used to cluster the thermal process alarm dataset to form a clustered dataset to distinguish different working conditions. The agglomerative hierarchical clustering algorithm is a bottom-up hierarchical clustering method. Its method for clustering the thermal process alarm data set is as follows: First, a time label constraint is added to the thermal process alarm dataset to make it continuous in the time dimension, which is consistent with the actual operation characteristics of thermal process data. Then, the thermal process alarm dataset is divided into initial clusters, and the clusters are merged according to the inter-cluster metric distance until the clustering requirements are met; The entire thermal process alarm data set is used as the cluster center, and there is no need to determine the cluster center in advance, so the clustering result will not be affected by the local optimal solution; The third step is to build a Gaussian process model to classify the clustered data set in the second step to obtain a classified data set; The specific method of using the Gaussian process model to classify clustered data sets is as follows: First, based on the characteristics of the alarm data of the thermal process, category labels are constructed. The category labels are divided into key alarm data labels and redundant alarm data labels. Then, a prediction function is established to effectively find the category label y of any input cluster data set, and divide the cluster data set into key alarm data and redundant alarm data; Where y∈{-1,1}; Step 4: For the classified data set in step 3, a data filtering model is constructed by combining the posterior alarm probability estimate; The method for constructing the data filtering model is as follows: First, the posterior alarm probability estimate of each data point in the classification data set is calculated; Then, according to the size of the posterior alarm probability estimate, the classification data set is filtered to obtain the key alarm data points; The calculation formula of the posterior alarm probability estimate is as follows: p(and * ={-1,1}∣X,y,x * ) (23) Among them, X is the classification data set, y, y * is the category label, x * is the unlabeled test set sample, p is the alarm probability estimate; The standard deviation of the posterior alarm probability estimate σ * and mean μ * The calculation formula is as follows: k * =k(X,x * ) (17) k ** =k(x * ,x * ) (18) Among them: K is the kernel matrix of the classification data set, k * is the vector value of the new classification data set in the kernel function matrix, I represents an n×n identity matrix, σ n is the standard deviation of the nth data point; The fifth step is to use the data filtering model in the fourth step to achieve accurate filtering of thermal process alarm data.

2. The thermal process alarm data filtering method based on the AHC-GP hybrid model according to claim 1, characterized in that: In the second step, the method for the neighbor propagation algorithm to determine the optimal number of clusters is as follows: The optimal number of clusters is detected on the thermal process alarm data set to determine the optimal number of clusters.

3. The thermal process alarm data filtering method based on the AHC-GP hybrid model according to claim 1, characterized in that: The data points of the classification dataset are x i (i=1,2,…n), the corresponding category label variable is t=(t1,…,t n ) T ; For a data point x in a specific classification dataset N+1 , whose category label is t N+1 , its posterior probability distribution is p(t N+1 |t); and introduce vector a N+1 The Gaussian process prior on the N+1 ); Vector a N+1 It is a Gaussian process. Based on the existing research experience and the characteristics of the data set, the Gaussian process model uses the variational inference method to perform Gaussian approximation and uses the logistic sigmoid function for transformation. After transformation, the vector a N+1 is a non-Gaussian random process, vector a N+1 The value becomes (0,1); Thus, t is determined accordingly. N+1 A non-Gaussian process on the training data t N , solve the data point x in the classification data set N+1 The posterior alarm probability estimate of .

4. The thermal process alarm data filtering method based on the AHC-GP hybrid model according to claim 3, characterized in that: The a n+1 The Gaussian process prior on is calculated as follows: Among them, C N+1 is the covariance matrix; The posterior alarm probability estimate p(t N+1 =1|t N ) is calculated as follows: p(t N+1 ∣t N )=∫p(t N+1 =1∣a N+1 )p(a N+1 ∣t N )da N+1 (21) p(t N+1 =1∣a N+1 )=σ(a N+1 (22); Among them, σ is the core parameter, which represents the core radius and controls the radial range of action.

5. The thermal process alarm data filtering method based on the AHC-GP hybrid model according to claim 1, characterized in that: At the same time, the accuracy, false positive rate and missed detection rate are used to evaluate the filtering effect of the data filtering model; The calculation formula of the accuracy is as follows: ACR=correct number / actual alarm data number; The calculation formula of the misjudgment rate is as follows: MDR = number of false positives / number of true alarm data; The calculation formula of the missed detection rate is: MJR = number of missed detections / number of actual alarm data.

6. A thermal process alarm data filtering method based on the AHC-GP hybrid model, characterized in that: The following steps are involved: Step 1: For the thermal process alarm data set, the nearest neighbor propagation clustering algorithm is used to determine the optimal number of clusters n; Step 2: Based on the optimal number of clusters n in step 1, repeat steps 3 to 6 to cluster the thermal process alarm dataset; Step 3: Calculate the category and proximity value ESS of each data point in the thermal process alarm data set; Among them, x i It is the data point of thermal process alarm data set; Step 4: Enumerate all binomial clusters and merge the data according to the ESS in step 3, and then calculate the total ESS after merging; Step 5: Select the two clusters with the smallest total proximity value ESS in step 4 and merge them; Step 6: Loop step 3 to step 5 until the category meets the requirements; Step 7: For each cluster n∈N selected in step 6, repeat the following steps 8 to 9; Step 8: Select kernel function RBF: The prior mean is set to 0 and the Gaussian process classification algorithm is used for training; Step 9: For the data trained in step 8, use the variational inference method to obtain the Gaussian approximation, use the logistic sigmoid function to perform discrete transformation, and calculate the posterior alarm probability estimate: P(t N+1 =1∣X N+1 ,t N )=∫P[t N+1 =1∣f(x N+1 )] P[(x N+1 )∣X N+1 ,t N ]df(x N+1 ) Among them, X N+1 is a classification data set, x N+1 is the data point of the classification dataset, t N is the corresponding category label; The function f(x) is a Gaussian process, which is transformed by the logistic sigmoid function. After the transformation, the function becomes a non-Gaussian random process, and the function value becomes (0,1); Step 10: Based on the size of the posterior alarm probability estimate in Step 9, the thermal process alarm data set is filtered to obtain key alarm data points.

7. A thermal process alarm data filtering system based on the AHC-GP hybrid model, characterized in that: It includes: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a thermal process alarm data filtering method based on an AHC-GP hybrid model as described in any one of claims 1 to 6.