Industrial coal-fired boiler operating condition classification method and system based on data integration and clustering

Through hybrid representation of nearest neighbor similarity, bipartite graph partitioning and third-order tensor integration, the problem of high computational complexity in traditional coal-fired boiler data processing is solved, efficient and accurate working condition division is achieved, and efficient operation and energy conservation and emission reduction of boilers are supported.

CN115169456BActive Publication Date: 2025-10-03UNIV OF JINAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210780261.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-04
Publication Date
2025-10-03
Estimated Expiration
2042-07-04

AI Technical Summary

Technical Problem

Traditional integrated clustering algorithms have high computational time complexity in coal-fired boiler data processing and cannot effectively construct affinity matrices, resulting in inaccurate operating condition division of boiler data, affecting the efficient operation of the boiler and energy conservation and emission reduction.

Method used

A mixed representative nearest neighbor method is used to construct a sparse affinity submatrix. Through bipartite graph partitioning and third-order tensor integration, the computational complexity is reduced, the clustering accuracy and robustness are improved, and efficient division of coal-fired boiler operating conditions is achieved.

Benefits of technology

By using hybrid representation of nearest neighbor similarity, bipartite graph partitioning, and third-order tensor integration, the computational time complexity of traditional clustering methods is reduced, the accuracy and robustness of coal-fired boiler operating condition division are improved, and auxiliary decision-making for parameter adjustment is provided for workers, achieving more effective fault detection and energy conservation and emission reduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115169456B_ABST
    Figure CN115169456B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of data clustering technology, and provides a method and system for industrial coal-fired boiler operating condition division based on data integration clustering to obtain the operating status data of the industrial coal-fired boiler; the present invention performs integrated clustering operations on the industrial coal-fired boiler data through three steps of mixed representative nearest neighbor similarity, bipartite graph segmentation and third-order tensor integration, and realizes effective operating condition division of the boiler data; specifically, by constructing a sparse affinity submatrix through mixed representative nearest neighbor similarity, it can solve the problem that the traditional clustering method has too high computational time complexity and cannot effectively construct the affinity matrix of the coal-fired boiler data; by dividing the bipartite graph, the time for solving the characteristic problem is reduced; and by integrating the multi-base clustering results into a unified integrated clustering framework, the accuracy and robustness of the clustering are further improved while maintaining high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data clustering, and in particular relates to a method and system for dividing industrial coal-fired boiler operating conditions based on data integration clustering. Background Art

[0002] As society's demands for safety, energy conservation, and environmental protection continue to rise, accelerating energy conservation and emission reduction efforts while promoting the healthy and steady development of the industry and ensuring the safe operation of industrial coal-fired boilers has become a new topic and a serious challenge facing the future development of the industrial boiler industry. my country's energy structure is primarily coal-based, and the large-scale and widespread use of industrial coal-fired boilers consumes significant amounts of coal and is also a major source of soot pollution. However, the average operating efficiency of industrial coal-fired boilers is relatively low, and the industry has significant potential for energy conservation. For future development, energy conservation and emission reduction are both an urgent and long-term strategic priority for the industrial coal-fired boiler industry.

[0003] The inventors discovered that during coal-fired boiler production, every aspect of the boiler equipment continuously generates status data, which is collected and stored in real time in a backend database. However, many companies fail to tap into the hidden value of this data "treasure trove," simply storing it in computers and losing the value it generates. Coal-fired boiler data is characterized by large volumes and high dimensionality, making traditional ensemble clustering algorithms inadequate. Traditional clustering methods are computationally prohibitive and cannot effectively construct affinity matrices for coal-fired boiler data. Summary of the Invention

[0004] In order to solve the above problems, the present invention proposes a method and system for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering. The present invention has made corresponding technical innovations in the method for dividing the operating conditions of coal-fired boilers. By integrating and clustering large-scale industrial coal-fired boiler data, the relationship between various parameters and boiler operating conditions in the production process of coal-fired boilers is more clearly displayed, providing reliable technical support for subsequent workers to adjust parameters and maintain efficient operation of coal-fired boilers.

[0005] In order to achieve the above object, the present invention is implemented through the following technical solutions:

[0006] In a first aspect, the present invention provides a method for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering, comprising:

[0007] Obtain operating status data of industrial coal-fired boilers;

[0008] Based on the operation status data, an affinity matrix is ​​constructed;

[0009] For constructing affinity matrix, a sparse affinity sub-matrix is ​​constructed based on the mixed representation nearest neighbor method;

[0010] The sparse affinity submatrix is ​​used as a bipartite graph, and the bipartite graph is partitioned to obtain multiple base clustering results;

[0011] Based on the third-order tensor, the base clustering results are integrated into a unified integrated clustering framework to obtain the final clustering results and realize the operating condition division of industrial coal-fired boilers.

[0012] Furthermore, a plurality of candidate representatives are randomly selected from the running status data, and for the plurality of candidate representatives, a density peak clustering method is used to obtain a plurality of cluster center points as representative points.

[0013] Furthermore, the density peak clustering method determines the cluster center based on the local density of the data point and the distance from the data point to the nearest data point with a local density greater than that of the data point.

[0014] Furthermore, the number of representative points is further screened, the nearest representative is found on the screened representative points, and the K nearest neighbor method is used to find the K nearest representatives in the adjacent area of ​​the nearest representative.

[0015] Furthermore, representative points are further screened by adjusting the threshold of the density peak clustering method.

[0016] Furthermore, the bipartite graph has N+p nodes. Through the eigenvalue and eigenvector transfer formula, the bipartite graph with N+p nodes is transferred and split into a bipartite graph with p nodes. Let the first k feature pairs of the bipartite graph with p nodes when solving the feature problem be expressed as The first k feature pairs when solving the feature problem for a bipartite graph with N+p nodes are expressed as The eigenvalue and eigenvector transfer formula is:

[0017] γ i (2-γ i )=λ i ,

[0018] in, Transfer Matrix D X is a diagonal matrix and B is a sparse matrix.

[0019] Furthermore, an initial tensor is constructed by superimposing multiple base clustering results; the initial tensor is expanded into a matrix, and the initial tensor is folded into its three-dimensional mode to obtain three expanded matrices; the three expanded matrices are subjected to singular value decomposition:

[0020] Y (t) =A (t) ∑ (t) (B (t) )T

[0021] Among them, Y (t) is the t-th expansion matrix; A (t) is a left singular matrix; B (t) is a right singular matrix;

[0022] Reduce the matrix A to include the left singular (t) The dimension of each array of ; based on Σ (t) The singular values ​​in , keep each corresponding left singular matrix A (t) The dominant e t Left singular vector, the resulting matrix is the final clustering result.

[0023] In a second aspect, the present invention further provides an industrial coal-fired boiler operating condition classification system based on data integration and clustering, comprising:

[0024] The data acquisition module is configured to: obtain operating status data of the industrial coal-fired boiler;

[0025] The affinity matrix building module is configured to: build an affinity matrix based on the operating status data;

[0026] The sparse affinity submatrix construction module is configured to: construct a sparse affinity submatrix according to a mixed representation nearest neighbor method for constructing an affinity matrix;

[0027] The clustering module is configured to: use the sparse affinity submatrix as a bipartite graph, partition the bipartite graph, and obtain multiple base clustering results;

[0028] The partitioning module is configured to integrate the base clustering results into a unified integrated clustering framework based on the third-order tensor to obtain the final clustering results and realize the operating condition partitioning of industrial coal-fired boilers.

[0029] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for dividing industrial coal-fired boiler operating conditions based on data integration clustering described in the first aspect.

[0030] In a fourth aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering as described in the first aspect are implemented.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] 1. The present invention performs integrated clustering operations on industrial coal-fired boiler data through three steps: mixed representative nearest neighbor similarity, bipartite graph segmentation, and third-order tensor integration, thereby achieving effective operating condition division of the boiler data. Specifically, by constructing a sparse affinity submatrix through mixed representative nearest neighbor similarity, the problem that the traditional clustering method has too high computational time complexity and cannot effectively construct an affinity matrix for coal-fired boiler data can be solved. By partitioning the bipartite graph, the time for solving the characteristic problem is reduced. By integrating the multi-base clustering results into a unified integrated clustering framework, the accuracy and robustness of clustering are further improved while maintaining high efficiency. By using the three steps of mixed representative nearest neighbor similarity, bipartite graph segmentation, and third-order tensor integration, the computational time complexity of the traditional clustering method is greatly reduced.

[0033] 2. The final clustering result of the present invention can provide auxiliary decision-making for workers to adjust boiler parameters, achieve more effective fault detection, save energy and reduce emissions, and improve the safety of boiler production. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings constituting a part of the specification of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions of this embodiment are used to explain this embodiment and do not constitute an improper limitation on this embodiment.

[0035] Figure 1 This is a flow chart of Example 1 of the present invention;

[0036] Figure 2 The DPC of Example 1 of the present invention is calculated once and twice to select representative points for interpretation;

[0037] Figure 3 is a representative set R and an object x in embodiment 1 of the present invention i ∈X;

[0038] Figure 4 is a representative point after further screening in Example 1 of the present invention and an object x i ∈X;

[0039] Figure 5 is x in Example 1 of the present invention i The distance between the selected representative points;

[0040] Figure 6 is x in Example 1 of the present invention i The nearest representative point r l ;

[0041] Figure 7 is x in Example 1 of the present invention i and r l The distance between the K' nearest neighbors;

[0042] Figure 8 is x in Example 1 of the present invention i The K nearest representatives of (K=3);

[0043] Figure 9 This is a visualization of the third-order tensor integration process of Example 1 of the present invention. DETAILED DESCRIPTION

[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0045] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0046] The main steps in boiler production are: first, a coal feeder continuously feeds coal into the furnace for combustion. Then, a blower continuously supplies oxygen required for combustion, and the furnace temperature is controlled by controlling the air volume. The heat generated by the furnace combustion heats the steam drum, generating large amounts of high-temperature, high-pressure steam inside the drum, which then drives the steam turbine to generate electricity. Each of these production steps constitutes the operating status data of a coal-fired boiler, which includes data such as steam drum pressure, main steam temperature, bed temperature, primary air volume, and flue gas oxygen content.

[0047] According to the applicant's research on boiler production and communication with front-line workers, it was found that although the same boiler is used for thermal power production every year, the operating conditions of the boiler equipment are actually different every year due to the aging and maintenance of the boiler equipment. The patterns and models obtained through the analysis of historical data from previous years may not be applicable to the equipment operating conditions of the boiler in the second year. It is crucial to ensure the stability of the boiler's working state, so workers actually need to monitor the status of the boiler at all times during the production process and make timely adjustments. For example, the pressure of the steam drum, the temperature of the fuel layer, the oxygen content of the flue gas, the speed of the coal feeder, etc. are all very important links to the stability of the boiler state. They are also the guarantee and prerequisite for the normal and safe operation of the boiler. These parameters need to be monitored at all times and adjusted accordingly.

[0048] Example 1:

[0049] This embodiment provides a method for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering, including:

[0050] Obtain operating status data of industrial coal-fired boilers;

[0051] Based on the operation status data, an affinity matrix is ​​constructed;

[0052] For constructing an affinity matrix, a sparse affinity sub-matrix is constructed according to the hybrid representative nearest neighbor approximation method.

[0053] The sparse affinity sub-matrix is regarded as a bipartite graph, and the bipartite graph is partitioned to obtain multiple basic clustering results.

[0054] Based on a third-order tensor, the basic clustering results are integrated into a unified integrated clustering framework to obtain the final clustering result, realizing the condition division of industrial coal-fired boilers.

[0055] It can be understood that the operating state data can be the state data generated during the production process of the coal-fired boiler collected; based on the state data, the condition division of the coal-fired boiler data is carried out, which can improve the accuracy of the condition division of the boiler data and play an auxiliary optimization role in the parameter regulation operation of the workers during the operation of the coal-fired boiler; after collecting the data, the specific content of this embodiment is:

[0056] S1. Hybrid representative nearest neighbor approximation: A hybrid representative selection strategy combining random selection and the clustering by fast search and find of density peaks (DPC) algorithm is used to select representative points, thereby reducing the scale of the data. Then, the K nearest neighbor approximation method is used to quickly and effectively obtain the K nearest representatives of the sample points, and other representatives are set to zero, further sparsifying the affinity sub-matrix.

[0057] For the large-scale coal-fired boiler data after collection and integration, an affinity matrix needs to be constructed according to the similarity between the data. However, the traditional clustering method has too high computational time complexity and cannot effectively construct the affinity matrix of the coal-fired boiler data. Therefore, this embodiment provides a hybrid representative nearest neighbor approximation method to quickly and effectively construct a sparse affinity sub-matrix; first, a hybrid representative selection strategy combining random selection and the DPC algorithm is used to select representative points, and then the K nearest neighbor approximation method is used to quickly and effectively obtain the K nearest representatives of the sample points, and other representatives are set to zero, further sparsifying the affinity sub-matrix.

[0058] Let X = {x1, x2,..., x N} represent a coal-fired boiler data set containing N objects, where x i ∈R d is the i-th object and d is the dimension. First, P` candidate representatives are randomly selected from the boiler data to reduce the time cost (p` << N). Then, for the P` candidate points, this embodiment uses the DPC method to obtain P clustering center points and uses the clustering center points as representative points. DPC determines the clustering centers according to ρ i and δ i [[ID=i , the distance from a data point to the nearest data point with a local density greater than it is δ i .

[0059]

[0060]

[0061] ρ i ·δ i >ε (3)

[0062] Among them, d ij is the distance between data point i and data point j; d c is the cutoff distance; χ is the logic judgment function; ε is the threshold value, which is defined in this embodiment. When a data point ρ i and δ i When the product of is greater than the threshold ε, the point is the cluster center. Therefore, the number of cluster centers can be controlled by adjusting the threshold ε. In form, this embodiment expresses the set of selected representatives as:

[0063] R={r1,r2,..r p} (4)

[0064] After obtaining p representatives, the next goal is to encode the pairwise relationships of the entire dataset through a small set of representatives. In this embodiment, a K-nearest neighbor approximation based on a coarse-to-fine mechanism is used to construct a sparse affinity submatrix. The main idea of ​​the K-nearest representative approximation in this embodiment is to first further filter the number of representative points, and then find the nearest representative on the filtered representative points, denoted as r l , then in r l Find the K nearest representatives in the adjacent area. This embodiment further screens representative points by adjusting the DPC threshold ε, so that two selections can be performed in one calculation, reducing time complexity. By obtaining the K nearest representatives of each object, a sparse N×p sparse affinity submatrix can be constructed. This embodiment uses a Gaussian kernel as the similarity kernel. Therefore, the sparse affinity submatrix can be expressed as:

[0065] B={b ij} N×p (5)

[0066]

[0067] Among them, N K (x i ) represents x iThe kernel parameter σ is set to the average Euclidean distance between the objects and their K nearest representatives; B is a sparse matrix that contains only NK non-zero entries.

[0068] like Figure 2 As shown in the figure, the DPC calculation is performed twice to explain the selection of representative points. The first time the threshold ε is set to 4000, 1000 representative points are selected, as shown in the upper right area of ​​the blue line. The second time the threshold ε is set to 12000, and 7 representative points are further selected, as shown in the upper right area of ​​the red line. This method reduces the amount of calculation required to find the nearest representative point to the sample point.

[0069] S2. Bipartite graph partitioning: The sparse affinity submatrix reflects the relationship between data objects and representative points, and can be naturally interpreted as a bipartite graph. By utilizing the structure of the bipartite graph, the transitive cut can be used to effectively partition the graph and obtain the final clustering result.

[0070] The sparse affinity matrix B reflects the relationship between the objects in X and their representatives in R. It can be naturally interpreted as a bipartite graph G = {X, R, B}, where XUR is the node set (U represents the union of the two sets X and R) and B is the cross-affinity matrix. Leveraging the structure of the bipartite graph, transfer cuts can be used to effectively partition the graph and obtain the final clustering results.

[0071] First, if we regard the bipartite graph G as a general graph with N+p nodes, then its full affinity matrix can be written as:

[0072]

[0073] Clustering attempts to partition a graph by solving the following generalized characteristic problem:

[0074] Lu=γDu (8)

[0075] Where L = DE is the Laplace operator of the bipartite graph G, D = R (N+p)×(N+p) is the degree matrix; γ is the eigenvalue; u is the eigenvector. By treating the bipartite graph G as a general graph, it takes O((N+p) 3 ) to solve the feature problem equation (8), which is not feasible on a very large-scale boiler dataset. O is the representative symbol of time complexity.

[0076] Therefore, in this embodiment, transfer cutting is used to reduce the time of solving the characteristic problem, and the characteristic problem of solving the bipartite graph G is transferred to solving a smaller graph G R The characteristic problem on the bipartite graph G has N+p nodes, and the graph G R There are p nodes. Specifically, the graph G R ={R,ER}, where R is the representative point set; is the affinity matrix, D X ∈R N×N It is a diagonal matrix, and the numbers on the diagonal are the sum of the numbers in each row of B. R =D R -E R is the Laplace operator, D R ∈R p×p It's G R Then, the graph G R The generalized characteristic problem on can be expressed as

[0077] L R v=λD R v (9)

[0078] Assume that the first k feature pairs of feature problem (9) are expressed as The first k feature pairs of feature problem (8) are expressed as The eigenvalue and eigenvector transfer formula is:

[0079] γ i (2-γ i )=λ i (10)

[0080]

[0081]

[0082] in, is the transfer matrix; after solving the eigenvalue problem, the k eigenvectors are stacked into a (N+p)×k matrix. Each row of the matrix is ​​used as a new eigenvector, and k-means is performed on this basis using the N rows corresponding to the N original objects to obtain the final clustering result.

[0083] S3. Third-Order Tensor Integration: To further improve clustering accuracy and robustness while maintaining high efficiency, this embodiment integrates multi-base clustering results into a unified ensemble clustering framework to further improve clustering accuracy and robustness. Third-Order Tensor Integration primarily leverages the advantages of tensor decomposition to mine hidden layer information in the data, resulting in better and more robust consistent clustering.

[0084] This embodiment provides an integration algorithm based on third-order tensors, which integrates the results of multi-base clustering into a unified integrated clustering framework, aiming to further improve the accuracy and robustness of clustering while maintaining high efficiency. Third-order tensor integration mainly uses the advantages of tensor decomposition to mine the hidden layer information of data, thereby obtaining better and more robust consistent clustering. Each base cluster consists of a certain number of clusters. In this embodiment, the clusters in each base cluster are represented as

[0085] C={C1,C2,…,C k} (13)

[0086] Among them, C i is the i-th cluster in the base cluster, and k is the number of clusters in the base cluster. The result of base clustering can be expressed as:

[0087] W={w ij} N×k (14)

[0088]

[0089] Where W is a sparse matrix containing only N non-zero entries. According to five main steps, this embodiment obtains a better clustering result:

[0090] S3.1. Initial construction of tensor y: This embodiment can obtain an initial third-order tensor in, Construct a tensor y∈R by superimposing P base clustering results N×K×P .

[0091] S3.2, Matrix expansion of tensor y: by arranging the unit elements corresponding to y into Y (t) The columns of t = 1, 2, 3, the tensor y can be expanded into a matrix. The original tensor y can be folded into its three-dimensional mode. (1) ∈R N×KP , Y (2) ∈R K×NP , Y (3) ∈R NK×P .

[0092] S3.3. Singular value decomposition (SVD) on each unfolded matrix: SVD decomposition of three unfolded matrices Y(t) (t=1, 2, 3):

[0093] Y (t) =A (t) ∑ (t) (B (t) ) T t=1,2,3 (16)

[0094] In order to better show the potential correlation between clustering results of different bases, it is necessary to reduce the left singular matrix A (t) The dimension of each array of (t=1, 2, 3). Based on Σ (t) The singular values ​​in , keep each corresponding matrix A (t) The dominant e in (t=1, 2, 3) t The left singular vector, the resulting matrix is Parameter e t The value is usually obtained by Σ (t) Choose based on the knowledge in .

[0095] S3.4. Construction of the core tensor G: The correlation between the three different modalities can be dominated by the core tensor G and performed:

[0096]

[0097] Here y is the initial tensor, × t represents modular multiplication of third-order tensors, yes The transpose of , G is a tensor of e1×e2×e3.

[0098] S3.5, Reconstructed tensor y`: This step verifies that the reconstructed tensor is similar to the original tensor. Reconstructed tensor y`:

[0099]

[0100] y` is approximately equal to y, which means is small. In addition, the potential correlation between clustering results of different bases is contained in a subset of left singular vectors. Finally, the clustering result WF is expressed as

[0101] The result of the third-order tensor integration is similar to that of fuzzy clustering. The value of each row in the matrix is ​​the probability that the sample point belongs to a cluster, and the cluster corresponding to the maximum value in a row is the cluster to which the sample point belongs. Therefore, the final clustering result is expressed as

[0102] WF={wf ij} N×k (19)

[0103]

[0104] Among them, wf ij The value of 1 means that the i-th sample point belongs to the j-th cluster, wf ij The value of 0 means that the i-th sample point does not belong to the j-th cluster, and finally WF contains only N non-zero items. It can be understood that formula (20) is to change the maximum value of each row of the matrix WF to 1 and the rest to 0.

[0105] This embodiment processes state data generated during the production process of industrial coal-fired boilers. A corresponding algorithmic process is designed to capture the relationship between various parameters and operating conditions during the coal-fired boiler production process, providing reasonable scientific and technological support for workers to adjust parameters and maintain efficient coal-fired boiler operation. A nearest representative proximity method is provided for coal-fired boiler data, reducing the data size and quickly and efficiently constructing sparse affinity submatrices between the data. Clustering results are obtained through bipartite graph partitioning. A third-order tensor-based ensemble method is then used to optimize the clustering results, improving the accuracy of coal-fired boiler operating condition classification.

[0106] Example 2:

[0107] This embodiment provides an industrial coal-fired boiler operating condition classification system based on data integration and clustering, including:

[0108] The data acquisition module is configured to: obtain operating status data of the industrial coal-fired boiler;

[0109] The affinity matrix building module is configured to: build an affinity matrix based on the operating status data;

[0110] The sparse affinity submatrix construction module is configured to: construct a sparse affinity submatrix according to a mixed representation nearest neighbor method for constructing an affinity matrix;

[0111] The clustering module is configured to: use the sparse affinity submatrix as a bipartite graph, partition the bipartite graph, and obtain multiple base clustering results;

[0112] The partitioning module is configured to integrate the base clustering results into a unified integrated clustering framework based on the third-order tensor to obtain the final clustering results and realize the operating condition partitioning of industrial coal-fired boilers.

[0113] The working method of the system is the same as the method for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering in Example 1, and will not be described in detail here.

[0114] Example 3:

[0115] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method for dividing the operating conditions of industrial coal-fired boilers based on data integration and clustering described in Example 1 are implemented.

[0116] Example 4:

[0117] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method for dividing the operating conditions of industrial coal-fired boilers based on data integration clustering described in Example 1 are implemented.

[0118] The above description is merely a preferred embodiment of this embodiment and is not intended to limit this embodiment. Those skilled in the art will readily appreciate that this embodiment may be modified and varied in various ways. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this embodiment shall be within the scope of protection of this embodiment.

Claims

1. A method for classifying industrial coal-fired boiler operating conditions based on data integration and clustering, characterized by: include: Obtain operating status data of industrial coal-fired boilers; Based on the operation status data, an affinity matrix is ​​constructed; To construct the affinity matrix, a sparse affinity sub-matrix is ​​constructed based on the mixed representation nearest neighbor method, specifically: Multiple candidate representatives are randomly selected from the running status data. For each candidate representative, a density peak clustering method is used to obtain multiple cluster center points as representative points. The number of representative points is further filtered by adjusting the threshold of the density peak clustering method. The nearest representative is found on the filtered representative points, and the K nearest neighbor method is used to find the K nearest representatives in the adjacent area of ​​the nearest representative. The sparse affinity submatrix is ​​used as a bipartite graph, and the bipartite graph is partitioned to obtain multiple base clustering results; Based on the third-order tensor, the base clustering results are integrated into a unified integrated clustering framework to obtain the final clustering results and realize the operating condition division of industrial coal-fired boilers.

2. The method for dividing industrial coal-fired boiler operating conditions based on data integration and clustering according to claim 1 is characterized in that: The density peak clustering method determines the cluster center based on the local density of the data point and the distance from the data point to the nearest data point with a local density greater than that of the data point.

3. The method for dividing industrial coal-fired boiler operating conditions based on data integration and clustering according to claim 1 is characterized in that: A bipartite graph has N + p Node, passing the formula through eigenvalues ​​and eigenvectors, will have N + p The bipartite graph with nodes is partitioned into p nodes; let p The previous step of solving the characteristic problem of a bipartite graph with 10 nodes k The feature pairs are represented as ; have N + p The previous step of solving the characteristic problem of a bipartite graph with 10 nodes k The feature pairs are represented as ; The eigenvalue and eigenvector transfer formula is: , in, , the transfer matrix , is a diagonal matrix, is a sparse matrix.

4. The method for dividing industrial coal-fired boiler operating conditions based on data integration and clustering according to claim 1 is characterized in that: An initial tensor is constructed by superimposing multiple basis clustering results; the initial tensor is expanded into a matrix, and the initial tensor is folded into its three-dimensional mode to obtain three expanded matrices; the three expanded matrices are subjected to singular value decomposition: in, For the t An expansion matrix; is a left singular matrix; is a right singular matrix; Reduce to include left singular matrices The dimensions of each array based on Σ (t) The singular values ​​in , keep each corresponding left singular matrix The leading e t Left singular vector, the resulting matrix is the final clustering result.

5. The industrial coal-fired boiler operating condition classification system based on data integration and clustering is characterized by: include: The data acquisition module is configured to: obtain operating status data of the industrial coal-fired boiler; The affinity matrix building module is configured to: build an affinity matrix based on the operating status data; The sparse affinity submatrix construction module is configured to: construct a sparse affinity submatrix based on the constructed affinity matrix and according to the mixed representation nearest neighbor method, specifically: Multiple candidate representatives are randomly selected from the running status data. For each candidate representative, a density peak clustering method is used to obtain multiple cluster center points as representative points. The number of representative points is further filtered by adjusting the threshold of the density peak clustering method. The nearest representative is found on the filtered representative points, and the K nearest neighbor method is used to find the K nearest representatives in the adjacent area of ​​the nearest representative. The clustering module is configured to: use the sparse affinity submatrix as a bipartite graph, partition the bipartite graph, and obtain multiple base clustering results; The partitioning module is configured to integrate the base clustering results into a unified integrated clustering framework based on the third-order tensor to obtain the final clustering results and realize the operating condition partitioning of industrial coal-fired boilers.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for dividing the operating conditions of industrial coal-fired boilers based on data integration clustering as described in any one of claims 1 to 4 are implemented.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the industrial coal-fired boiler operating condition division method based on data integration clustering as described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Self-adaptive clustering center density peak value clustering algorithm based on weighted shared nearest neighbor

    CN113222027A

  • Text clustering integration method and system based on three-layer weighting model

    CN114281994A