Grate cooler working condition identification method for cement production based on Gaussian mixture model and K-means fusion
By combining the fusion strategy of Gaussian mixture model and K-means, and using the XGboost model for grate cooler operating condition identification, the problem of difficult accurate division of operating conditions in traditional methods is solved, achieving more efficient and accurate operating condition identification, and supporting intelligent control and optimization of grate coolers.
Patent Information
- Application Number
- CN202510890778.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional grate cooler operating condition identification methods cannot achieve precise control when dealing with complex and variable industrial environments, especially due to the difficulty in accurately classifying operating conditions caused by nonlinearity, multivariable coupling, and hysteresis effects.
A method based on the fusion of Gaussian mixture model and K-means was adopted. By collecting grate cooler operating data, preprocessing, feature extraction and clustering were performed. The XGboost model was then used for operating condition identification, including Z-score algorithm to filter out outliers, local weighted regression and sliding window method to extract features, K-means initial clustering and Gaussian mixture model optimization, and XGboost model was constructed for training and identification.
It improves the accuracy and efficiency of grate cooler operating condition identification, can better handle ambiguous boundary areas, avoid hard classification errors, improve the stability and accuracy of clustering results, and provide support for intelligent control and optimization.
Smart Images

Figure CN120873658A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operating condition identification methods for grate coolers used in cement production, specifically a method for identifying operating conditions of grate coolers used in cement production based on the fusion of Gaussian mixture model and K-means. Background Technology
[0002] In the research of grate coolers used in cement production, the identification and classification of operating conditions are crucial to ensuring their efficient operation and energy optimization. Traditional methods for identifying grate cooler operating conditions typically rely on expert experience or simple physical models, which are often insufficiently accurate for complex and variable industrial environments. For example, while parameters such as grate thickness, under-grate pressure, and kiln main unit current can indirectly reflect the equipment's operating status, the nonlinearity, multivariate coupling, and hysteresis effects in industrial processes often prevent precise control through the classification and identification of operating conditions based on these traditional methods.
[0003] In recent years, with the development of data mining and machine learning technologies, data-driven clustering methods have been proposed and gradually applied to the identification of grate cooler operating conditions. These methods automatically extract features from large amounts of real-time data and perform clustering, thereby achieving accurate identification of different operating conditions of grate coolers.
[0004] While traditional clustering methods can effectively address the problem of identifying the operating conditions of grate coolers, several limitations remain. Grate cooler operating data exhibits significant volatility and nonlinearity, including periodic fluctuations and abrupt changes, and the data is distributed across different density regions. Traditional K-means clustering assumes spherical clusters with uniform variance, which is insufficient for handling complex, asymmetric data like grate cooler data. K-means is sensitive to outliers and requires a pre-set number of clusters K, making it unsuitable for data with strong volatility and pronounced temporal characteristics. In contrast, Gaussian Mixture Models (GMMs) model cluster shape using a Gaussian distribution, handling non-spherical clusters and providing the probability of each data point's affiliation, adapting to density differences and outliers in the data. However, GMMs have higher computational complexity, are sensitive to initial parameters, and are prone to getting trapped in local optima. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for identifying the operating conditions of a grate cooler used in cement production based on the fusion of Gaussian mixture model and K-means, in order to solve the problem that the operating conditions of the grate cooler are difficult to accurately classify due to the nonlinearity and strong coupling characteristics of the cooling process.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for identifying the operating conditions of a grate cooler used in cement production, based on the fusion of Gaussian mixture model and K-means, includes the following steps:
[0008] Step 1: Collect the operating condition data of the grate cooler at the production site, and select the key variable operating condition data;
[0009] Step 2: Preprocess the selected key variable operating condition data to filter out outlier values at shutdown points;
[0010] Step 3: Using a Gaussian mixture model, the key variable operating condition data after filtering out outliers in Step 2 are divided into two categories: typical operating condition data and abnormal operating condition data.
[0011] Step 4: Extract typical features from the typical operating condition data obtained in Step 3;
[0012] Step 5: Use a fusion model consisting of K-means clustering algorithm and Gaussian mixture model to cluster the typical working condition data obtained in Step 3. During clustering, first use K-means clustering algorithm to cluster the typical working condition data to obtain coarse working condition clustering results, and then use Gaussian mixture model to further optimize the coarse working condition clustering results obtained by K-means clustering algorithm, thereby obtaining the final working condition clustering results.
[0013] Step 6: Construct the XGboost model. Use the key features of the typical working condition data obtained in Step 4 as the input features of the XGboost model, and use the final working condition clustering results obtained in Step 5 as the target label to train the XGboost model. The trained XGboost model is then used as the working condition recognition model.
[0014] Step 7: Input the operating condition data of the grate cooler to be identified into the operating condition identification model obtained in Step 6, and the operating condition identification model outputs the operating condition identification result.
[0015] In further step 1, the key variables include grate pressure, grate speed, kiln main unit current, raw material feed rate, and tertiary air temperature.
[0016] In the further step 2, the Z-score algorithm is used to preprocess the selected key variable operating condition data to filter out outlier values at shutdown points.
[0017] In the further step 4, the trend characteristics, slope characteristics, and fluctuation characteristics of the grate pressure data in the typical working condition data obtained in step 3 are extracted as the typical characteristics.
[0018] Furthermore, a locally weighted regression algorithm is used to extract trend features from the grate data.
[0019] Furthermore, the sliding window method is used to extract slope features from the grate compression data.
[0020] Furthermore, the fluctuation characteristics are obtained by calculating the standard deviation and variance of the grate compression data.
[0021] This invention can reasonably handle outliers at shutdown points. By dividing abnormal and typical operating conditions into data, it extracts typical features based on typical operating conditions and uses a hybrid clustering strategy of Gaussian mixture model and K-means to achieve detailed clustering of typical operating conditions.
[0022] By combining the K-means algorithm with the Growing Model (GMM), K-means is first used for initial clustering, followed by GMM for further optimization. This approach ensures computational efficiency while improving clustering accuracy and robustness. This combined strategy better handles regions with ambiguous boundaries, avoids hard classification errors, effectively identifies outliers, and enhances the stability and accuracy of clustering results. Therefore, combining the advantages of both approaches has significant application potential for clustering analysis of complex grate cooler operating data.
[0023] This invention presents a method for identifying the operating conditions of a grate cooler used in cement production, based on the fusion of Gaussian mixture model and K-means. This method is an effective way to improve the accuracy and efficiency of operating condition identification for grate coolers and can provide strong support for the intelligent control and optimization of grate coolers. Attached Figure Description
[0024] Figure 1 This is a flowchart of the method according to an embodiment of the present invention. Figure 2 This is the clustering result obtained from the fusion model. Detailed Implementation
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0026] like Figure 1 As shown in the figure, this embodiment discloses a method for identifying the operating conditions of a grate cooler used in cement production based on the fusion of Gaussian mixture model and K-means, including the following steps:
[0027] Step 1: Collect the operating condition data of the grate cooler at the production site, and select the key variable operating condition data.
[0028] In this embodiment, operating data of the grate cooler at the cement production site was collected over a month at 40-second intervals. Operating data for several key variables were then selected for subsequent operating condition identification and analysis.
[0029] In this embodiment, the key variables selected include grate pressure, grate speed, kiln main unit current, raw material feed rate, and tertiary air temperature. Among them:
[0030] The undergrate pressure reflects the cooling effect of the cooling air on the clinker and is closely related to the thickness of the material layer, making it an important indicator for evaluating the operating status of the grate cooler. The grate speed is inversely proportional to the undergrate pressure, directly affecting the material layer thickness and cooling efficiency. The kiln main unit current is a crucial indicator of the stability of the rotary kiln's operation, thus also influencing the grate cooler. The raw material feed rate is a key factor determining the load on the grate cooler, affecting cooling efficiency and kiln temperature changes, and therefore plays a vital role in operational analysis. The tertiary air temperature is primarily used to evaluate the heat recovery effect of the grate cooler, helping to further understand the heat exchange efficiency of the cooling air and thus guiding adjustments to the grate cooler's operating conditions.
[0031] By selecting the operating data of these key variables, more accurate data support can be provided for the classification and identification of the operating conditions of the grate cooler.
[0032] Step 2: Use the Z-score algorithm to preprocess the selected key variable operating condition data to filter out outliers at shutdown points and ensure the accuracy of subsequent analysis.
[0033] In this embodiment, the Z-score algorithm determines whether a data point is an outlier by calculating the standard deviation distance between each data point in the key variable operating condition data and the data mean. The calculation formula is as follows:
[0034]
[0035] Where: μ is the mean of the data, σ is the standard deviation of the data, and x i Let i be the i-th data point in the key variable operating condition data.
[0036] A threshold of 3 is set, and outliers are determined based on this threshold. If the absolute value of a data point's Z-score exceeds the threshold of 3, that data point is considered an outlier, potentially indicating a shutdown point or a data acquisition error. Finally, the outliers are removed from the critical variable operating condition data to ensure its accuracy and improve the reliability of subsequent analysis and model training.
[0037] Using the Z-score algorithm to process outliers at shutdown points can effectively clean the data and ensure the accuracy of subsequent operating condition identification.
[0038] Step 3: Using a Gaussian mixture model, the key variable operating condition data after outlier filtering in Step 2 are divided into two categories: typical operating condition data and abnormal operating condition data.
[0039] In this embodiment, abnormal operating conditions include shutdown points and material shortage recovery conditions. A shutdown point typically indicates that the equipment has stopped operating. Near the shutdown point are conditions of material shortage and recovery, where key variables such as grate pressure and kiln main unit current may exceed reasonable ranges, exhibiting characteristics inconsistent with normal operating conditions. Typical operating conditions refer to the operating data of various key variables of the grate cooler under normal operating conditions.
[0040] Because the Z-score algorithm in step 2 alone cannot completely separate the material interruption recovery condition from the normal operating conditions. Furthermore, the data density of the material interruption recovery condition is not consistent with that of the normal typical operating conditions, while the Gaussian Mixture Model (GMM) clustering process can distinguish between the two and automatically identify the abnormal material interruption recovery condition. Therefore, in this embodiment, the Gaussian Mixture Model is used to distinguish between the abnormal and typical operating conditions.
[0041] By using the soft partitioning mechanism of the Gaussian mixture model, typical and abnormal operating conditions can be accurately distinguished based on the probability distribution of the key variable operating condition data after outlier filtering in step 2, providing an accurate data foundation for subsequent analysis.
[0042] Step 4: Extract typical features from the typical working condition data obtained in Step 3.
[0043] In this embodiment, considering that the grate pressure data is a crucial indicator reflecting the grate cooler's condition during operation, effectively revealing its operational status, this embodiment extracts the trend, slope, and fluctuation characteristics of the grate pressure data from typical operating condition data. These are used as typical features of the typical operating condition data for comprehensive analysis of the grate cooler's operating conditions. The specific extraction process of these typical features is explained below:
[0044] (4.1) The local weighted regression algorithm is used to fit the data and extract trend features from the grate pressure data in typical working conditions. The trend can reveal the long-term direction of data change, help identify the rising, falling or stable state of the grate cooler, eliminate noise interference, and highlight the main trend in the operation of the grate cooler.
[0045] The Locally Weighted Regression (LOWESS) algorithm is used to extract trend features from grate data. It performs regression fitting around each data point using weighted least squares, focusing on the changes of the target point and its neighboring data.
[0046] In the locally weighted regression algorithm, a weight function w(x) is first defined. i ,x j The `tricube` function is typically used, as shown in the following formula:
[0047] When |x j-x i |≤h
[0048] Where: h is the window size, used to control the influence range of neighboring data points; x i Indicates the current test point; x j This represents the training sample points.
[0049] Then, for each data point xi, the local data is fitted using weighted least squares (WLS) to minimize the weighted sum of squared residuals, yielding the slope β1 of the local trend. Slope β1 represents the local trend centered at the current test point xi, obtained through the weight function w(x... i ,x j The rate of change of training data within the local neighborhood defined by is shown in the following formula:
[0050]
[0051] in: y is the predicted value for the test sample xi; β is the parameter vector (coefficients) of the local linear model, which needs to be solved through optimization; w(xi,x) j ) is the weight function, representing the training sample x. j The test weights for test point xi; n is the weight of the training samples.
[0052] This locally weighted regression algorithm extracts trend features that can capture local nonlinear changes in grate compression data, providing strong support for operating condition identification and equipment status monitoring.
[0053] (4.2) The slope characteristics of the grate pressure data in typical operating conditions are calculated using the sliding window method. The sliding window method can reflect the changes in the operating rate of the grate cooler by calculating the rate of change of the grate pressure data in each window, which is of great significance for monitoring rapid fluctuations and rates of change.
[0054] (4.3) The fluctuation characteristics are obtained by calculating the standard deviation and variance of the grate pressure data. The fluctuation characteristics can be used to describe the fluctuation intensity of the grate pressure data and can effectively identify whether the grate cooler is in a stable operating state.
[0055] In this embodiment, by extracting the trend characteristics, slope characteristics, and fluctuation characteristics of the grate pressure data, as typical features of typical working condition data, rich information can be provided for subsequent cluster analysis, thereby dividing the operating status of the grate cooler into different working condition modes, helping to monitor the operating status of the equipment in real time, and discover potential faults or abnormal fluctuations in advance.
[0056] Furthermore, these extracted typical features can provide strong support for the maintenance and optimization of grate coolers, ensuring that the grate coolers operate in optimal condition and avoiding prolonged periods of unstable operation, thereby improving the efficiency and reliability of the grate coolers. Therefore, the typical feature extraction in this embodiment provides a reliable technical means for the intelligent monitoring and optimized control of grate coolers, promoting the precision of grate cooler management and the intelligent development of industrial production.
[0057] Step 5: Use a fusion model consisting of K-means clustering algorithm and Gaussian mixture model to cluster the typical working condition data obtained in Step 3. During clustering, first use K-means clustering algorithm to cluster the typical working condition data to obtain coarse working condition clustering results, and then use Gaussian mixture model to further optimize the coarse working condition clustering results obtained by K-means clustering algorithm, thereby obtaining the final working condition clustering results.
[0058] K-means clustering is a distance-based unsupervised learning algorithm that aims to divide a dataset into K clusters, maximizing the similarity between data points within a cluster and minimizing the similarity between clusters. The basic idea of the K-means algorithm is to iteratively optimize the dataset to eventually obtain the centroid of each cluster. Gaussian Mixture Model (GMM) is a probabilistic generative model used to represent a weighted combination of multiple Gaussian distributions. It assumes that data points come from multiple Gaussian distributions, each representing a cluster. Compared to K-means clustering, GMM does not simply assign data points to a specific cluster; instead, it assigns each data point a probability of belonging to a different cluster, making it suitable for handling data with complex structures.
[0059] In this embodiment, the K-means clustering algorithm is used to initially classify the typical operating condition data obtained in step 3. The K-means algorithm divides the data into several clusters by minimizing the Euclidean distance from each sample to the cluster center. At this stage, K-means performs coarse clustering of the data, classifying the grate cooler's operating conditions into several approximate categories (e.g., high grate pressure, medium grate pressure, and low grate pressure operating conditions). This step can quickly provide a preliminary classification of operating conditions.
[0060] In this embodiment, a Gaussian mixture model is used to further refine the clustering. Since K-means clustering assumes spherical clusters with no overlap, but in actual industrial data, the boundaries between operating conditions may be ambiguous or non-spherical, this embodiment uses a Gaussian mixture model to fit a Gaussian distribution to each cluster, thereby further optimizing the clustering results based on the probability distribution of data points. The Gaussian mixture model uses the expectation-maximization (EM) algorithm to optimize the distribution of data within clusters, thus improving the accuracy and robustness of clustering. In this way, the Gaussian mixture model can handle overlapping regions and complex shapes between clusters, thereby further refining the division of operating conditions, especially when the boundaries between operating conditions are ambiguous or have significant overlap.
[0061] In this embodiment, a fusion model of K-means and GMM is adopted, which can improve the accuracy and stability of clustering results while ensuring efficient computation. The fast clustering results of K-means provide good initial conditions for GMM, while GMM can effectively avoid the misclassification problem that may be caused by hard clustering through probability allocation mechanism, further improving the accuracy of working condition classification. In addition, this strategy can adapt to different working condition structures, especially dynamically changing working condition patterns, enabling the clustering process to accurately capture the evolution characteristics of working conditions, providing strong support for subsequent working condition identification and optimized control.
[0062] The K-means+GMM fusion model in this embodiment performs the following clustering process on the typical working condition data obtained in step 3:
[0063] (5.1) Initialize K-means cluster centers
[0064] Based on the opinions of experts on site, the operating conditions of the grate cooler are divided into three categories: high grate pressure, medium grate pressure, and low grate pressure. Three initial cluster centers are selected, and three sample points are randomly selected from the typical operating condition data obtained in step 3 as initial centroids.
[0065] (5.2) Calculate the distance between the sample point and the cluster center.
[0066] For each data point x in the typical working condition data obtained in step 3 i Use Euclidean distance to calculate data point x i With each cluster center c k The distance between the sample point and the cluster center is used to measure the similarity between the sample point and the cluster center. The calculation formula is:
[0067]
[0068] in, Let x represent the i-th sample point. im Represents the i-th sample point x i m-dimensional features; It is the centroid of the k-th cluster, c km d(x) represents the m-th eigenvalue of the k-th cluster centroid; i ,c k ) represents the sample point x i With cluster centroid c k The Euclidean distance between them.
[0069] (5.3) Assign the sample points to the nearest cluster
[0070] Based on the calculated Euclidean distance, each sample point x i Assigned to the nearest cluster center c k Cluster label l i Represents sample point x i The cluster it belongs to, cluster label l i The formula is:
[0071]
[0072] This step assigns the sample point to the nearest cluster k, which is called hard clustering.
[0073] (5.4) Update cluster center
[0074] After all sample points have been assigned, the mean of all sample points within each cluster is calculated and used as the new centroid of that cluster. Let c be the new centroid of cluster k. k ′, then the new centroid c k ′ is the mean of all sample points in the cluster, calculated using the following formula:
[0075]
[0076] Among them, C k Let |C| represent the set of sample points in cluster k. k | is the number of sample points in cluster k, x i It is the i-th sample point in cluster k.
[0077] (5.5) Repeat steps (5.2) and (5.3).
[0078] Repeat steps (5.2) and (5.3) until the cluster centers no longer change significantly, or the maximum number of iterations is reached. Each time the cluster centers are updated, the distance from each data point to each cluster center is recalculated, and cluster labels are reassigned. When the algorithm converges, the cluster centers are stable, and the sample points no longer change their cluster affiliation.
[0079] (5.6) Calculate the total squared error (SSE)
[0080] This embodiment calculates the total squared error (SSE) to evaluate the clustering effect, which is the sum of the squares of the distances from all data points within a cluster to their cluster centers. The formula for calculating SSE is as follows:
[0081]
[0082] Where K is the number of clusters, C k c is the set of sample points in cluster k. k It is the centroid of cluster k.
[0083] (5.7) Add the K_means clustering results back to the typical working condition data obtained in step 3 to prepare for GMM clustering.
[0084] (5.8) Initialize the parameters of the Gaussian Mixture Model (GMM), including:
[0085] mean (μ) k ): The mean vector of each Gaussian distribution.
[0086] Covariance matrix (Σ) k ): The covariance matrix of each Gaussian distribution, used to describe the shape of the cluster.
[0087] Weight (π) k ): The mixture weight of each Gaussian distribution, representing the proportion of that Gaussian distribution in the entire data.
[0088] (5.9) E-step (desired step)
[0089] Calculate the probability that each data point belongs to each cluster. Given a data point xi, calculate its posterior probability (responsibility) of belonging to cluster k, representing the responsibility γ of the data point to cluster k. ik The calculation formula is as follows:
[0090]
[0091] Where: N(x) i |μ k ,Σ k ) is the probability density function of a Gaussian distribution, representing the probability density function of a data point x. i The probability in cluster k, and we have:
[0092]
[0093] Where: m is the dimension of the data with added features of slope variance and k-means label, Σ k It is the covariance matrix of cluster k, μ k It is the mean of cluster k.
[0094] (5.10) M-steps (maximizing steps)
[0095] Responsibility level γ calculated using E-step ik Update the parameters of the GMM, including the mean, covariance, and weights.
[0096] in:
[0097] mean μ k The updated formula is as follows:
[0098]
[0099] Covariance matrix Σ k The updated formula is as follows:
[0100]
[0101] weight π k The updated formula is as follows:
[0102]
[0103] Where N is the number of sample points, γ ik It is the sample point x i The degree of responsibility belonging to cluster k.
[0104] (5.11) Repeat steps E and M.
[0105] Repeat step E of step (5.9) and step M of step (5.10) until the log-likelihood function converges or the maximum number of iterations is reached. The log-likelihood function L is used to measure the goodness of fit between the current model and the data, and is calculated using the following formula:
[0106]
[0107] The convergence condition is that the change in the log-likelihood function L is less than a certain threshold, or the maximum number of iterations has been reached.
[0108] (5.12) Principal Component Analysis (PCA) is applied to reduce the dimensionality of the standardized data, compressing it from a high-dimensional space to a three-dimensional space. Standardized data refers to data clustered using a K-means + GMM fusion model, where the GMM clustering results replace the K-means clustering results (the K-means clustering results are added for GMM processing; once GMM processing is complete, K-means is no longer needed as a feature and can be removed). In other words, the standardized data consists of typical working condition data with outliers removed, supplemented with trend slope variance extracted from the grate pressure and GMM clustering results.
[0109] PCA extracts the three principal components with the largest variance from the standardized data, preserving the most informative directions. After dimensionality reduction, the coordinates of each data point in three-dimensional space are obtained, and a three-dimensional scatter plot is drawn using visualization tools such as Matplotlib.
[0110] To visually represent the GMM clustering results, data points are colored according to cluster labels, giving different clusters distinct colors to clearly show the distribution of each cluster and thus evaluate the clustering effectiveness. This process facilitates intuitive analysis of the clustering results through visualization, allowing observation of the separation and density between different clusters.
[0111] Step 6: Construct the XGboost model. Use the typical features of the typical working condition data obtained in Step 4 as the input features of the XGboost model, and use the final working condition clustering results obtained in Step 5 as the target labels to train the XGboost model. The trained XGboost model is then used as the working condition recognition model.
[0112] XGBoost (Extreme Gradient Boosting) is an efficient gradient boosting algorithm that performs well in many machine learning tasks and is suitable for identifying and processing clustering scenarios.
[0113] The XGBoost model trains multiple decision trees and optimizes its performance using gradient boosting. During training, the XGBoost model progressively adjusts the tree weights using weighted averaging and gradient information from samples, reducing the model's error at each step to achieve higher prediction accuracy. The advantage of the XGBoost model lies in its ability to effectively avoid overfitting through regularization during training, thus improving the model's generalization ability.
[0114] During training, the XGBoost model optimizes the objective function. By adjusting hyperparameters such as the learning rate, tree depth, and subsampling ratio, it can effectively improve the training efficiency and accuracy of the model. Hyperparameters are typically tuned using grid search or random search to ensure optimal model performance on the training set while avoiding overfitting on the test set.
[0115] In this embodiment, the typical features of the typical operating condition data obtained in step 4 are used as the input features of the XGboost model, and the final operating condition clustering result obtained by the fusion model composed of K-means and GMM in step 5 is used as the target label. The feature vector of each data sample includes the operating parameters and operating condition characteristics of the grate cooler, while the label represents different operating condition categories. 80% of the data is used for training, and the remaining 20% is used for testing to evaluate the generalization ability of the model.
[0116] The core idea of the XGBoost model is to improve model performance by continuously optimizing the residuals. During training, the XGBoost model starts with an initial prediction, typically the mean of the training data. The output F0(x) of this initial model can be expressed as:
[0117]
[0118] in: is the mean of the training data, used as the initial prediction value; x is the feature vector.
[0119] Next, the XGBoost model progressively optimizes itself by minimizing a loss function. This loss function consists of two parts: one is the training error, representing the difference between the current prediction and the true label; the other is a regularization term, used to control the model's complexity and prevent overfitting. The specific loss function L(F) can be expressed as follows:
[0120]
[0121] Where: L(y) i ,F(x i ) represents the error of the current model; Ω(F) is the regularization term, which helps control the complexity of the model; y i Represents the true label; F(x) i ) represents the model's predicted value.
[0122] The regularization term Ω(F) is typically the sum of the tree's complexity and the squares of the leaf node weights, expressed as follows:
[0123]
[0124] Where: T is the number of trees, γ and λ are regularization parameters that control the complexity of the trees; w j This represents the leaf weight.
[0125] During training, the XGBoost model continuously updates and optimizes the residuals in each iteration using gradient boosting. In each iteration, the XGBoost model calculates the current prediction error (i.e., the residuals) and improves the model by learning from these residuals. The residuals calculated in each iteration can be obtained by applying gradient descent to the loss function, as shown in the following equation:
[0126]
[0127] Where, r i (t)F represents the residual of the t-th iteration; (t-1) (x i ) represents the cumulative prediction value of the model for all samples after the (t-1)th iteration.
[0128] Finally, the XGBoost model uses a decision tree to fit these residuals, thus obtaining the prediction output of the current decision tree, as shown in the following equation:
[0129]
[0130] Where: η is the learning rate. This is the output of the current decision tree; This represents the cumulative predicted value of the model after the t-th iteration.
[0131] In this way, the XGBoost model continuously corrects previous prediction errors in each iteration, making the final trained model more accurate.
[0132] In each iteration, XGBoost updates the model by minimizing the loss function, ultimately resulting in a powerful model that combines multiple decision trees. Each iteration allows the model to achieve a better fit on the training data, gradually approaching the optimal classification result.
[0133] After multiple iterations, the XGBoost model generates the final prediction model. For samples in the training set, the model makes predictions based on the output of each tree, and the final prediction is a weighted sum of the outputs of all trees. For new samples, the XGBoost model makes predictions in the same way, combining the results of all trees to obtain the final output.
[0134]
[0135] After training, the trained XGBoost model is used as the working condition identification model. This model is used to predict unknown data and classify new data based on learned patterns.
[0136] Step 7: Input the operating condition data of the grate cooler to be identified into the operating condition identification model obtained in Step 6, and the operating condition identification model outputs the operating condition identification result.
[0137] In the following embodiment, the clustering effect of the fusion model composed of K-means and GMM, as well as the performance of the final working condition identification model, are analyzed and explained.
[0138] (1) The clustering effect of the fusion model was evaluated using the silhouette coefficient, Davies-Bouldin index and Calinski-Harabasz index.
[0139] The silhouette coefficient S(i) is used to measure the quality of the clustering results, and it is defined as follows:
[0140]
[0141] Where: a(i) is the average distance between sample i and other samples in the same cluster, and b(i) is the minimum average distance between sample i and samples in its nearest neighbor cluster. The silhouette coefficient ranges from -1 to 1. A value closer to 1 indicates a better clustering effect, a value closer to 0 indicates a more ambiguous clustering, and a negative value indicates that the sample may have been incorrectly assigned to a cluster.
[0142] The Davies-Bouldin index (DBI) is used to measure the separation between clusters and the compactness within clusters. It is defined as follows:
[0143]
[0144] Where Si represents the average distance between samples within cluster i, d(ci,cj) represents the distance between the cluster centers of cluster i and cluster j, and N is the total number of clusters. A smaller DBI value indicates a better clustering effect.
[0145] The Calinski-Harabasz index (CH) is used to assess the tightness and separation of clusters. The Calinski-Harabasz index (CH) is defined as follows:
[0146]
[0147] Where: Tr(Bk) is the trace of the inter-cluster scatter matrix, Tr(Wk) is the trace of the intra-cluster scatter matrix, k is the number of clusters, and n is the total number of samples. A larger CH value indicates a better clustering effect.
[0148] In traditional unsupervised clustering methods, K-means has become the preferred choice for industrial big data analysis due to its low computational complexity and simple implementation. This algorithm iteratively minimizes the Euclidean distance from samples to the centroids, enabling hard partitioning of large-scale samples in O(NK) time, demonstrating significant real-time performance in engineering. However, K-means inherently assumes that clusters are equivariant spheres and that there is no overlap between clusters. When the data is elongated, ellipsoidal, or has significant density differences, this strong assumption can lead to over-pruning of cluster boundaries or mis-clustering. Furthermore, the algorithm is extremely sensitive to the initial centroids and the preset number of clusters K, and lacks robustness to outliers.
[0149] Gaussian Mixture Models (GMMs), within a probabilistic generative framework, treat each cluster as a parameter-variable Gaussian distribution. They can characterize non-spherical clusters and assign soft labels to samples using mean vectors and covariance matrices, thus naturally handling overlaps and transition states between clusters. However, GMMs also have significant drawbacks: parameter estimation relies on EM iterations, which are prone to getting trapped in local optima and converge slowly when high-quality initial values are lacking; and the repeated inversion of high-dimensional covariance matrices significantly increases computational overhead.
[0150] To address the aforementioned limitations, this embodiment proposes a K-means+GMM fusion strategy. First, K-means is used to rapidly and coarsely partition the samples, obtaining stable centroids and cluster labels. Then, this label information is used as a warm-start input to the GMM, performing EM fine-tuning only on the mean and covariance within local neighborhoods. This two-stage process combines the advantages of a "hard-first, soft-later" approach: First, the centroids provided by K-means significantly reduce the sensitivity of the GMM to initial values, accelerating convergence and avoiding inferior solutions; second, the probabilistic modeling flexibility of the GMM effectively compensates for the strong sphericity assumption of K-means, allowing cluster boundaries to conform to the non-spherical and overlapping characteristics of the real data; third, the soft allocation of outliers by the fusion model improves overall robustness. As shown in Table 1:
[0151] Table 1 Comparison of Clustering Effects in Experiments
[0152]
[0153] Experimental comparisons show that the three algorithms exhibit a pattern of increasing performance and improving efficiency in their internal performance evaluation metrics. K-means has a higher silhouette coefficient and CH index than GMM, while its DBI index is lower, indicating that under the hard partitioning assumption, K-means can maintain a certain level of inter-cluster separation and intra-cluster compactness. However, limited by the spherical cluster assumption, it still falls short for data with complex shapes or overlapping boundaries. Although GMM introduces covariance to characterize non-spherical distributions, its initialization randomness and high-dimensional covariance estimation errors result in a lower silhouette coefficient and CH index, and a higher DBI, leading to a relatively loose overall clustering structure. In contrast, the model fusion algorithm achieves the best values in all three metrics: a higher silhouette coefficient indicates its ability to simultaneously compress intra-cluster errors and increase inter-cluster distances; a lower DBI reflects smoother interfaces and minimal overlap between clusters; and a higher CH index further confirms its outstanding performance in comprehensively balancing intra-cluster and inter-cluster variances.
[0154] This demonstrates that the rapid coarse segmentation provided by K-means and the probabilistic flexibility of GMM complement each other within the fusion framework. This effectively alleviates the morphological assumption bias and local optima problems inherent in single algorithms, resulting in a final clustering result that outperforms both benchmark methods in terms of robustness, interpretability, and characterization of complex operating conditions. This result validates the applicability and advantages of the model fusion strategy in multi-condition identification of grate coolers and lays a reliable data foundation for subsequent identification models and predictive control.
[0155] (2) The performance of the working condition identification model is evaluated using common classification metrics such as accuracy, precision, recall and F1 score.
[0156] Accuracy measures the proportion of samples correctly classified by the model out of the total samples. The formula is:
[0157]
[0158] Where TP is the number of samples correctly classified as positive, TN is the number of samples correctly classified as negative, FP is the number of samples misclassified as positive, and FN is the number of samples misclassified as negative.
[0159] Precision measures the proportion of samples that are actually predicted to be positive, and its formula is:
[0160]
[0161] High accuracy means that the model accurately predicts the positive class, reducing false positives.
[0162] Recall measures the proportion of samples that are correctly predicted as positive out of all samples that are actually positive. The formula is:
[0163]
[0164] High recall means that the model can capture as many positive samples as possible, reducing false negatives.
[0165] The F1 score (F1-Score) is the harmonic mean of precision and recall, taking into account both the model's accuracy and sensitivity. Its formula is:
[0166]
[0167] A high F1 score indicates a good balance between precision and recall, and is especially suitable for class imbalance situations, avoiding the bias of evaluating the model solely based on accuracy.
[0168] In this embodiment, the XGBoost model used to form the working condition recognition model is an ensemble learning algorithm developed based on the gradient boosting tree framework. Compared with traditional decision trees or ordinary GBDT, XGBoost shows outstanding performance in both efficiency and generalization ability: it utilizes second-order gradient information to accelerate loss optimization, introduces L1 / L2 regularization terms to suppress overfitting, and improves model robustness through mechanisms such as column sampling, learning rate reduction, and automatic handling of missing values; combined with block data storage and parallel splitting strategies, the training speed can be linearly scaled to large-scale features and samples. As shown in Table 2:
[0169] Table 2 Comparison of Model Performance in Experiments
[0170]
[0171] The experimental results are shown in the table. All three operating condition recognition models achieved high scores on the four evaluation metrics—accuracy, precision, recall, and F1 score—but significant performance differences remained. Compared to Extreme Learning Machine (ELM) and Support Vector Machine (SVM), XGBoost achieved the highest scores across all four metrics, demonstrating the most outstanding overall performance and fully proving its generalization ability and robustness in the grate cooler operating condition recognition task. Therefore, this embodiment ultimately selected XGBoost as the primary recognition model for the subsequent deployment of control strategies.
[0172] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. These embodiments are merely descriptions of preferred embodiments and are not intended to limit the scope or concept of the invention. The specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. Such combinations, as long as they do not violate the spirit of the present invention, should also be considered as part of this disclosure. To avoid unnecessary repetition, the present invention will not further describe the various possible combinations.
[0173] This invention is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this invention and without departing from the design idea of this invention, all modifications and improvements made by those skilled in the art to the technical solutions of this invention should fall within the protection scope of this invention. The technical content for which protection is sought in this invention has been fully described in the claims.
Claims
1. A method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means, characterized in that, Includes the following steps: Step 1: Collect the operating condition data of the grate cooler at the production site, and select the key variable operating condition data; Step 2: Preprocess the selected key variable operating condition data to filter out outlier values at shutdown points; Step 3: Using a Gaussian mixture model, the key variable operating condition data after filtering out outliers in Step 2 are divided into two categories: typical operating condition data and abnormal operating condition data. Step 4: Extract typical features from the typical operating condition data obtained in Step 3; Step 5: Use a fusion model consisting of K-means clustering algorithm and Gaussian mixture model to cluster the typical working condition data obtained in Step 3. During clustering, first use K-means clustering algorithm to cluster the typical working condition data to obtain coarse working condition clustering results, and then use Gaussian mixture model to further optimize the coarse working condition clustering results obtained by K-means clustering algorithm, thereby obtaining the final working condition clustering results. Step 6: Construct the XGboost model. Use the key features of the typical working condition data obtained in Step 4 as the input features of the XGboost model, and use the final working condition clustering results obtained in Step 5 as the target label to train the XGboost model. The trained XGboost model is then used as the working condition recognition model. Step 7: Input the operating condition data of the grate cooler to be identified into the operating condition identification model obtained in Step 6, and the operating condition identification model outputs the operating condition identification result.
2. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 1, characterized in that, In step 1, the key variables include grate pressure, grate speed, kiln main unit current, raw material feed rate, and tertiary air temperature.
3. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 1, characterized in that, In step 2, the Z-score algorithm is used to preprocess the selected key variable operating condition data to filter out outlier values at shutdown points.
4. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 1, characterized in that, In step 4, the trend characteristics, slope characteristics, and fluctuation characteristics of the grate pressure data obtained in step 3 are extracted as the typical characteristics.
5. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 4, characterized in that, A local weighted regression algorithm is used to extract trend features from the grate compression data.
6. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 4, characterized in that, The sliding window method was used to extract slope features from the grate compression data.
7. The method for identifying the operating conditions of a grate cooler in cement production based on the fusion of Gaussian mixture model and K-means as described in claim 4, characterized in that, The fluctuation characteristics are obtained by calculating the standard deviation and variance of the grate compression data.