A Machine Learning-Based Classification and Backfilling Method and System for Spontaneous Combustion Tendency of Coal Gangue
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-08-14
AI Technical Summary
由于煤和煤矸石之间的硫含量、吸氧能力和交叉温度存在差异,因此用于煤的方法不适用于煤矸石
[0028]本发明中,对训练集采用合成少数类过采样算法SMOTE进行数据增强,对测试集采用基于原始数据分布的合成样本扩充方法进行数据增强。对增强后的数据从样本数量、特征分布、数据离散程度及特征相关性四个维度开展对比分析。结果表明,数据增强后其特征分布等与原始数据高度契合,可用于后续分类模型的训练与测试。
Smart Images

Figure CN122571357A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and resources and environment technology, and in particular to a method and system for classifying and backfilling coal gangue based on machine learning-based spontaneous combustion tendency. Background Technology
[0002] Coal gangue is a solid waste generated during coal mining and washing. It is a heterogeneous mixture with low calorific value and complex composition, containing small amounts of combustible components such as carbon and pyrite. Its emissions account for approximately 15%–30% of raw coal production, but its overall resource utilization rate is only about 30%. It is one of the largest and most harmful industrial wastes in my country in terms of both emissions and cumulative volume. According to incomplete statistics, my country currently has over 6 billion tons of coal gangue stockpiled, forming 1600-1800 gangue mountains. These mountains occupy and pollute large amounts of land resources. Untreated, spontaneous dumping, coupled with long-term heat accumulation, has caused over 300 of these mountains to remain in a state of chronic spontaneous combustion. Each square meter of burning area from a continuously burning gangue mountain will emit 10.8g of CO, approximately 2g of H2S, and NO. X Approximately 6.5g of SO2, along with polycyclic aromatic hydrocarbons (PAHs) such as benzo[a]pyrene and benzo[a]fluoranthene, as well as particulate matter, pollute the atmosphere and endanger health. More than 50 explosions have occurred, including nearly 10 major accidents. Furthermore, landslides, mudslides, collapses, gas poisoning, and burials are prone to occur, causing casualties, economic losses, threatening residents' lives, and increasing social instability. Therefore, the treatment and disposal of coal gangue in my country is urgently needed.
[0003] The comprehensive utilization of coal gangue mainly includes power generation, building material preparation, recycling of valuable minerals, production of chemical products, soil improvement and fertilizer production. It can also be used in construction, backfilling of low-lying areas and mine subsidence areas, reclamation of subsidence areas, and road construction. Against the backdrop of the country's strong advocacy for the resource utilization of bulk solid waste and green development, the comprehensive utilization of coal gangue is ushering in a favorable development opportunity and broad market space. According to the relevant requirements of the management regulations for the comprehensive utilization of coal gangue, its utilization should follow the principle of giving equal importance to reduction and resource recovery, adhering to the principles of local utilization, classified utilization, large-scale utilization, and high-value utilization, and achieving a synergistic unity of economic, social, and environmental benefits through technological innovation. The high-intensity, large-scale mining in my country's main coal-producing areas has caused serious ecological damage, resulting in problems such as surface subsidence and groundwater depletion. This has led to the continuous deterioration of the already fragile and sensitive ecological environment of the region, with significant degradation of grasslands and vegetation, further exacerbating the vicious cycle between resources and ecology, and aggravating soil erosion in wind-blown sand areas, desertified areas, and the Loess Plateau. Furthermore, ecological disturbances such as mining waste discharge, heavy metal and water pollution can easily spread from the mining area to surrounding areas through multiple ecological cycles, leading to regional and even city-wide ecological and environmental problems. Therefore, using coal gangue backfill for ecological restoration of wind and water erosion areas has significant practical implications.
[0004] In the process of backfilling wind-eroded areas with coal gangue, its flammability is particularly important for the design of backfilling schemes. To date, considerable research has been conducted on the mechanism, influencing factors, and prevention of spontaneous combustion of coal gangue. It is generally believed that spontaneous combustion of coal gangue is the result of the combined action of pyrite and the organic components of coal. On the one hand, coal gangue contains combustible components similar to those in coal. These combustible components adsorb oxygen from the air and react with it, generating a large amount of heat. On the other hand, pyrite in coal gangue reacts more readily with oxygen in humid environments, producing ferrous sulfate and ferric hydroxide, which continue to undergo exothermic redox reactions in aqueous solutions. If left uncontrolled, both factors can promote spontaneous combustion, as the heat accumulated from these reactions increases the internal temperature of the gangue. These studies have played an important role in advancing our understanding of the combustion behavior and thermodynamics of coal gangue. Traditional coal gangue treatment methods rely on a large amount of experimental data, which is not only time-consuming and labor-intensive but also involves high economic costs. Furthermore, these traditional methods have many shortcomings, including low resource utilization efficiency, complex and costly processing procedures, insufficient control over environmental impact, lack of intelligent decision support, and difficulty in adapting to large-scale and customized needs. These limitations often lead to low utilization rates of coal gangue, high environmental pollution risks, and poor processing efficiency. Moreover, to date, no effective method has been established for evaluating the spontaneous combustion tendency of coal gangue. Establishing a rigorous method for assessing the spontaneous combustion hazard of coal gangue will be beneficial for its prevention and control. Because of the differences in sulfur content, oxygen absorption capacity, and cross-temperature between coal and coal gangue, methods used for coal are not applicable to coal gangue. Therefore, developing a quantifiable and universally applicable method to describe the spontaneous combustion tendency of coal gangue is crucial.
[0005] To address this gap, this invention establishes a quantitative assessment stacking model for the spontaneous combustion tendency of coal gangue. This model can predict the intrinsic and fundamental spontaneous combustion category (not easily spontaneously combustible, easily spontaneously combustible, extremely easily spontaneously combustible) of coal gangue based solely on its material composition, excluding interference from external environmental factors (temperature, oxygen concentration, particle size, etc.). Subsequently, a corresponding backfilling plan is generated based on the corresponding spontaneous combustion category of the coal gangue. The backfilling plan generally focuses on the thickness, particle size, compaction coefficient, and soil thickness of the coal gangue. Summary of the Invention
[0006] This invention addresses the problems and shortcomings of existing technologies by providing a machine learning-based method and system for classifying and backfilling the spontaneous combustion tendency of coal gangue.
[0007] The present invention solves the above-mentioned technical problems through the following technical solution:
[0008] This invention provides a machine learning-based method for classifying and backfilling the spontaneous combustion tendency of coal gangue, characterized by comprising the following steps:
[0009] S1. Data Collection: Obtain a dataset with category labels for the spontaneous combustion tendency level of coal gangue based on the K-means clustering algorithm, and divide the dataset into training set and test set. The data input consists of 9 cleaned and standardized features, and the output consists of category labels. The category labels include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content, sulfur content, moisture content, ash content, volatile matter, alumina, silicon dioxide, calcium oxide, and sulfur trioxide.
[0010] S2. Data Augmentation: The training set is augmented using a synthetic minority class oversampling algorithm. The number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible in the training set is expanded to a uniform target number. The consistency of the training set distribution before and after augmentation is verified and analyzed from three dimensions: single feature distribution, distribution differences between categories, and multi-feature association patterns. The analysis shows that the distribution is consistent. The test set is augmented using a synthetic sample expansion method based on metadata distribution. During the augmentation and expansion process, the distribution ratio of metadata categories in the test set remains unchanged. The consistency of the test set distribution before and after augmentation is verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization deviation. The analysis shows that the distribution is consistent.
[0011] S3. Model Construction: Five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—were selected as base models to construct a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set was input into each base model to complete the fitting and parameter tuning. Subsequently, the prediction output classification results were performed on the preprocessed augmentation test set. The classification results included class labels and probability values of each class, obtaining the optimized base models after tuning. The Stacking ensemble learning framework was used to construct a Stacking two-layer ensemble learning model. The five optimized base models were used as the first layer learners, and their prediction results were used as new fusion features. Logistic regression was used as the second layer meta-learner. Layered 5-fold cross-validation was used to complete the training to ensure that each class was evenly distributed in each fold, obtaining the optimized Stacking ensemble model after tuning.
[0012] S4. Optimal Model Selection: Perform SHAP interpretability analysis on all base models to quantify the positive and negative contributions of each feature to the classification results, identify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance indicators of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model.
[0013] S5. Model Prediction: Construct a classification decision module. Input the nine measured features of the study area into the coal gangue spontaneous combustion feature value input area to predict the spontaneous combustion category of coal gangue. Output the spontaneous combustion tendency level of coal gangue in the study area and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a backfill scheme for coal gangue that is not easy to spontaneously combust, a backfill scheme for coal gangue that is easy to spontaneously combust, and a backfill scheme for coal gangue that is extremely easy to spontaneously combust.
[0014] This invention also provides a machine learning-based system for classifying and backfilling the spontaneous combustion tendency of coal gangue, characterized in that it includes:
[0015] The data collection unit is used to acquire a dataset with category labels of coal gangue spontaneous combustion tendency levels based on the K-means clustering algorithm, and divide the dataset into training set and test set. The data input consists of 9 cleaned and standardized features, and the output is the category labels, which include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content, sulfur content, moisture content, ash content, volatile matter, alumina, silicon dioxide, calcium oxide, and sulfur trioxide.
[0016] The data augmentation unit is used to augment the training set using a synthetic minority class oversampling algorithm, expanding the number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible to a unified sampling target number. The consistency of the training set distribution before and after augmentation is verified and analyzed from three dimensions: single-feature distribution, distribution differences between categories, and multi-feature association patterns. The analysis shows that the distribution is consistent. For the test set, a synthetic sample augmentation method based on metadata distribution is used for data augmentation. During the augmentation process, the distribution ratio of metadata categories in the test set remains unchanged. The consistency of the test set distribution before and after augmentation is verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization deviation. The analysis shows that the distribution is consistent.
[0017] The model building unit selects five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—as base models to build a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set is input into each base model to complete the fitting and achieve parameter tuning. Subsequently, the prediction output classification results are performed on the preprocessed augmentation test set, including class labels and probability values of each class, to obtain the optimized base models. The Stacking ensemble learning framework is used to build a Stacking two-layer ensemble learning model, using the five optimized base models as the first layer learner and their prediction results as new fusion features. Logistic regression is used as the second layer meta-learner, and training is completed using hierarchical 5-fold cross-validation to ensure that each class is evenly distributed in each fold, thus obtaining the optimized Stacking ensemble model.
[0018] The optimal model selection unit is used to perform SHAP interpretability analysis on all base models, quantify the positive and negative contributions of each feature to the classification results, clarify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance index of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model.
[0019] The model prediction unit is used to input nine measured features of the study area into the coal gangue spontaneous combustion feature value input area, predict the spontaneous combustion category of coal gangue, output the spontaneous combustion tendency level of coal gangue in the study area, and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a non-spontaneous combustion coal gangue backfill scheme, a spontaneous combustion coal gangue backfill scheme, and an extremely spontaneous combustion coal gangue backfill scheme.
[0020] This invention also provides a coal gangue spontaneous combustion tendency classification and backfill decision support system, characterized by adopting a B / S architecture + front-end and back-end separation + three-layer logical layering design pattern to achieve decoupling of interface interaction, business logic and data storage. The three-layer logical layering consists of a presentation layer, an application layer and a data layer. The front-end adopts the Vue 3 + Vite + VueRouter + Pinia + Axios + ECharts + Gaode Map technology stack, the back-end adopts the Python 3.11 + FastAPI + Pydantic + SQLAlchemy + Uvicorn technology stack, and the database adopts the SQLite lightweight database, with data persistence achieved through SQLAlchemy ORM.
[0021] The system includes a registration and login module, a backfill point monitoring module, a classification decision module, a history module, and a backfill knowledge module;
[0022] The registration and login module is used for user registration and login, and implements user identity authentication;
[0023] The backfill point monitoring module is used to display the geographical location of each mining area in a certain region on an electronic map, and to specifically display the visual distribution of the study area in a certain mining area. It displays the core statistical data of the mining area in the form of a data panel, and uses a bar chart to visualize the weight distribution of nine characteristics of the spontaneous combustion tendency of coal gangue. Based on the correlation between the internal temperature of coal gangue and the risk of spontaneous combustion, a three-level early warning visualization system is constructed.
[0024] The classification decision module is used to input nine measured features of the study area into the coal gangue spontaneous combustion feature value input area, predict the spontaneous combustion category of coal gangue, output the spontaneous combustion tendency level of coal gangue in the study area, and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a non-spontaneous combustion coal gangue backfill scheme, a spontaneous combustion coal gangue backfill scheme, and an extremely spontaneous combustion coal gangue backfill scheme.
[0025] The history module is used for centralized management of historical prediction results saved during the classification decision-making stage;
[0026] The backfilling knowledge module provides users with a standardized knowledge base and specification support covering the entire backfilling process.
[0027] The positive and progressive effects of this invention are as follows:
[0028] In this invention, the Synthetic Minority Oversampling (SMOTE) algorithm is used to augment the training set, while a synthetic sample augmentation method based on the original data distribution is used to augment the test set. The augmented data is compared and analyzed from four dimensions: sample size, feature distribution, data dispersion, and feature correlation. The results show that the feature distribution and other characteristics of the augmented data are highly consistent with the original data, and it can be used for subsequent training and testing of classification models.
[0029] In this invention, five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—are selected as base models to construct a basic classification model library. Under a unified augmented dataset and experimental settings, the classification performance of each base model is comprehensively evaluated through repeated training and parameter tuning. Based on this, an ensemble model is constructed using the Stacking ensemble learning framework. The five base models serve as the first-layer learners, and their prediction results are used as new fusion features. Logistic regression is selected as the second-layer meta-learner to integrate and optimize the advantages of multiple models. Simultaneously, to overcome the "black box" dilemma of machine learning models, SHAP interpretability analysis is conducted on all base models. Visualization techniques such as feature importance heatmaps, scatter plots, and waterfall plots are used to quantify and analyze the positive and negative contributions of each feature to the classification results, clarifying the core variables affecting the classification of coal gangue spontaneous combustion tendency. Finally, by comparing the classification performance metrics of all models, the optimal Stacking ensemble model was selected. This model not only ensures classification accuracy and generalization ability, but also has good interpretability, providing core algorithmic support for the subsequent construction of a coal gangue spontaneous combustion tendency classification and backfill decision support platform.
[0030] In this invention, during the system construction phase, a coal gangue spontaneous combustion tendency classification and prediction page is developed using the optimized Stacking classification model as the core algorithm. This page supports user input of nine features of coal gangue and outputs real-time classification results such as "not easily spontaneously combustible," "easily spontaneously combustible," and "extremely easily spontaneously combustible." Based on this, a backfilling decision page is developed simultaneously, matching corresponding backfilling process parameters according to the spontaneous combustion classification level, achieving a closed-loop output from spontaneous combustion risk identification to backfilling disposal plan. During system development, a front-end and back-end separation architecture is adopted to ensure user-friendly interface interaction and algorithm call stability. Each functional module undergoes multiple rounds of testing and iterative optimization to ensure the system can efficiently support the engineering application requirements for coal gangue spontaneous combustion category prediction and backfilling plan output.
[0031] This system has developed five core functional modules: registration and login, backfill point monitoring, classification and decision-making, historical records, and backfill knowledge. It has achieved full business process coverage, including user identity authentication, visualization of the study area, intelligent prediction of spontaneous combustion tendency and matching of backfill schemes, historical record management and knowledge base query, providing complete functional support for the classification of spontaneous combustion tendency of coal gangue and backfill decision-making.
[0032] This system enables intelligent and visualized classification of coal gangue spontaneous combustion tendency and backfilling decisions, providing coal mining enterprises with an efficient and reliable decision support tool. It has significant engineering application value for improving the resource utilization level of coal gangue and reducing the safety risks of spontaneous combustion. Attached Figure Description
[0033] Figure 1 This is a technical roadmap for the present invention.
[0034] Figure 2 A comparison chart showing the number of samples before and after augmentation of the training set.
[0035] Figure 3 A comparison chart of kernel density between the original data of the nine features in the training set and the newly added data in SMOTE.
[0036] Figure 4 Violin plot of the original training data and the feature distribution after SMOTE enhancement.
[0037] Figure 5 Scatter plots of the original and new SMOTE data for pairwise combinations of the nine features in the training set (This figure is a matrix of 81 SMOTE scatter plots generated by pairwise combinations of the nine features in the training set. Only the subplots of the three feature combinations involved in the analysis are magnified and annotated locally to clearly show the details of the data distribution; circles represent the original data, crosses represent the new data, blue represents non-flammable, orange represents flammable, and green represents extremely flammable).
[0038] Figure 6 A comparison chart showing the number of samples before and after the test set expansion.
[0039] Figure 7 A comparison of the raw data of the nine features in the test set with the kernel density (KDE) of the synthetic samples.
[0040] Figure 8 A violin plot showing the distribution of features between the original test set data and the synthetic samples.
[0041] Figure 9 A heatmap showing the relative deviations of the mean and standard deviation before and after test set expansion.
[0042] Figure 10 This is a diagram of the overall system architecture.
[0043] Figure 11 This is the monitoring interface for backfill points.
[0044] Figure 12 This is the interface for classification and decision-making.
[0045] Figure 13 This is the history view interface. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] like Figure 1 As shown, this embodiment of the invention provides a method for classifying and backfilling the spontaneous combustion tendency of coal gangue based on machine learning, which includes the following steps:
[0048] S1. Data Collection: Obtain a dataset with category labels for the spontaneous combustion tendency level of coal gangue based on the K-means clustering algorithm (see Table 1). Divide the dataset into training set and test set. The data input consists of 9 cleaned and standardized features, and the output consists of category labels. The category labels include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content (FC), sulfur content (S), moisture content (M), ash content (Ad), volatile matter (Vad), alumina (Al2O3), silicon dioxide (SiO2), calcium oxide (CaO), and sulfur trioxide (SO3).
[0049] Table 1. Sample datasets with cluster labels.
[0050]
[0051] S2. Data Augmentation: The training set was augmented using the Synthetic Minority Oversampling (SMOTE) algorithm, expanding the number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible to a uniform target number. The consistency of the training set distribution before and after augmentation was verified and analyzed from three dimensions: single-feature distribution, inter-class distribution differences, and multi-feature association patterns, demonstrating distribution consistency. The test set was augmented using a synthetic sample expansion method based on metadata distribution, maintaining the metadata class distribution ratio in the test set during augmentation. The consistency of the test set distribution before and after augmentation was verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization bias, demonstrating distribution consistency. The augmented training and test sets were stored in Excel format.
[0052] The coal gangue dataset used in this invention contains only 85 samples, a typical example of a small dataset. After dividing it into training and testing sets in an 8:2 ratio, the training set has only 68 samples and the testing set only 17. This insufficient sample size leads to inadequate model training and poor generalization ability. Furthermore, the original training set exhibits an uneven distribution of samples across the three spontaneous combustion tendencies, resulting in class imbalance, which directly impacts the learning performance and prediction accuracy of the classification model. To address this, targeted data augmentation strategies are implemented for both the training and testing sets: the training set, responsible for model learning, uses SMOTE oversampling to interpolate and generate samples across different classes, thereby balancing the class distribution and improving the model's learning ability across different classes; the testing set is used to objectively evaluate the model's generalization performance. Changing its original data distribution would distort the evaluation results; therefore, sample expansion is performed only while maintaining the original data distribution pattern, without adjusting the class ratio. Through differentiated augmentation, the number of samples is increased while ensuring data authenticity and the rationality of model training and evaluation.
[0053] To fundamentally address the class imbalance problem and ensure balanced model training, this invention sets a unified sampling target, expanding the number of samples in each of the three categories—not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible—to 400. This target eliminates model bias caused by differences in the number of samples in each category in the original training set, while avoiding data redundancy and model overfitting due to oversampling, achieving a balance between sample size and data accuracy. After SMOTE enhancement, the total number of samples in the training set expands from the original 68 to 1200, achieving a completely balanced distribution of the three spontaneous combustion tendency samples. This effectively alleviates the training defects caused by small samples and class imbalance. Simultaneously, the synthesized samples follow the characteristic distribution patterns of the original data, ensuring the effectiveness and reliability of the enhanced dataset and providing balanced and high-quality data support for subsequent classification model training. Figure 2The comparison of the number of samples before and after training set augmentation visually demonstrates the effect of the SMOTE algorithm in correcting the imbalanced distribution of the three classes of samples to a balanced distribution. The original training set had fewer than 30 samples in each of the three classes, which increased to 400 samples after augmentation. This fully verifies the role of data augmentation strategies in improving the problems of small sample size and class imbalance.
[0054] To further verify that the newly generated samples by the SMOTE algorithm do not disrupt the feature distribution patterns of the original data and to ensure the effectiveness and reliability of the augmented dataset, this embodiment analyzes the consistency of the distribution between the original training set and the newly generated SMOTE data from three dimensions: single feature distribution, distribution differences between categories, and multi-feature association patterns.
[0055] Figure 3 The image shows a comparison of the kernel density (KDE) of the original data with nine features in the training set and the newly added data from SMOTE. As can be seen from the image, the peak positions of the kernel density curves of the newly added SMOTE data and the original data for the three types of spontaneous combustion tendency samples highly overlap, and the distribution patterns are basically the same, with only a reasonable extension at the tail. This indicates that the newly added samples completely preserve the characteristic distribution patterns of the original data, without introducing abnormal samples that deviate from the original data distribution, effectively maintaining the inherent statistical characteristics of the data. Taking the Vad feature as an example, the "susceptible to spontaneous combustion" samples in the original data show a unimodal distribution, while the "extremely susceptible to spontaneous combustion" samples show a clear bimodal distribution. After SMOTE enhancement, the kernel density curve of the newly added data completely replicates the bimodal shape of the original data, with the peak positions corresponding to the original data. This further verifies the algorithm's ability to accurately preserve complex distribution features, ensuring that the statistical characteristics of the enhanced data are highly consistent with the original data.
[0056] To further verify that the statistical characteristics of the data after SMOTE enhancement have not shifted, this embodiment uses a violin plot to compare and analyze the distribution structure of the original data and the enhanced data. Figure 4 The image shows a violin plot comparing the original data of the nine features in the training set with the data augmented by SMOTE. Blue represents the original training set, pink represents the dataset augmented by SMOTE, white hollow circles mark the median of each feature category, and the thick black line range corresponds to the interquartile range, representing the distribution interval of the middle 50% of the samples.
[0057] Overall, the three types of spontaneous combustion tendency samples showed a high degree of consistency in their violin shape, median position, and interquartile range before and after enhancement. Taking the Al2O3 feature as an example, the median of the "spontaneously combustible" samples in the original data was approximately 25, and the interquartile range was concentrated in the 21-28 range. After SMOTE enhancement, the positions of the pink white dots and the range of the thick black lines overlapped with the blue ones, indicating that the core statistical features were not changed, and only a slight distribution extension occurred after the sample size increased. Taking the Vad feature as another example, the original distribution of the "spontaneously combustible" samples showed a narrow peak shape and a narrow interquartile range. After enhancement, the distribution shape and median position of the data were precisely aligned with the original data, indicating that the SMOTE algorithm did not destroy the structural differences of the single feature distribution during the sample expansion process.
[0058] In summary, the violin plot results clearly demonstrate that the newly added samples in SMOTE not only achieve balance in quantity but also maintain consistency with the original data in key statistical indicators such as median, interquartile range, and distribution pattern, without any abnormal shifts or distribution distortions. This result statistically validates the authenticity and reliability of the augmented dataset, providing a stable and high-quality data foundation for subsequent model training.
[0059] Figure 5 This is a scatter plot matrix of the original SMOTE data and newly added data, representing pairwise combinations of the nine features in the training set. Solid dots represent the original data, crosses represent newly added SMOTE data, blue represents "not easily spontaneously combustible", orange represents "easily spontaneously combustible", and green represents "extremely easily spontaneously combustible".
[0060] As shown in the figure, the distribution area of the newly added samples is completely nested within the feature space of the original data, with no outliers deviating from the original distribution, consistent with the interpolation principle of the SMOTE algorithm. Furthermore, the newly added data and the original data exhibit consistent correlation patterns under multiple feature combinations. For example, the approximately negative correlation trend between Al2O3 and SiO2, and the approximately positive correlation trend between FC and M, are completely preserved and remain unchanged despite the sample expansion. Taking the combination of CaO and SO3 as an example, the clustering region and linear trend of the green "highly flammable" samples in the original data are accurately continued in the newly added data, fully verifying that the SMOTE algorithm only interpolates within the feature neighborhood of the original samples, effectively maintaining the coupling relationship and overall distribution structure between multiple features, ensuring the authenticity and usability of the enhanced dataset.
[0061] Based on the combined visualization results of the three types, the SMOTE oversampling strategy adopted in this embodiment not only achieves class balance and data augmentation, but also fully preserves the distribution pattern of the original data, providing high-quality and highly reliable dataset support for subsequent classification model training.
[0062] After completing the 8:2 hierarchical partitioning of the training and test sets, this invention addresses the issue of the original test set having only 17 samples and being relatively small in scale. It employs a categorical conditional univariate normal distribution sampling method, specifically a sample synthesis and expansion method based on metadata distribution, to process the test set. This method differs from the SMOTE interpolation strategy used in the training set. Its core logic lies in strictly inheriting the statistical distribution patterns of the metadata, constructing a normal distribution model by fitting the mean and standard deviation of the physicochemical characteristics under each category, rather than introducing new distribution biases through interpolation. This ensures that the test set maintains the distribution ratio of the metadata after expansion.
[0063] To ensure the objectivity of the model's generalization performance evaluation, this invention strictly maintains the original distribution ratio of the test set metadata categories during the expansion process, without making any form of category balancing adjustment. Based on the actual proportions of "not easily ignited," "easily ignited," and "extremely easily ignited" in the metadata test set (not easily ignited: 47.1%, easily ignited: 35.3%, extremely easily ignited: 17.6%), the number of synthetic samples for each category after expansion was determined, ultimately expanding the total number of test set samples from 17 metadata entries to 9933 entries. This quantity setting satisfies the basic requirements for sample size in model evaluation while avoiding distribution distortion caused by blind expansion, achieving a reasonable balance between test set size and data authenticity.
[0064] After being expanded by synthesizing based on metadata distribution, the test set not only significantly increased in the total number of samples, but more importantly, the proportion of samples in each category was completely consistent with that of the metadata test set. Furthermore, the synthesized samples were all generated within the statistical distribution range of the metadata, without any abnormal deviations or data redundancy. Figure 6 The comparison of the number of samples before and after the test set enhancement intuitively demonstrates the effect of the method in significantly increasing the size of the test set while fully preserving the distribution ratio of its metadata categories. The number of samples in the three categories of the metadata test set were 8, 6, and 3, respectively. After the expansion, the sample size of each category was increased proportionally to a stable size, which effectively solved the problem of unstable evaluation results in small sample scenarios and laid a solid data foundation for the objective and reliable evaluation of the generalization performance of the model in the future.
[0065] To verify whether the test set effectively inherited the distribution pattern of the original data under the synthetic sample expansion strategy that preserves the original category distribution and proportion, the analysis was carried out from three dimensions: morphological consistency, statistical consistency, and quantification bias.
[0066] First, the distribution patterns of the original data and the synthetic sample are compared using kernel density maps. For example... Figure 7As shown, the peak positions of the kernel density curves of the three types of spontaneous combustion tendency samples basically coincide with those of the original data, and the distribution patterns are basically consistent, with only reasonable extension at the tail. This indicates that the synthetic samples completely retain the characteristic distribution patterns of the original test set, without introducing abnormal samples that deviate from the original data distribution, and effectively maintaining the inherent statistical characteristics of the data. Taking the S feature as an example, the green "not easily spontaneously combustible" samples in the original data show a unimodal distribution, while the orange "easily spontaneously combustible" samples show a clear bimodal distribution. After the synthetic samples are expanded, the kernel density curves of the newly added data completely replicate the unimodal and bimodal shapes of the original data, and the peak positions basically correspond to those of the original data. This further verifies the accurate preservation capability of this expansion strategy for complex distribution characteristics, ensuring that the statistical characteristics of the expanded test set are highly consistent with the original data.
[0067] To further verify that the statistical characteristics of the data after the test set expansion have not shifted, this embodiment uses a violin plot to compare and analyze the distribution structure of the original data and the synthetic sample. Figure 8 The violin plots compare the original data of the nine features of the test set with those of the synthetic samples. Blue represents the original test set, pink represents the synthetic sample dataset, white hollow circles mark the median of each category feature, and the thick black lines correspond to the interquartile range, representing the distribution interval of the middle 50% of the samples.
[0068] Overall, the violin shape, median position, and interquartile range of the three types of spontaneous combustion tendency samples were largely consistent before and after expansion. Only for some features, there were slight differences in the violin shape, median position, and interquartile range before and after expansion for certain spontaneous combustion tendency categories. Taking the SiO2 feature as an example, the median of the "not easily spontaneously combustible" samples in the original data was about 60, and the interquartile range was concentrated in the 55-65 range. After the synthetic sample expansion, the positions of the white dots in the pink area and the range of the thick black line largely overlapped with the blue area, indicating that the core statistical features were not changed, and only a slight distribution extension occurred after the sample size increased. Taking the Vad feature as another example, the synthetic sample distribution of the "not easily spontaneously combustible" category showed an extremely narrow linear shape, the median had a slight shift from the original data, and the interquartile range had little overlap. This phenomenon was mainly due to the small total number of original samples in the test set, the imbalance of the number of samples in each category, and the relatively concentrated numerical distribution and small variance of each feature in its corresponding category. Under small sample conditions, supplementing a small number of synthetic samples only through nearest-neighbor interpolation can easily cause significant fluctuations in statistics such as the median and quartiles, resulting in significant differences in the violin distribution pattern. Although some features show distribution shifts in specific categories, from the perspective of data augmentation, the synthetic samples are still strictly generated within the numerical range of the original data, do not exceed the value range of the original features, do not introduce out-of-distribution abnormal samples, and overall remain within the reasonable fluctuation range of the original data.
[0069] In summary, the violin plot results clearly demonstrate that the expanded test set not only supplements the sample size but also maintains consistency with the original data in key statistical indicators such as median, interquartile range, and distribution pattern. This result statistically validates the authenticity and reliability of the expanded test set, providing a stable and high-quality data foundation for subsequent model evaluation.
[0070] To quantify the shift in feature distribution before and after test set expansion, this embodiment uses the original test set as a benchmark to calculate the relative deviation of the mean and standard deviation of each feature under different spontaneous combustion tendency categories, and visualizes the deviation level through a heatmap, such as... Figure 9 As shown.
[0071] From the perspective of mean deviation, the relative deviation of the mean for most features is controlled within 10%, with the deviations for stable components such as SiO2, Al2O3, and FC all being <2%, indicating that the mean distribution pattern of the core features is well inherited. High deviations are mainly concentrated in Vad (19.2%) and M (16%) in the non-spontaneously ignitable category, and CaO (13.9%) in the spontaneously ignitable category. The mean deviations of most features in the small sample category "extremely spontaneously ignitable" are significantly lower, indicating that the expansion strategy has a better fit to the mean of the small sample category.
[0072] From the perspective of standard deviation, the relative standard deviations of stable features such as SiO2, Al2O3, and FC are all <7%, reflecting a high degree of consistency in the dispersion of the distribution. High deviations are concentrated in Vad (78.3%) and Aad (75.4%) of the "not easily spontaneously combustible" category, which is highly consistent with the distribution pattern changes observed in the violin plot, confirming that the statistical fluctuations of small sample and low variance features are significant after expansion. The standard deviations of the easily spontaneously combustible and extremely easily spontaneously combustible categories are generally controllable, indicating that the dispersion of the data has not been substantially damaged after expansion.
[0073] In summary, the test set expansion strategy effectively preserved the distribution and statistical regularity of the original data at the overall level, with the mean and standard deviation of stable features remaining at extremely low levels. High biases at the local level were concentrated only in the volatile features of small sample categories, which are inherent statistical fluctuations in small sample data expansion and do not introduce out-of-distribution abnormal samples, remaining within the reasonable fluctuation range of the original data. This strategy effectively improved the data stability of small sample categories while ensuring overall distribution consistency, providing a reliable data foundation for subsequent model evaluation.
[0074] In this step, the SMOTE oversampling algorithm was used to augment the training set data, increasing the number of data entries from 68 to 1200. The reliability of the data augmentation was verified from three dimensions: single feature distribution, inter-category distribution differences, and multi-feature association patterns, using kernel density plots, violin plots, and scatter plot matrices. For the test set, a sample synthesis augmentation method based on metadata distribution was used to augment the test set data, increasing the number of data entries from 17 to 9933. The rationality of the data augmentation was verified from three dimensions: data distribution pattern, statistical regularity, and mean and standard deviation deviation of stable features, using kernel density plots, violin plots, and heatmaps of relative deviations between the mean and standard deviation.
[0075] S3. Model Construction: Five representative algorithms—Random Forest (RF), Extreme Gradient Boosting (XGBoost), CatBoost, Lightweight Gradient Boosting Machine (LightGBM), and Support Vector Machine (SVC)—were selected as base models to construct a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set was input into each base model to complete the fitting and parameter tuning. Subsequently, predictions were performed on the preprocessed augmentation test set to output classification results, including class labels and probability values for each class. Optimized base models were obtained after each tuning. Through repeated training and parameter tuning, the classification performance of each base model was comprehensively evaluated. A Stacking ensemble learning framework was used to construct a Stacking two-layer ensemble learning model. The five optimized base models were used as the first-layer learner, and their prediction results were used as new fusion features. Logistic regression was used as the second-layer meta-learner to integrate and optimize the advantages of multiple models. Layered 5-fold cross-validation was used to complete the training to ensure that each class was evenly distributed in each fold, resulting in an optimized Stacking ensemble model.
[0076] Random forest is built and trained using Python's scikit-learn library. In the data processing stage, it takes 1,200 training data points balanced by SMOTE and 9,933 test data points synthesized based on metadata distribution as input. The target variable is the three-class label of coal gangue spontaneous combustion tendency.
[0077] The system reads the enhanced training and test sets in Excel format, automatically removes redundant columns between them to ensure complete uniformity of feature dimensions, performs adaptation transformation on non-numerical features, performs standardization on all features using StandardScaler to eliminate dimensional differences, and finally uses LabelEncoder to convert the spontaneous combustion tendency level into a digital code to obtain the preprocessed enhanced training and test sets, which meet the input requirements of the multi-classification model.
[0078] Extreme gradient boosting, classification boosting trees, lightweight gradient boosting machines, and support vector machines are similar to random forests.
[0079] To further improve the overall classification performance of the coal gangue spontaneous combustion tendency classification task, this embodiment constructs a Stacking ensemble model. This model is based on a hierarchical framework of multi-model fusion, integrating the feature advantages of different base learners through serial training and secondary learning to achieve more stable and accurate prediction results.
[0080] This Stacking ensemble model employs a two-layer structure, using Random Forest, XGBoost, CatBoost, LightGBM, and SVC as the first-layer base learners, and Logistic Regression as the second-layer meta-learner. It improves classification performance by integrating the decision-making advantages of multiple models. The model training process, data input, and evaluation system remain consistent with each base model. Training and validation are completed using a 1200-piece training set balanced by SMOTE and a 9933-piece test set synthesized based on metadata distribution, under the same feature space and label encoding rules.
[0081] In this step, the Stacking two-layer ensemble learning model employs hierarchical 5-fold cross-validation for training to ensure a balanced distribution of each category within each fold. First, the preprocessed augmented training set is divided into five parts. Four parts are used sequentially to train and optimize the base model, while the remaining part generates prediction probabilities. This process is repeated across all folds, and the resulting matrix is concatenated to obtain base model probability meta-features of equal length to the augmented training set. Then, the category probabilities of the five optimized base models are horizontally integrated to form a high-dimensional meta-feature matrix, which is input into the logistic regression model to train and obtain the optimal fusion weights for each optimized base model. Finally, the optimized base models predict the probabilities of the preprocessed augmented test set, inputting this matrix into the trained meta-model to output the final ensemble classification result and category probabilities.
[0082] All six models were constructed using the same random seed, random_state=42, ensuring that the experimental process was fully reproducible, the results were compared fairly and rigorously, and errors caused by random factors were eliminated.
[0083] S4. Optimal Model Selection: Perform SHAP interpretability analysis on all base models to quantify the positive and negative contributions of each feature to the classification results, identify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance indicators of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model.
[0084] To comprehensively and objectively evaluate the performance of each model in the task of classifying the spontaneous combustion tendency of coal gangue, all models were evaluated using classification performance indicators, including accuracy, precision, recall, F1 score, and AUC score.
[0085] Table 2 shows the comparison results of the six models on the full evaluation metrics of the test set. In terms of overall performance, the Stacking ensemble model achieves comprehensive superiority in all metrics: accuracy 0.9744, precision 0.9750, recall 0.9744, F1 score 0.9743, and AUC 0.9990, significantly outperforming other single-base models.
[0086] Compared to individual models, CATBoost and XGBoost have the closest performance, with accuracies ranging from 0.9701 to 0.9702, making them the second-best performing models. SVC and LightGBM have slightly inferior overall performance, with LightGBM having the lowest AUC (0.9969) among all models. RandomForest is in the lower-middle range across all metrics; although its computational efficiency is high, its classification accuracy is lower than that of boosted tree models and ensemble models.
[0087] Table 2 Comparison of Full Assessment Indicators
[0088]
[0089] More intuitively, it can be observed that the Stacking ensemble model has the highest accuracy, precision, recall, F1 score, and AUC, demonstrating the performance gain of ensemble learning over a single model and verifying the effectiveness of the ensemble strategy in the classification task of spontaneous combustion tendency of coal gangue.
[0090] Perform SHAP interpretability analysis on all base models, including category-level SHAP feature importance, SHAP bee colony graph, interaction graph, waterfall plot, and heatmap.
[0091] SHAP-based interpretability analysis effectively breaks down the "black box" barrier of traditional machine learning models, making the previously difficult-to-interpret model decision-making process open, transparent, and traceable. The base models show high consistency in identifying key features of spontaneous combustion tendency in coal gangue, all using Vad as the core feature, with S, SO3, FC, and M ranking highly in importance, closely matching prior knowledge of the mechanism, indicating reasonable and reliable model decisions. The influence of features on spontaneous combustion tendency exhibits significant nonlinearity and threshold effects, with S and Vad playing a dominant role in the highly spontaneously combustible category. The interaction of FC with other components demonstrates cross-model stability, positively synergistically interacting with sulfur-containing components and negatively antagonistically interacting with inert components. The base models show consistent feature effects under the highly spontaneously combustible category, with differences only in response intensity and amplitude distribution. Overall, the results verify the dominant role of key components in prediction, providing reliable interpretability support for model decision-making and subsequent optimization. This improvement in interpretability has completely changed the "black box" dilemma of model decision-making, which is "invisible and intangible," making the role of features and decision-making logic clear and ensuring the reliability of model decision-making.
[0092] S5. Model Prediction: Construct a classification decision module. Input the nine measured features of the study area into the coal gangue spontaneous combustion feature value input area to predict the spontaneous combustion category of coal gangue. Output the spontaneous combustion tendency level of coal gangue in the study area and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a backfill scheme for coal gangue that is not easy to spontaneously combust, a backfill scheme for coal gangue that is easy to spontaneously combust, and a backfill scheme for coal gangue that is extremely easy to spontaneously combust.
[0093] This invention also provides a machine learning-based system for classifying and backfilling the spontaneous combustion tendency of coal gangue, which includes a data collection unit, a data augmentation unit, a model building unit, an optimal model selection unit, and a model prediction unit.
[0094] The data collection unit is used to acquire a dataset with category labels of coal gangue spontaneous combustion tendency levels based on the K-means clustering algorithm. The dataset is divided into training set and test set. The data input consists of 9 cleaned and standardized features, and the output is the category labels. The category labels include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content, sulfur content, moisture content, ash content, volatile matter, alumina, silicon dioxide, calcium oxide, and sulfur trioxide.
[0095] The data augmentation unit is used to augment the training set using a synthetic minority class oversampling algorithm, expanding the number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible to a unified target number. The consistency of the training set distribution before and after augmentation is verified and analyzed from three dimensions: single-feature distribution, distribution differences between categories, and multi-feature association patterns. The analysis shows that the distribution is consistent. For the test set, a synthetic sample augmentation method based on metadata distribution is used for data augmentation. During the augmentation process, the distribution ratio of metadata categories in the test set remains unchanged. The consistency of the test set distribution before and after augmentation is verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization deviation. The analysis shows that the distribution is consistent.
[0096] The model building unit selects five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—as base models to construct a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set is input into each base model to complete the fitting and parameter tuning. Subsequently, predictions are executed on the preprocessed augmentation test set to output classification results, including class labels and probability values for each class, thus obtaining optimized base models. A Stacking ensemble learning framework is used to construct a Stacking two-layer ensemble learning model, using the five optimized base models as the first-layer learner and their prediction results as new fusion features. Logistic regression is used as the second-layer meta-learner, and training is completed using hierarchical 5-fold cross-validation to ensure that each class is evenly distributed in each fold, resulting in the optimized Stacking ensemble model.
[0097] The optimal model selection unit is used to perform SHAP interpretability analysis on all base models, quantify the positive and negative contributions of each feature to the classification results, identify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance index of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model.
[0098] The model prediction unit is used to input nine measured features of the study area into the coal gangue spontaneous combustion feature value input area, predict the spontaneous combustion category of coal gangue, output the spontaneous combustion tendency level of coal gangue in the study area, and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a backfill scheme for coal gangue that is not easy to spontaneously combust, a backfill scheme for coal gangue that is easy to spontaneously combust, and a backfill scheme for coal gangue that is extremely easy to spontaneously combust.
[0099] This invention also provides a coal gangue spontaneous combustion tendency classification and backfilling decision support system. This system adopts a "B / S architecture + front-end and back-end separation + three-layer logical layering" design pattern to decouple interface interaction, business logic and data storage, and ensure the scalability, maintainability and stability of the system.
[0100] This system is implemented based on a B / S (Browser / Server) architecture, allowing users to access all functions through mainstream browsers without the need for client installation, facilitating deployment and maintenance. It also employs a front-end / back-end separation architecture, with the front-end responsible for interface rendering and interaction logic, and the back-end handling business processing and model inference. The two communicate via a RESTful API in JSON format, enabling independent iteration and efficient collaboration.
[0101] The overall architecture design of this system is as follows: Figure 10 As shown, at the logical level, the system is divided into three layers: presentation layer, application layer, and data layer.
[0102] 1. Presentation Layer: Responsible for user interaction and data visualization. Built on Vue 3, this single-page application (SPA) includes five modules: registration and login, backfill point monitoring, category decision-making, history, and backfill knowledge. It utilizes a technology stack of Vue 3 + Vite + VueRouter + Pinia + Axios + ECharts + Gaode Maps. Data requests and result rendering are completed by calling the backend RESTful API in JSON format via HTTPS / HTTP protocol.
[0103] 2. Application Layer: The core business logic layer of the system, receiving requests from the presentation layer and handling business rules and model inference. Backend services are built on FastAPI, providing five API modules: authentication, display, prediction, history, and backfilling knowledge. Identity verification is implemented through JWTToken, the Stacking ensemble learning model engine is called to complete spontaneous combustion tendency classification, and CRUD operations are implemented for prediction records. Backfilling technical specifications can be queried and specification files can be downloaded. The model engine uses RandomForest, XGBoost, CatBoost, LightGBM, and SVC as base models, fusing results through a multi-class logistic regression meta-model. The technology stack is Python 3.11 + FastAPI + Pydantic + SQLAlchemy + Uvicorn.
[0104] 3. Data Layer: Responsible for data persistence and storage management, providing data read and write services to the application layer. It uses the lightweight SQLite database and implements data persistence through SQLAlchemy ORM. It includes three core data tables: users (stores user identity information such as phone number, password, and creation time), prediction_records (stores prediction requests and results, feature parameters, classification results, and operation time), and backfill_knowledge (stores backfill technical specification entries and file information). Data is stored locally as files to ensure data security and read / write efficiency.
[0105] This architecture achieves the design goal of high cohesion and low coupling through layered decoupling, which not only ensures the ease of use and compatibility of the system, but also provides good support for subsequent functional expansion and model optimization.
[0106] This system includes a registration and login module, a backfill point monitoring module, a classification decision module, a history module, and a backfill knowledge module.
[0107] The registration and login module serves as the core identity authentication entry point of the system, integrating user login and registration functions. The interface adopts a unified login and registration design, achieving a closed-loop operation process where unregistered users are redirected to registration, and registered users log in directly. During user registration, users must sequentially fill in their mobile phone number, a login password of at least 6 characters, a confirmation password, and a local verification code. The system will verify each piece of information in real time. If the mobile phone number has already been registered, the two passwords do not match, or the verification code is incorrect, corresponding prompts will be triggered. After all information is verified successfully, the system will indicate successful registration and automatically redirect to the login page, completing the account registration and login process. During user login, users must enter their mobile phone number, login password, and local verification code to complete identity verification. The password input box supports plaintext and encrypted text switching for easy verification. The system will automatically verify the legality of the input information. If the mobile phone number is not registered, the password is incorrect, or the verification code is mismatched, corresponding prompts will appear. After all information is verified successfully, the system will indicate successful login and automatically redirect to the system homepage.
[0108] The backfill point monitoring module, with its monitoring page serving as the system's core visualization and data display platform, integrates four core modules—spatial distribution display, mine area overview statistics, spontaneous combustion characteristic weight analysis, and spontaneous combustion risk level early warning—to meet the management needs of coal gangue backfill points in a specific region (such as Ordos City). This achieves comprehensive data presentation from macro-regional assessment to micro-risk identification, providing fundamental data support for subsequent classification of coal gangue spontaneous combustion tendencies and backfilling decisions. The backfill point monitoring interface is shown below. Figure 11 As shown.
[0109] The top left area of the main page displays the geographical location of the mining areas in Ordos City, based on the Gaode Map API. This visualization of the Ordos City mining area research area is achieved through a selection of key coal mine locations covered in this invention (some are not fully displayed due to geolocation limitations). Users can hover their mouse over blue fluorescent markers on the map to quickly view the name of the corresponding coal mine and intuitively understand its approximate location and spatial layout within the Ordos City area. This module integrates "location—identification—spatial analysis" of the research area, providing a visual basis for spatial planning in the regional backfilling work.
[0110] The upper right corner of the main page: This module presents core statistical data for the Ordos mining area as of mid-2025 in a data panel format. Specifically, it includes annual coal gangue disposal volume (10,000 tons), open-pit mine land reclamation area (square kilometers), subsidence area remediation area (square kilometers), comprehensive utilization rate of coal gangue (including resource utilization dimension), number of green mines, and number of monitored coal mines. Through this centralized display of quantitative data, users can quickly grasp the overall status of regional coal gangue resource utilization, ecological restoration, and monitoring coverage, providing data references for macro-level decision-making.
[0111] The bottom left area of the main page: This module uses a bar chart visualization to clearly present the weight distribution of the nine core influencing features (such as S, M, FC, etc.) of coal gangue spontaneous combustion tendency. This module serves as a prerequisite for the subsequent "Classification Decision" module, intuitively quantifying the impact of each feature on spontaneous combustion risk, identifying core influencing factors, and helping users to understand the key driving factors of coal gangue spontaneous combustion in advance. This provides an intuitive basis for the importance of features in subsequent spontaneous combustion tendency classification models based on multi-feature inputs, strengthening the logical connection between functional modules.
[0112] The bottom right area of the main page: This module constructs a "three-level early warning" visualization system based on the correlation between the internal temperature of coal gangue and the risk of spontaneous combustion. Specific early warning rules are: Green warning (low risk), internal temperature 60-120℃, gangue slowly oxidizing, with potential heat accumulation hazards; Yellow warning (medium risk), internal temperature 120-220℃, gangue internal heat accumulation and temperature rise, in a critical spontaneous combustion state; Red warning (high risk), internal temperature >220℃, active stage of spontaneous combustion of coal gangue, with open flames or severe oxidation, requiring immediate control measures. The module clearly marks the color-coded levels and temperature ranges to achieve rapid identification and intuitive early warning of coal gangue spontaneous combustion risks, providing direct guidance for risk control in on-site backfilling operations.
[0113] This interface adopts a four-in-one design logic of "spatial monitoring - data statistics - feature analysis - risk warning". It not only realizes the precise positioning of the research area and the comprehensive display of the mining area overview, but also connects the subsequent classification and decision-making functions through the pre-presentation of feature weights and risk warnings, forming a complete business closed loop of "data visualization - feature cognition - risk assessment - decision support", which fully reflects the practicality and logical coherence of the system.
[0114] The classification and decision-making module is the core functional unit of the coal gangue spontaneous combustion tendency classification and backfilling decision support system. It adopts a left-right column layout and integrates intelligent classification of coal gangue spontaneous combustion tendency and backfilling scheme decision-making functions, realizing an integrated closed-loop operation of "feature input—risk prediction—scheme generation," providing direct decision support for the safe backfilling and resource utilization of coal gangue. The system classification and decision-making interface is as follows: Figure 12 As shown.
[0115] The left-hand area is the input area for spontaneous combustion characteristics of coal gangue: it provides input interfaces for nine characteristics, including fixed carbon content (FC), sulfur content (S), moisture content (M), ash content (Ad), volatile matter (Vad), alumina (Al2O3), silicon dioxide (SiO2), calcium oxide (CaO), and sulfur trioxide (SO3). Users can enter measured data and click "Start Prediction" to predict the spontaneous combustion category of coal gangue.
[0116] The right-hand area is the output area for prediction results and backfill schemes: After the user inputs complete feature values and clicks the "Start Prediction" button, the system uses a built-in classification model to intuitively present the spontaneous combustion tendency level in three colors: "green (not likely to spontaneously combust), yellow (prone to spontaneously combust), and red (extremely prone to spontaneously combust)". At the same time, without any additional operation, the system automatically matches the backfill scheme corresponding to the risk level and simultaneously displays detailed decision-making content such as backfill suggestions, layering structure, compaction coefficient, and reference cases.
[0117] The module supports result archiving and iterative prediction: Clicking the "Save Results" button will archive the current prediction and plan to the history module for easy retrospective review; the "Reset" button on the left or the "New Prediction" button on the right can be used to clear the data and results with one click and restore the initial blank interface, supporting continuous decision analysis in multiple batches and scenarios.
[0118] This module encapsulates complex classification models and engineering decisions into a simple operational process, achieving seamless integration from risk assessment to backfilling scheme implementation, effectively improving the scientific nature and efficiency of coal gangue spontaneous combustion prevention and backfilling operations.
[0119] The historical data module is the core of the coal gangue spontaneous combustion tendency classification and backfilling decision support system, used for centralized management of historical prediction results saved during the classification decision-making stage, enabling visualized statistics, full display, and efficient management of decision data. The system's historical data interface is shown below. Figure 13 As shown.
[0120] The top of the module displays the total number of records and the number of records in three risk levels—not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible—in the form of data cards, distinguished by green, yellow, and red colors to help users quickly grasp the risk distribution characteristics of historical decisions. Below, a timeline is displayed for each historical record, including the record number, generation time, and spontaneous combustion tendency tag. It supports refresh operations to synchronize the latest saved prediction data, and also provides single-record deletion and batch deletion functions for selecting all records, allowing for flexible cleanup of redundant records. After expanding a single record, users can fully trace back the nine input characteristic parameters, supporting backfilling plans, and construction parameters of this decision, achieving full-chain traceability from "input to output."
[0121] The backfilling knowledge module is the core of the technical specifications and knowledge accumulation of the coal gangue spontaneous combustion tendency classification and backfilling decision support system. It provides users with a standardized knowledge base and specification support covering the entire backfilling process, realizing the structured display, dynamic expansion and compliance management of technical knowledge.
[0122] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these are merely illustrative examples, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but all such changes and modifications fall within the scope of protection of the present invention.
Claims
1. A machine learning-based method for classifying and backfilling the spontaneous combustion tendency of coal gangue, characterized in that, It includes the following steps: S1. Data Collection: Obtain a dataset with category labels for the spontaneous combustion tendency level of coal gangue based on the K-means clustering algorithm, and divide the dataset into training set and test set. The data input consists of 9 cleaned and standardized features, and the output consists of category labels. The category labels include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content, sulfur content, moisture content, ash content, volatile matter, alumina, silicon dioxide, calcium oxide, and sulfur trioxide. S2. Data Augmentation: The training set is augmented using a synthetic minority class oversampling algorithm. The number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible in the training set is expanded to a uniform target number. The consistency of the training set distribution before and after augmentation is verified and analyzed from three dimensions: single feature distribution, distribution differences between categories, and multi-feature association patterns. The analysis shows that the distribution is consistent. The test set is augmented using a synthetic sample expansion method based on metadata distribution. During the augmentation and expansion process, the distribution ratio of metadata categories in the test set remains unchanged. The consistency of the test set distribution before and after augmentation is verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization deviation. The analysis shows that the distribution is consistent. S3. Model Construction: Five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—were selected as base models to construct a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set was input into each base model to complete the fitting and parameter tuning. Subsequently, the prediction output classification results were performed on the preprocessed augmentation test set. The classification results included class labels and probability values of each class, obtaining the optimized base models after tuning. The Stacking ensemble learning framework was used to construct a Stacking two-layer ensemble learning model. The five optimized base models were used as the first layer learners, and their prediction results were used as new fusion features. Logistic regression was used as the second layer meta-learner. Layered 5-fold cross-validation was used to complete the training to ensure that each class was evenly distributed in each fold, obtaining the optimized Stacking ensemble model after tuning. S4. Optimal Model Selection: Perform SHAP interpretability analysis on all base models to quantify the positive and negative contributions of each feature to the classification results, identify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance indicators of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model. S5. Model Prediction: Construct a classification decision module. Input the nine measured features of the study area into the coal gangue spontaneous combustion feature value input area to predict the spontaneous combustion category of coal gangue. Output the spontaneous combustion tendency level of coal gangue in the study area and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a backfill scheme for coal gangue that is not easy to spontaneously combust, a backfill scheme for coal gangue that is easy to spontaneously combust, and a backfill scheme for coal gangue that is extremely easy to spontaneously combust.
2. The method for classifying and backfilling the spontaneous combustion tendency of coal gangue based on machine learning as described in claim 1, characterized in that, In S2, the augmented training set and the augmented test set are stored in Excel format; In S3, the enhanced training set and enhanced test set in Excel format are read, and redundant columns between the enhanced training set and enhanced test set are automatically removed to ensure that the feature dimensions are completely consistent. Then, the non-numerical features are adapted and transformed. StandardScaler is used to perform standardization processing on all features to eliminate the difference in units. Finally, LabelEncoder is used to convert the spontaneous combustion tendency level into digital code to obtain the preprocessed enhanced training set and enhanced test set.
3. The machine learning-based method for classifying and backfilling the spontaneous combustion tendency of coal gangue as described in claim 1, characterized in that, In S3, the Stacking two-layer ensemble learning model uses hierarchical 5-fold cross-validation for training to ensure that each class is evenly distributed in each fold. First, the preprocessed augmented training set is divided into 5 parts. Four parts are used to train and optimize the base model, and the remaining part is used to generate the prediction probability. After iterating through all the folds, the base model probability meta-features with the same length as the augmented training set are obtained. Subsequently, the class probabilities of the five optimized base models are horizontally integrated to form a high-dimensional meta-feature matrix, which is then input into the logistic regression model to train and obtain the optimal fusion weights for each optimized base model. Finally, the optimized base models perform probability prediction on the preprocessed augmented test set, which is then input into the trained meta-model to output the final integrated classification result and class probabilities.
4. The machine learning-based method for classifying and backfilling the spontaneous combustion tendency of coal gangue as described in claim 1, characterized in that, In S4, classification performance metrics are evaluated for all models, including accuracy, precision, recall, F1 score, and AUC score. Perform SHAP interpretability analysis on all base models, including category-level SHAP feature importance, SHAP bee colony graph, interaction graph, waterfall plot, and heatmap.
5. A machine learning-based system for classifying and backfilling the spontaneous combustion tendency of coal gangue, characterized in that, It includes: The data collection unit is used to acquire a dataset with category labels of coal gangue spontaneous combustion tendency levels based on the K-means clustering algorithm, and divide the dataset into training set and test set. The data input consists of 9 cleaned and standardized features, and the output is the category labels, which include non-spontaneous combustion, spontaneous combustion, and extremely spontaneous combustion. The 9 features are fixed carbon content, sulfur content, moisture content, ash content, volatile matter, alumina, silicon dioxide, calcium oxide, and sulfur trioxide. The data augmentation unit is used to augment the training set using a synthetic minority class oversampling algorithm, expanding the number of samples that are not easily spontaneously combustible, easily spontaneously combustible, and extremely easily spontaneously combustible to a unified sampling target number. The consistency of the training set distribution before and after augmentation is verified and analyzed from three dimensions: single-feature distribution, distribution differences between categories, and multi-feature association patterns. The analysis shows that the distribution is consistent. For the test set, a synthetic sample augmentation method based on metadata distribution is used for data augmentation. During the augmentation process, the distribution ratio of metadata categories in the test set remains unchanged. The consistency of the test set distribution before and after augmentation is verified and analyzed from three dimensions: morphological consistency, statistical consistency, and quantization deviation. The analysis shows that the distribution is consistent. The model building unit selects five representative algorithms—random forest, extreme gradient boosting, classification boosting tree, lightweight gradient boosting machine, and support vector machine—as base models to build a basic classification model library. Under a unified augmentation dataset and experimental settings, during model training, the preprocessed augmentation training set is input into each base model to complete the fitting and achieve parameter tuning. Subsequently, the prediction output classification results are performed on the preprocessed augmentation test set, including class labels and probability values of each class, to obtain the optimized base models. The Stacking ensemble learning framework is used to build a Stacking two-layer ensemble learning model, using the five optimized base models as the first layer learner and their prediction results as new fusion features. Logistic regression is used as the second layer meta-learner, and training is completed using hierarchical 5-fold cross-validation to ensure that each class is evenly distributed in each fold, thus obtaining the optimized Stacking ensemble model. The optimal model selection unit is used to perform SHAP interpretability analysis on all base models, quantify the positive and negative contributions of each feature to the classification results, clarify the core features affecting the classification of coal gangue spontaneous combustion tendency, evaluate the classification performance index of all models, select the optimized Stacking ensemble model as the optimal Stacking ensemble model, and derive the optimal Stacking ensemble model. The model prediction unit is used to input nine measured features of the study area into the coal gangue spontaneous combustion feature value input area, predict the spontaneous combustion category of coal gangue, output the spontaneous combustion tendency level of coal gangue in the study area, and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a non-spontaneous combustion coal gangue backfill scheme, a spontaneous combustion coal gangue backfill scheme, and an extremely spontaneous combustion coal gangue backfill scheme.
6. The machine learning-based coal gangue spontaneous combustion tendency classification and backfilling system as described in claim 5, characterized in that, In the data augmentation unit, the augmented training set and the augmented test set are stored in Excel format; In the model building unit, the enhanced training set and enhanced test set in Excel format are read, and redundant columns between the enhanced training set and enhanced test set are automatically removed to ensure that the feature dimensions are completely consistent. Then, the non-numerical features are adapted and transformed. StandardScaler is used to perform standardization processing on all features to eliminate differences in units. Finally, LabelEncoder is used to convert the spontaneous combustion tendency level into digital code to obtain the preprocessed enhanced training set and enhanced test set.
7. The machine learning-based coal gangue spontaneous combustion tendency classification and backfilling system as described in claim 5, characterized in that, In the model building unit, the Stacking two-layer ensemble learning model uses hierarchical 5-fold cross-validation for training to ensure a balanced distribution of each class in each fold. First, the preprocessed augmented training set is divided into 5 parts. Four parts are used to train and optimize the base model, and the remaining part is used to generate the prediction probability. After iterating through all the folds, the base model probability meta-features with the same length as the augmented training set are obtained. Subsequently, the class probabilities of the five optimized base models are horizontally integrated to form a high-dimensional meta-feature matrix, which is then input into the logistic regression model to train and obtain the optimal fusion weights for each optimized base model. Finally, the optimized base models perform probability prediction on the preprocessed augmented test set, which is then input into the trained meta-model to output the final integrated classification result and class probabilities.
8. The machine learning-based coal gangue spontaneous combustion tendency classification and backfilling system as described in claim 5, characterized in that, In the optimal model selection unit, all models are evaluated for classification performance metrics, including accuracy, precision, recall, F1 score, and AUC score. Perform SHAP interpretability analysis on all base models, including category-level SHAP feature importance, SHAP bee colony graph, interaction graph, waterfall plot, and heatmap.
9. A decision support system for classifying and backfilling spontaneous combustion tendencies of coal gangue, characterized in that, The design pattern adopts a B / S architecture + front-end and back-end separation + three-layer logical layering to decouple the interface interaction, business logic and data storage. The three-layer logical layering is divided into presentation layer, application layer and data layer. The front-end uses Vue 3 + Vite + Vue Router + Pinia + Axios + ECharts + Gaode Map technology stack, and the back-end uses Python 3.11 + FastAPI + Pydantic + SQLAlchemy + Uvicorn technology stack. The database uses the lightweight SQLite database, and data persistence is achieved through SQLAlchemy ORM. The system includes a registration and login module, a backfill point monitoring module, a classification decision module, a history module, and a backfill knowledge module; The registration and login module is used for user registration and login, and implements user identity authentication; The backfill point monitoring module is used to display the geographical location of each mining area in a certain region on an electronic map, and to specifically display the visual distribution of the study area in a certain mining area. It displays the core statistical data of the mining area in the form of a data panel, and uses a bar chart to visualize the weight distribution of nine characteristics of the spontaneous combustion tendency of coal gangue. Based on the correlation between the internal temperature of coal gangue and the risk of spontaneous combustion, a three-level early warning visualization system is constructed. The classification decision module is used to input nine measured features of the study area into the coal gangue spontaneous combustion feature value input area, predict the spontaneous combustion category of coal gangue, output the spontaneous combustion tendency level of coal gangue in the study area, and automatically match the corresponding coal gangue classification backfill scheme. The coal gangue classification backfill scheme includes a non-spontaneous combustion coal gangue backfill scheme, a spontaneous combustion coal gangue backfill scheme, and an extremely spontaneous combustion coal gangue backfill scheme. The history module is used for centralized management of historical prediction results saved during the classification decision-making stage; The backfilling knowledge module provides users with a standardized knowledge base and specification support covering the entire backfilling process.