Method for rapidly predicting volatile components of coal quality

Through the combination of a transmissive terahertz time domain spectrometer and machine learning regression model, the problems of time-consuming and noise-influence of traditional coal quality volatile components are solved, and fast, lossless high-precision volatile components are achieved.

CN120507312APending Publication Date: 2025-08-19NAT ENERGY COAL & COKING GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351473.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional coal-quality volatile components detection methods require high temperature heating, which consumes time and is prone to damage samples. The existing terahertz time domain spectroscopy technology is susceptible to noise, making it difficult to achieve rapid non-destructive detection and effective mining of nonlinear correlation between volatile components and spectral data.

Method used

A transmissive terahertz time domain spectrometer was used to collect spectral data, combine standardized processing and dimensionality reduction technology to build a machine learning regression model, optimize hyperparameters through grid search method, and build a final prediction model for volatile fraction prediction.

Benefits of technology

It realizes rapid and non-destructive detection of volatile coal-quality components, improves the efficiency of spectrum data feature extraction, achieves high accuracy and strong generalization capabilities, and has a decisive coefficient of 0.985 and a root mean square error of 1.949.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120507312A_ABST
    Figure CN120507312A_ABST
Patent Text Reader

Abstract

The invention discloses a coal volatile component rapid prediction method, and belongs to the technical field of coal volatile component prediction. Comprising the following steps: acquiring spectral data of a to-be-detected sample through a transmission-type terahertz time-domain spectrometer; performing standardization processing and dimensionality reduction on the spectral data in sequence, and screening out principal component characteristics related to volatile components to form a data set; constructing a first machine learning regression model, training the first machine learning regression model through the data set, and optimizing hyper-parameters of the first machine learning regression model by adopting a grid search method; and constructing a second machine learning regression model based on the optimized hyper-parameters, calculating a learning curve and a residual image of the second machine learning regression model, screening out a final prediction model, and performing volatile component prediction on the to-be-detected sample. The method has the beneficial effects that nondestructive testing is realized based on the terahertz time-domain spectroscopy and the machine learning regression model, the feature extraction efficiency of spectral data is improved, and high precision and strong generalization ability of volatile component prediction are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of coal volatile matter prediction, and in particular to a method for rapid prediction of volatile matter. Background Art

[0002] Coal occupies a central position in my country's energy system, and its quality parameters (such as volatile matter content) have a direct impact on combustion efficiency and pollutant emissions. Traditional volatile matter analysis methods require heating coal samples in a high-temperature, oxygen-free environment and measuring mass loss, which is complex and time-consuming. In recent years, terahertz time-domain spectroscopy (THz-TDS) technology and machine learning algorithms have been applied in the field of coal analysis.

[0003] In existing technologies, traditional coal volatile matter detection methods require high-temperature heating of coal samples to measure mass loss, which not only destroys the samples but also takes a long time, making it difficult to meet the actual needs of rapid and non-destructive testing on coal production lines. Although terahertz time-domain spectroscopy technology can obtain coal sample spectral data through non-destructive testing, it is easily affected by the surface roughness, internal scattering and instrument noise of the coal samples. Existing spectral feature extraction mostly relies on a single principal component analysis method, which can only capture linear relationships and cannot effectively explore the potential nonlinear correlation between volatiles and spectral data. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for quickly predicting coal volatile matter to solve the above technical problems;

[0005] A method for quickly predicting coal volatile matter, comprising:

[0006] Step S1, collecting spectral data of the sample to be tested by a transmission terahertz time-domain spectrometer;

[0007] Step S2, performing standardization and dimensionality reduction on the spectral data in sequence, screening out principal component features related to volatiles, and forming a data set;

[0008] Step S3: constructing a first machine learning regression model, training the first machine learning regression model using the data set, and optimizing hyperparameters of the first machine learning regression model using a grid search method;

[0009] Step S4: construct a second machine learning regression model based on the optimized hyperparameters, calculate the learning curve and residual graph of the second machine learning regression model, and screen out the final prediction model based on the learning curve and the residual graph to predict the volatile matter of the sample to be tested.

[0010] Preferably, the transmission terahertz time-domain spectrometer in step S1 includes a femtosecond laser, a terahertz wave generator, a delay device and a detector, and the spectral data collection step includes:

[0011] Step S11, emitting a first pulse laser by the femtosecond laser;

[0012] Step S12, the first pulse laser is divided into a pump pulse and a probe pulse by a beam splitter;

[0013] Step S13, the pump pulse acts on the terahertz wave generator to generate a terahertz wave;

[0014] Step S14, the terahertz wave passes through the sample to be tested, the detection pulse is combined with the terahertz wave after the delay is adjusted by the delay device, and the terahertz pulse time domain signal is obtained by the detector;

[0015] Step S15: The terahertz pulse time domain signal is transformed and calibrated to obtain the spectral data.

[0016] Preferably, step S2 uses one of principal component analysis, kernel principal component analysis, sequence principal component analysis and incremental principal component analysis to reduce the dimension of the spectral data.

[0017] Preferably, step S3 includes,

[0018] Step S31, constructing the first machine learning regression model;

[0019] Step S32: dividing the data set into a training set and a test set according to a preset ratio, using the training set to train the first machine learning regression model, and using the test set to test the first machine learning regression model;

[0020] Step S33: using the grid search method to optimize the hyperparameters of the first machine learning regression model.

[0021] Preferably, the preset ratio in step S32 is 4:1, and the hyperparameters of the first machine learning regression model in step S33 include the maximum tree depth, the minimum number of leaf node samples, and the minimum number of split samples;

[0022] The maximum tree depth is 10-20, the minimum number of leaf node samples is 1-5, and the minimum number of split samples is 5-15.

[0023] Preferably, the second machine learning regression model in step S4 includes principal component analysis random forest, kernel principal component analysis random forest, sequence principal component analysis random forest and incremental principal component analysis random forest.

[0024] Preferably, step S4 includes,

[0025] Step S41, constructing the second machine learning regression model based on the optimized hyperparameters;

[0026] Step S42, evaluating the second machine learning regression model through ten-fold cross validation according to the first evaluation index to obtain the learning curve and the residual graph;

[0027] Step S43: Filter out the second machine learning regression model with the highest first evaluation index as the final prediction model to perform volatile content prediction on the sample to be tested.

[0028] Preferably, in step S42, the first evaluation index includes the coefficient of determination, the root mean square error and the mean absolute error.

[0029] Preferably, the coefficient of determination of the final prediction model is 0.985, the root mean square error is 1.949, and the mean absolute error is 0.913.

[0030] Preferably, the sample to be tested in step S1 and step S4 is coal.

[0031] The beneficial effects of the present invention are: non-destructive testing is achieved based on terahertz time-domain spectroscopy and machine learning regression model, the efficiency of feature extraction of spectral data is improved, and high precision and strong generalization ability of volatile matter prediction are achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a step diagram of the method for quickly predicting coal volatile matter of the present invention;

[0033] Figure 2 is a schematic diagram of step S1 of the present invention;

[0034] Figure 3 is a schematic diagram of step S3 of the present invention;

[0035] Figure 4 is a schematic diagram of step S4 of the present invention;

[0036] Figure 5 is a spectrum data acquisition flow chart of the present invention;

[0037] Figure 6 This is a flow chart of the method for rapid prediction of coal volatile matter of the present invention;

[0038] FIG7( a ) is a schematic diagram of the principal component analysis scatter plot of the present invention;

[0039] FIG7( b ) is a schematic diagram of the kernel principal component analysis scatter plot of the present invention;

[0040] FIG7( c ) is a schematic diagram of the scatter plot of the principal component analysis of the sequence of the present invention;

[0041] FIG7( d ) is a schematic diagram of the incremental principal component analysis scatter plot of the present invention;

[0042] FIG8( a ) is a diagram showing the model training results of the principal component analysis random forest of the present invention;

[0043] FIG8( b ) is a diagram showing the model test results of the principal component analysis random forest of the present invention;

[0044] FIG9( a ) is a diagram showing the model training results of the kernel principal component analysis random forest of the present invention;

[0045] FIG9( b ) is a diagram showing the model test results of the kernel principal component analysis random forest of the present invention;

[0046] FIG10( a ) is a diagram showing the model training results of the random forest method for sequence principal component analysis according to the present invention;

[0047] FIG10( b ) is a diagram showing the model test results of the random forest sequence principal component analysis of the present invention;

[0048] FIG11( a ) is a diagram showing the model training results of the incremental principal component analysis random forest of the present invention;

[0049] FIG11( b ) is a graph showing the model test results of the incremental principal component analysis random forest of the present invention;

[0050] FIG12( a ) is a learning curve diagram of the principal component analysis random forest of the present invention;

[0051] FIG12( b ) is a graph showing the learning curve of the kernel principal component analysis random forest of the present invention;

[0052] FIG12( c ) is a learning curve diagram of the random forest for sequence principal component analysis of the present invention;

[0053] FIG12( d ) is a graph showing the learning curve of the incremental principal component analysis random forest of the present invention;

[0054] FIG13( a ) is a residual plot of the principal component analysis random forest of the present invention;

[0055] FIG13( b ) is a residual plot of the kernel principal component analysis random forest of the present invention;

[0056] FIG13( c ) is a residual graph of the random forest of sequence principal component analysis of the present invention;

[0057] FIG13( d ) is a residual graph of the incremental principal component analysis random forest of the present invention. DETAILED DESCRIPTION

[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0059] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0060] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.

[0061] A method for rapid prediction of coal volatile matter, such as Figure 1 、 Figure 6 Shown, including,

[0062] Step S1, collecting spectral data of the sample to be tested by a transmission terahertz time-domain spectrometer;

[0063] Step S2, performing standardization and dimensionality reduction on the spectral data in sequence, screening out the principal component features related to volatiles, and forming a data set;

[0064] Step S3: constructing a first machine learning regression model, training the first machine learning regression model using the data set, and optimizing the hyperparameters of the first machine learning regression model using a grid search method;

[0065] Step S4: construct a second machine learning regression model based on the optimized hyperparameters, calculate the learning curve and residual graph of the second machine learning regression model, and screen out the final prediction model based on the learning curve and residual graph to predict the volatile content of the sample to be tested.

[0066] Specifically, the present invention provides a method for rapid prediction of coal volatile matter, which realizes non-destructive testing of the sample to be tested based on terahertz time-domain spectroscopy and machine learning regression model, improves the efficiency of feature extraction of spectral data, and achieves high precision and strong generalization ability of volatile matter prediction.

[0067] In a preferred embodiment, referring to Figure 2 In step S1, the transmission terahertz time-domain spectrometer includes a femtosecond laser, a terahertz wave generator, a delay device and a detector. The spectrum data collection step includes:

[0068] Step S11, emitting a first pulse laser by a femtosecond laser;

[0069] Step S12, the first pulse laser is divided into a pump pulse and a probe pulse by a beam splitter;

[0070] Step S13, the pump pulse acts on the terahertz wave generator to generate a terahertz wave;

[0071] Step S14: The terahertz wave passes through the sample to be tested, and the detection pulse is combined with the terahertz wave after the time is adjusted by the delay device, and the terahertz pulse time domain signal is obtained by the detector;

[0072] Step S15: The terahertz pulse time domain signal is transformed and calibrated to obtain spectrum data.

[0073] Specifically, refer to Figure 5 The spectral acquisition device is a transmission-mode terahertz time-domain spectrometer consisting of a femtosecond laser, a terahertz wave generator, a detector, and a delay mechanism. The device's operation begins with an ultrashort laser pulse (i.e., the first pulse) emitted by the femtosecond laser (fiber laser source). This laser is split into two by a beam splitter, forming a fiber optical path consisting of a pump pulse and a probe pulse.

[0074] The pump pulse acts on the terahertz wave generator, inducing the generation of terahertz waves to form a terahertz optical path. The terahertz waves are then collimated and pass through a thin slice of the sample to be tested that is precisely centered on the sample stage.

[0075] At the same time, the probe pulse is time-adjusted through an adjustable delay system to achieve time-resolved detection of the terahertz wave. It then combines with the terahertz wave and is ultimately captured by the detector. The system utilizes lock-in amplification technology to synchronously acquire the detector output signal, optimizing the signal-to-noise ratio and obtaining a clear terahertz pulse time-domain signal. This signal is collected by the data collector and displayed on the main control screen, which is connected to the power supply system for power.

[0076] In a preferred embodiment, step S2 uses one of principal component analysis (PCA), kernel principal component analysis (KPCA), sequential principal component analysis (SPCA) and incremental principal component analysis (IPCA) to reduce the dimension of the spectral data.

[0077] Specifically, terahertz spectral data is typically high-dimensional and contains a large number of variables. This variable data contains a lot of redundant information, which not only increases computational complexity but also can interfere with the model's ability to capture key information. Principal component analysis uses linear transformations to project the original high-dimensional data onto a new set of orthogonal coordinate axes, known as principal components.

[0078] The new principal components are sorted by data variance, retaining those with larger variances. This significantly reduces the data's dimensionality while preserving its key features. When analyzing terahertz spectral data from a sample, the original data may contain hundreds of variables. After PCA dimensionality reduction, only 10-20 principal components may be retained, which summarizes the majority of the data, removes a significant amount of redundant information, and improves the efficiency and accuracy of subsequent model training.

[0079] PCA and its variants can uncover hidden features related to volatiles in spectral data. Complex correlations exist between volatiles in a sample and spectral data, and these algorithms can help uncover these hidden relationships. Kernel principal component analysis, by introducing a kernel function, maps data into a high-dimensional space, thereby capturing nonlinear relationships within the data. In sample analysis, nonlinear interactions exist between volatiles and spectral data. KPCA can better reveal these relationships, identifying spectral features closely related to volatiles and providing strong support for accurate volatile content prediction.

[0080] Data that has undergone dimensionality reduction can effectively improve the performance of machine learning models. High-dimensional data can easily lead to overfitting of the model, causing the model to perform well on the training set but poorly on the test set or new data.

[0081] By reducing dimensionality and removing noise and redundant information, the model can focus more on key features, thereby improving its generalization capabilities. When building a random forest regression model to predict coal volatiles, training on data that has undergone dimensionality reduction using PCA or other variants allows the model to more accurately learn the relationship between volatiles and spectral features. This allows for more stable predictions across different batches of coal samples, reducing errors.

[0082] In a preferred embodiment, referring to Figure 3 , step S3 includes,

[0083] Step S31, constructing a first machine learning regression model;

[0084] Step S32: Divide the data set into a training set and a test set according to a preset ratio, use the training set to train the first machine learning regression model, and use the test set to test the first machine learning regression model;

[0085] Step S33: Optimize the hyperparameters of the first machine learning regression model using a grid search method.

[0086] Specifically, hyperparameters, which are parameters that must be set before model training, are primarily responsible for regulating the model's learning ability and complexity. The goal of hyperparameter optimization is to identify the parameter combination that maximizes model performance. This not only aims to improve the model's accuracy and generalization ability, but also to avoid overfitting and underfitting, ensure model stability and robustness, and achieve efficient model training with limited computing resources.

[0087] In step S32, the dataset is divided into a training set and a test set according to a preset ratio. This allows the model to learn the characteristics and patterns of the data during training, while simultaneously evaluating its adaptability to new data during testing. The training set is used to train the model, allowing it to gradually grasp the relationship between spectral data and coal volatile matter by learning from the samples in the training set. The test set, independent of the training set, is used to test the model's predictive ability on unseen data. This prevents the model from overfitting—that is, from over-learning the noise and special cases in the training set—thus improving the accuracy of predictions for unknown data.

[0088] In step S33, a grid search method is used to optimize the hyperparameters of the first machine learning regression model. Hyperparameters are parameters that need to be set before model training. Different hyperparameter combinations will result in different model performance. The grid search method iterates through pre-set hyperparameter combinations, trains and evaluates the model for each combination, and selects the hyperparameter combination with the best performance.

[0089] This application uses a grid search approach for hyperparameter optimization, aiming to systematically explore and identify the optimal parameter combination within the hyperparameter space. Compared to random search, grid search ensures a comprehensive evaluation of each parameter combination by exhaustively traversing a predefined parameter grid, thereby increasing the likelihood of discovering the optimal hyperparameter combination.

[0090] Grid search is the preferred method for this study because of its advantages in small to medium-sized search spaces. We used the Python scikit-learn library to train and evaluate the machine learning model and optimize its hyperparameters.

[0091] In a preferred embodiment, the preset ratio in step S32 is 4:1, and the hyperparameters of the first machine learning regression model in step S33 include the maximum tree depth, the minimum number of leaf node samples, and the minimum number of split samples;

[0092] The maximum tree depth is 10-20, the minimum number of leaf node samples is 1-5, and the minimum number of split samples is 5-15;

[0093] The sample to be tested in step S1 and step S4 is coal.

[0094] Specifically, the test samples of coal samples were randomly divided into training set and test set in a ratio of 4:1.

[0095] First, the data from the training set is used to train the machine learning model, hoping that the model can effectively learn the characteristics and patterns of the data.

[0096] After training is completed, independent test set data is used to evaluate the performance of the model on new data to verify its generalization ability and prediction accuracy.

[0097] To ensure the reliability and stability of the evaluation process, a 10-fold cross-validation method was used. The training data was divided into 10 subsets of similar size, with one subset used as the validation set and the remaining nine used as the training set. This process was repeated 10 times, with a different validation set selected each time to reflect the diversity and variation of the dataset and thus reduce bias introduced by specific data splits.

[0098] Finally, by averaging the results of 10 cross-validation runs, a comprehensive and reliable evaluation of the model performance was obtained.

[0099] Predicting coal volatile matter is a regression task. When evaluating the performance of the prediction model, a variety of indicators are used for comprehensive analysis, including the coefficient of determination (R 2 ), root mean square error (RMSE) and mean absolute error (MAE). 2 It is used to measure the degree to which the data is explained by the model, that is, the degree of fit between the model prediction value and the actual value. RMSE reflects the average error between the predicted value and the actual value.

[0100] MAE provides the mean absolute error, further evaluating the accuracy of the model. Through these evaluation indicators, the performance of the model can be comprehensively evaluated to ensure its effectiveness and reliability in practical applications.

[0101] Specifically, PCA, KPCA, SPCA and IPCA were used to process the spectral data of coal samples respectively. By comparing the correlation between the absorbance and volatile matter of coal samples at different wavelengths, the spectral data dimension reduction and feature acquisition optimization were completed.

[0102] Figures 7(a)-(d) show the relationship between principal components 1 and 2 after dimensionality reduction using different algorithms. The color bars represent the volatile content. As can be seen from the figures, all four algorithms can perform well in clustering based on the volatile content.

[0103] Taking Figure 7(a) as an example, coal samples with a volatile matter percentage of 0-15 are mainly distributed between -60 and 5 of principal component 1 and -10 to 10 of principal component 2, while coal samples with a volatile matter percentage of 15-35 are mainly distributed between 50 and 90 of principal component 1 and -10 to 10 of principal component 2. The remaining volatile matter percentages greater than 35 are mainly distributed between 20 and 40 of principal component 1 and -30 to -15 of principal component 2. This shows that all four algorithms capture the connection between the features after dimensionality reduction and the volatile matter.

[0104] Specifically, a random forest regression model based on the optimized features was constructed. To ensure the model's predictive accuracy, a grid search method was used to adjust the model's hyperparameters. The model's three key hyperparameters were specifically adjusted: the maximum tree depth (max_depth), the minimum number of samples required for a node split (min_samples_leaf), and the minimum number of samples required for a leaf node (min_samples_split). These hyperparameter settings significantly impact the model's complexity and generalization ability.

[0105] By comprehensively testing different values of hyperparameters, the optimal parameter configuration is found to achieve a balance between model complexity and performance, thereby improving the model's prediction accuracy.

[0106] In a preferred embodiment, the second machine learning regression model in step S4 includes principal component analysis random forest, kernel principal component analysis random forest, sequence principal component analysis random forest and incremental principal component analysis random forest.

[0107] Specifically, principal component analysis is an unsupervised dimensionality reduction technique that projects the original data onto a new set of orthogonal coordinate axes through linear transformation. These coordinate axes are called principal components. The principal components are sorted by their variance; principal components with larger variances contain more data information. In the prediction of coal volatile matter, spectral data are usually high-dimensional and redundant. PCA can reduce these high-dimensional data to lower dimensions while retaining the main information of the data. Random forest is an ensemble learning method that makes predictions by constructing multiple decision trees and combining their results. Random forest has strong capabilities in processing high-dimensional data and complex nonlinear relationships. Combining PCA with random forest, PCA provides random forest with features that have been processed through dimensionality reduction and denoising, reducing the computational complexity and overfitting risk of random forest. At the same time, random forest can make full use of these features to make accurate predictions. The two work together to improve the performance of the model.

[0108] Kernel principal component analysis (KPCA) introduces a kernel function based on PCA. Kernel functions can map raw data into a high-dimensional feature space, making data that is nonlinearly separable in a low-dimensional space linearly separable in the high-dimensional space. The relationship between coal spectral data and volatiles is nonlinear. Traditional PCA can only handle linear relationships, while KPCA, by selecting an appropriate kernel function, can exploit nonlinear features in the data. In a high-dimensional feature space, random forests can better learn the relationship between these nonlinear features and volatiles, thereby improving prediction accuracy.

[0109] Sequential principal component analysis (SPCA) is suitable for processing data with time series or sequential characteristics. In coal quality analysis, samples may be collected in a specific order, and the characteristics of the data may change over time or over the sample sequence. SPCA can perform principal component analysis on data in a sequential manner, capturing dynamic changes in the data and updating the principal component information. When combined with SPCA results for prediction, random forests can better adapt to data changes and improve prediction accuracy.

[0110] Incremental principal component analysis (IPCA) is designed to handle large-scale data. In practical applications, the spectral data of coal samples may continue to grow. Performing principal component analysis on all data at once would be computationally intensive. IPCA can gradually update the principal components as data continues to grow, eliminating the need to recalculate the principal components for all data and improving computational efficiency. Furthermore, random forests combined with IPCA results can accurately predict new data while maintaining computational efficiency.

[0111] Four machine learning regression models were constructed using feature optimization algorithms: RF-PCA (Random Forest with Principal Component Analysis), RF-KPCA (Random Forest with Kernel Principal Component Analysis), RF-SPCA (Random Forest with Sequential Principal Component Analysis), and RF-IPCA (Random Forest with Incremental Principal Component Analysis). The following table lists the optimal hyperparameters determined after testing to ensure model stability and reliability.

[0112]

[0113] To accurately capture the complex mapping relationship between coal sample spectral characteristics and volatile components, the random forest algorithm was adopted as a machine learning model. As an ensemble learning method, random forest improves the model's prediction accuracy and robustness by constructing multiple decision trees.

[0114] The model training process involves independently training each decision tree using the raw data from the training set by randomly selecting features and samples. Each tree is generated on a different subset of samples and features, ensuring model diversity. During the construction of the random forest model, each decision tree is trained independently of the others, increasing the model's generalization ability. This allows the model to better learn from the data and make accurate predictions when faced with new, unseen data.

[0115] In a preferred embodiment, referring to Figure 4 , step S4 includes,

[0116] Step S41, constructing a second machine learning regression model based on the optimized hyperparameters;

[0117] Step S42, evaluating the second machine learning regression model through ten-fold cross validation according to the first evaluation index to obtain a learning curve and a residual graph;

[0118] Step S43, selecting the second machine learning regression model with the highest first evaluation index as the final prediction model, and performing volatile content prediction on the sample to be tested;

[0119] In step S42, the first evaluation index includes the coefficient of determination, the root mean square error, and the mean absolute error;

[0120] The final prediction model had a determination coefficient of 0.985, a root mean square error of 1.949, and a mean absolute error of 0.913.

[0121] Specifically, Figure 8(a)-Figure 11(b) The training and testing results of various random forest regression models are shown. In the figure, the gray dashed line represents the ideal curve y = x, and the color bar on the right represents the range of mean absolute error.

[0122] In addition to the RF-KPCA model training score R 2 In addition, the training results of RF-PCA, RF-SPCA and RF-IPCA models are very ideal, all showing scores exceeding 0.99.

[0123] From the test results, RF-SPCA still maintains a high R 2 The RF-SPCA regression model has an excellent prediction performance.

[0124] To ensure that the model avoids the risk of overfitting, the learning curve and residual graph of each regression model are plotted, as shown in Figure 12(a)-(d) and Figure 13(a)-(d). The learning curve is expressed in R 2As an evaluation indicator, the performance changes of the model under different training sample sizes are shown. It is observed that before the number of training samples reaches 40, the R 2 The scores rise rapidly, which reflects that the model needs enough data to learn the mapping between features and target attributes.

[0125] As the number of samples increases, the R 2 The scores are stabilizing, indicating that the model has captured the underlying patterns in the data. Careful observation reveals that the validation and training curves follow similar trends and gradually converge as training progresses. This phenomenon indicates that the model has been adequately trained and is learning effectively under the current dataset. The data points in both the training and test sets are evenly distributed around zero, indicating that the model's prediction errors at different prediction levels are random and that the model is effective.

[0126] Specifically, through multi-dimensional feature screening and dimensionality reduction optimization of principal component analysis (PCA) and its variant algorithms (KPCA, SPCA, IPCA), combined with the grid search method to systematically tune the hyperparameters of the random forest regression model (such as the maximum tree depth, the minimum number of leaf node samples, etc.), the problems of high dimensionality, noise interference and insufficient nonlinear relationship mining of spectral data were effectively solved, and the volatile matter prediction accuracy (R 2 ) reaches above 0.985, and the root mean square error (RMSE) and mean absolute error (MAE) are as low as 1.949 and 0.913 respectively, which is a leap forward compared with the existing technology (RMSE ≥ 5%).

[0127] The strong generalization ability of the model has been fully verified through ten-fold cross-validation, convergence of the learning curve and uniform distribution of residuals (the residual plot shows that the errors are randomly distributed on both sides of the zero point). It can stably adapt to the detection needs of different batches and multiple samples of coal samples, and avoid performance fluctuations caused by differences in data distribution.

[0128] In addition, this method has good scalability and versatility. By adjusting the feature screening algorithm and model parameters, it can be further applied to the high-precision prediction of other key quality parameters such as coal ash, sulfur content, and calorific value, providing efficient and reliable technical support for the intelligent and green development of the coal industry.

[0129] The present application also includes step S0, coal sample preparation: the coal sample is crushed by mechanical methods, and a pretreatment step is required for large or wet coal samples. The reduction operation adopts mechanical methods such as a divider to ensure that at least 20 sub-samples are taken for coal samples with a particle size of less than 13 mm to ensure representativeness. For coal samples with a particle size of less than 3 mm, after passing through a 3 mm sieve, at least 100 g is reduced for testing. Before crushing to a particle size of less than 0.2 mm, iron is removed first to avoid interfering with the analysis results. Finally, the coal sample is ensured to be air-dried and placed in a coal sample bottle that does not exceed 3 / 4 of the volume for subsequent use.

[0130] Coal Sample Volatile Matter Percentage Test: Take a certain amount of coal sample for routine analysis, place it in a sealed porcelain crucible, and heat it in an oxygen-free environment at 900±10°C for 7 minutes. The loss in mass of the coal sample, after deducting the moisture content, is used as the volatile matter content of the coal sample.

[0131] This invention achieves non-destructive and rapid detection of the volatile matter content of coal through the deep integration of terahertz time-domain spectroscopy technology and machine learning algorithms. The single-sample detection time is shortened to less than 5 minutes, which is significantly better than the traditional high-temperature heating method (which takes tens of minutes to several hours) and does not require sample destruction, greatly reducing detection costs and improving the real-time monitoring efficiency of coal production lines.

[0132] In summary, this application used terahertz time-domain spectroscopy to perform spectral analysis on 186 coal samples and determined the corresponding volatile content. During the data preprocessing phase, principal component analysis (PCA) and its variants, including kernel PCA, sequential PCA, and incremental PCA, were employed to effectively reduce the dimensionality and filter features of the spectral data. This not only significantly improved the clustering characteristics of the data but also verified the algorithm's applicability to coal sample analysis.

[0133] Furthermore, based on the optimized data, four random forest regression models were constructed: RF-PCA, RF-KPCA, RF-SPCA, and RF-IPCA to predict the volatile matter percentage of coal samples. Through careful hyperparameter tuning, the RF-SPCA model, in particular, showed excellent prediction performance, with R 2 The model achieved a mean error of 0.985, an RMSE of 1.949, and a MAE of 0.913, demonstrating its high accuracy and reliability. To ensure the model's generalization capabilities, a learning curve and residual plot were plotted. The convergence of the learning curve indicates that the model stabilizes with increasing training samples, while the residual plot shows that the prediction errors are evenly distributed on both sides of zero, further confirming the model's generalization performance.

[0134] The above description is only a preferred embodiment of the present invention and does not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for rapid prediction of coal volatile matter, characterized in that: include, Step S1, collecting spectral data of the sample to be tested by a transmission terahertz time-domain spectrometer; Step S2, performing standardization and dimensionality reduction on the spectral data in sequence, screening out principal component features related to volatiles, and forming a data set; Step S3: constructing a first machine learning regression model, training the first machine learning regression model using the data set, and optimizing hyperparameters of the first machine learning regression model using a grid search method; Step S4: construct a second machine learning regression model based on the optimized hyperparameters, calculate the learning curve and residual graph of the second machine learning regression model, and screen out the final prediction model based on the learning curve and the residual graph to predict the volatile matter of the sample to be tested.

2. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: The transmission terahertz time-domain spectrometer in step S1 includes a femtosecond laser, a terahertz wave generator, a delay device and a detector. The spectral data collection step includes: Step S11, emitting a first pulse laser by the femtosecond laser; Step S12, the first pulse laser is divided into a pump pulse and a probe pulse by a beam splitter; Step S13, the pump pulse acts on the terahertz wave generator to generate a terahertz wave; Step S14, the terahertz wave passes through the sample to be tested, the detection pulse is combined with the terahertz wave after the delay is adjusted by the delay device, and the terahertz pulse time domain signal is obtained by the detector; Step S15: The terahertz pulse time domain signal is transformed and calibrated to obtain the spectral data.

3. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: Step S2 uses one of principal component analysis, kernel principal component analysis, sequence principal component analysis and incremental principal component analysis to reduce the dimension of the spectral data.

4. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: Step S3 includes, Step S31, constructing the first machine learning regression model; Step S32: dividing the data set into a training set and a test set according to a preset ratio, using the training set to train the first machine learning regression model, and using the test set to test the first machine learning regression model; Step S33: using the grid search method to optimize the hyperparameters of the first machine learning regression model.

5. The method for rapid prediction of coal volatile matter according to claim 4, characterized in that: The preset ratio in step S32 is 4:1, and the hyperparameters of the first machine learning regression model in step S33 include the maximum tree depth, the minimum number of leaf node samples, and the minimum number of split samples; The maximum tree depth is 10-20, the minimum number of leaf node samples is 1-5, and the minimum number of split samples is 5-15.

6. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: The second machine learning regression model in step S4 includes principal component analysis random forest, kernel principal component analysis random forest, sequence principal component analysis random forest and incremental principal component analysis random forest.

7. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: Step S4 includes, Step S41, constructing the second machine learning regression model based on the optimized hyperparameters; Step S42, evaluating the second machine learning regression model through ten-fold cross validation according to the first evaluation index to obtain the learning curve and the residual graph; Step S43: Filter out the second machine learning regression model with the highest first evaluation index as the final prediction model to perform volatile content prediction on the sample to be tested.

8. The method for rapid prediction of coal volatile matter according to claim 7, characterized in that: In step S42, the first evaluation index includes the coefficient of determination, the root mean square error, and the mean absolute error.

9. The method for rapid prediction of coal volatile matter according to claim 8, characterized in that: The determination coefficient of the final prediction model is 0.985, the root mean square error is 1.949, and the mean absolute error is 0.

913.

10. The method for rapid prediction of coal volatile matter according to claim 1, characterized in that: The sample to be tested in step S1 and step S4 is coal.