Method for rapidly predicting coal quality parameters based on Raman spectrum in combination with machine learning

By systematically preprocessing and feature screening of Raman spectra, combined with machine learning models, the baseline drift and high-dimensional redundancy problems of Raman spectroscopy in coal quality parameter detection were solved, achieving rapid, non-destructive, and high-precision prediction of coal quality parameters.

CN121933495APending Publication Date: 2026-04-28HUAZHONG UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-01-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, Raman spectroscopy suffers from baseline drift, fluorescence background, and high-dimensional redundancy in coal quality parameter detection, leading to distorted feature extraction and high model complexity, making it difficult to achieve rapid, non-destructive, and high-precision prediction of coal quality parameters.

Method used

Baseline correction is performed using adaptive iterative reweighted least squares, combined with Savitzky-Golay filtering for noise reduction. Spectral data is then processed using max-min normalization or standard normalization. Subsequently, competitive adaptive reweighted sampling and continuous projection algorithms are used to reduce feature dimensionality, and machine learning models such as support vector machines and random forests are constructed for prediction.

Benefits of technology

It improves the spectral signal-to-noise ratio and data consistency, reduces model complexity, enhances the model's generalization ability and prediction accuracy, and achieves fast and lossless high-precision prediction of coal quality parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121933495A_ABST
    Figure CN121933495A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of coal quality detection, and particularly relates to a method for rapidly predicting coal quality parameters based on Raman spectrum in combination with machine learning. According to the method, a Raman spectrum technology and machine learning are combined, a Raman spectrum systematic preprocessing process is adopted, an adaptive iteration reweighted least square method Air-PLS is applied, spectral signals and baseline drift are accurately separated, Savitzky-Golay filtering is applied, peak shape features are reserved, random noise is effectively suppressed, and the method is suitable for large-scale popularization and application. Maximum-minimum normalization or Z-score standardization processing is provided, and the intensity difference between samples is eliminated; moreover, a data-driven Raman spectrum intelligent feature screening mechanism is introduced, and an optimal feature wavelength subset which is highly related to coal quality parameters and low in redundancy is screened out through automatic iteration; according to the method, the data quality in Raman spectrum coal quality detection is improved, the feature correlation is enhanced while the feature dimension is reduced, the model precision and generalization ability are improved, and rapid, lossless, high-precision and high-robustness coal quality parameter prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of coal quality testing technology, specifically relating to a method for rapidly predicting coal quality parameters based on Raman spectroscopy combined with machine learning. Background Technology

[0002] As an important basic energy source and industrial raw material in my country, coal's quality parameters (such as moisture, ash, volatile matter, and fixed carbon) directly determine its combustion efficiency, calorific value, and environmental pollution level, and have important guiding significance for coal mining, sorting, utilization, and trade settlement.

[0003] Traditional methods for detecting coal quality parameters mainly rely on chemical analysis methods specified in national standards, such as gravimetric determination of moisture and ash, and high-temperature combustion determination of volatile matter. While these methods provide accurate results, they suffer from drawbacks such as cumbersome operation, long testing cycles (typically requiring hours or even days), high sample destructiveness, and inability to achieve rapid batch testing, making them unsuitable for the real-time monitoring and efficient sorting requirements of modern coal industry production. Raman spectroscopy, as a rapid and non-destructive spectroscopic analysis technique, can reflect the molecular structure information of substances. The Raman spectral characteristics of coal are closely related to its chemical composition and molecular structure, and coal quality parameters are essentially a macroscopic manifestation of coal's chemical composition and structure; therefore, there is a strong coupling between Raman spectral data and coal quality parameters. On the other hand, machine learning technology has powerful data mining and nonlinear fitting capabilities, effectively handling high-dimensional, strongly coupled data. Therefore, combining Raman spectroscopy with machine learning to construct coal quality prediction models has attracted increasing attention.

[0004] Chinese patent CN119046684A proposes a multi-dimensional fusion coal quality detection method and system. This method integrates multiple information sources, including images, particle size analysis, laser-induced breakdown spectroscopy (LIBS), near-infrared spectroscopy, XRF, and Raman spectroscopy. It verifies the results by comparing and verifying the predictions with historical coal sample manual analysis results. A coal quality detection model based on neural networks, random forests, or support vector machines is constructed, and online iterative updates of the model are supported. Then, various coal quality characteristics of the coal sample to be tested are collected and input into the pre-trained coal quality detection model for coal quality detection, yielding the coal quality detection results. While this technical solution uses Raman spectra as one of the inputs, it treats them the same as other modal data, failing to design a dedicated preprocessing process for the unique problems of Raman spectra, such as baseline drift, fluorescence background, and high-dimensional redundancy. The raw spectra contain significant baseline drift and random noise, which can mask the true Raman peaks and lead to distorted feature extraction. On the other hand, although this technical solution uses a machine learning model, its feature input is raw or multimodal spliced ​​data. It does not perform targeted high-dimensional feature reduction or correlation screening on the Raman spectra, directly inputting the full spectrum or roughly segmented data into the model. Full spectrum modeling leads to high feature dimensionality, a lot of redundant information, and strong multicollinearity, increasing model complexity and making it prone to overfitting.

[0005] Chinese patent CN106198488B discloses a rapid coal quality detection method based on Raman spectroscopy. The core idea is to collect Raman spectra of various standard coal samples, and establish mapping relationships between these features and coal quality parameters such as moisture, ash, volatile matter, and fixed carbon by artificially pre-setting specific peaks (e.g., 2670 cm⁻¹, 2810 cm⁻¹, etc.) and their combination ratios, forming a static correlation database. For the coal sample to be tested, the same features are extracted, and coal quality is predicted by comparison with the database. This technical solution performs some preprocessing of the Raman spectra, but only mentions "segmentation" and "baseline removal," without addressing key steps such as noise smoothing and data normalization. Furthermore, the large differences in spectral intensity between different samples mean that without normalization, scale bias will be introduced, affecting the model's generalization ability. On the other hand, in terms of feature selection, this technical solution relies entirely on manually preset fixed peaks and their ratios, which is an experience-driven feature engineering. It cannot adapt to the spectral variation patterns of coal samples of different coal ranks and metamorphic degrees, and may miss key but atypical information. It has weak generalization ability, and the prediction accuracy may drop sharply when facing new coal types or complex mixed coals. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, the technical problem to be solved by this invention is to provide a method for rapidly predicting coal quality parameters based on Raman spectroscopy combined with machine learning. By combining Raman spectroscopy with machine learning, and through systematic spectral preprocessing and automatic feature screening of the raw Raman spectral data, the data quality in Raman spectroscopy coal quality detection is improved, the feature dimensionality is reduced while the feature correlation is enhanced, thereby improving the model accuracy and generalization ability, and achieving rapid, non-destructive, high-precision, and highly robust prediction of coal quality parameters.

[0007] The technical solution adopted in this invention is as follows:

[0008] A method for rapidly predicting coal quality parameters based on Raman spectroscopy combined with machine learning includes the following steps:

[0009] Step S1: Obtain the coal quality parameters and Raman spectrum of the coal sample to obtain the true values ​​of the coal quality parameters and the original Raman spectrum data;

[0010] Preferably, step S1 includes:

[0011] S1.1 Obtain the true values ​​of coal quality parameters from coal samples: Select coal samples of different coal types and different degrees of metamorphism, crush, grind, and sieve the coal samples to a particle size of no more than 75 μm, and use national standard methods to measure the coal quality parameters of the coal samples, including moisture, ash, volatile matter, and fixed carbon. Each coal sample is tested three times, and the average value is taken as the true value of the coal quality parameters to establish a coal quality parameter dataset; the national standard methods are specifically GB / T 211-2017 (determination of moisture) and GB / T 212-2008 (determination of ash, volatile matter, and fixed carbon).

[0012] S1.2, Obtain the raw Raman spectral data of the coal sample: Spread the coal sample processed in step S1.1 evenly on the sample stage of the Raman spectrometer, ensuring that the sample covers the laser spot and is free from impurities; set the acquisition parameters of the Raman spectrometer, start the spectral acquisition, obtain the raw Raman spectral data of each coal sample, and establish the raw Raman spectral dataset.

[0013] Preferably, in step S1.2, the acquisition parameters of the Raman spectrometer are set as follows: the excitation wavelength is 532nm or 785nm (selected according to the characteristics of the coal sample to avoid fluorescence interference), the spectral scanning range is 100~4000cm⁻¹, the scanning time is 10~30s, the number of scans is 3~5, and the laser power is 50~200mW (to avoid excessive power causing oxidation of the coal sample).

[0014] Step S2: Baseline correction, spectral noise removal and data normalization are performed on the raw Raman spectral data obtained in step S1 to obtain preprocessed spectral data.

[0015] Preferably, step S2 includes:

[0016] S2.1, Baseline Correction: The original Raman spectrum is corrected using the adaptive iterative reweighted least squares method Air-PLS. By iteratively optimizing the weighting coefficients, the spectral signal and baseline drift are accurately separated to obtain the baseline-corrected spectral data.

[0017] S2.2, Spectral noise removal: The Savitzky-Golay smoothing method is used to smooth the baseline-corrected spectral data obtained in S2.1. The smoothing window width is set to 5 to 15 points (adjusted according to the intensity of spectral noise, with the window width being an odd number). Random noise is eliminated by polynomial fitting, while spectral feature information is preserved, resulting in smoothed spectral data.

[0018] S2.3, Data Normalization Processing: The smoothed spectral data obtained in S2.2 is processed using max-min normalization or standard normalization to obtain preprocessed spectral data.

[0019] If maximum-minimum normalization is used to map the spectral data to the range [0,1], the calculation formula is as follows:

[0020] x'=(x-x_min) / (x_max-x_min),

[0021] Where x is the original spectral data, x_min is the minimum value of the spectral data, x_max is the maximum value of the spectral data, and x' is the normalized data;

[0022] If standard normalization is used to convert the spectral data into standard normal distribution data with a mean of 0 and a variance of 1, the calculation formula is as follows:

[0023] x'=(x-μ) / σ,

[0024] Where μ is the mean of the spectral data, σ is the standard deviation of the spectral data, and x' is the normalized data.

[0025] Step S3: Select characteristic wavelengths from the preprocessed spectral data obtained in step S2 to obtain a characteristic spectral dataset;

[0026] Preferably, in step S3, the preprocessed spectral data obtained in step S2 is subjected to feature screening by the feature wavelength selection module. Redundant information and irrelevant wavelengths are removed, and feature wavelengths with strong correlation to coal quality parameters are extracted. This reduces model complexity and improves model generalization ability. The specific process is as follows:

[0027] S3.1 Feature Selection Algorithm: One or more combinations of the following algorithms are selected as feature selection algorithms: Competitive Adaptive Reweighted Sampling (CARS), Continuous Projection Algorithm (SPA), Random Forest Feature Importance Ranking, Genetic Algorithm (GA), or Recursive Feature Elimination (RFE). Among them, Competitive Adaptive Reweighted Sampling (CARS) and Continuous Projection Algorithm are used to eliminate multicollinearity in spectral data, Random Forest Feature Importance Ranking is used to evaluate the contribution of each wavelength to coal quality parameters, Genetic Algorithm is used for global optimization to select the optimal feature subset, and Recursive Feature Elimination is used to gradually eliminate redundant features.

[0028] S3.2, Feature Wavelength Extraction: The preprocessed spectral data obtained in step S2 is correlated with the true values ​​of coal quality parameters obtained in step S1. The correlation coefficient (which measures the linear correlation between wavelength and coal quality parameters) and root mean square error (which measures the prediction error of the feature subset) are used as evaluation indicators. The feature selection algorithm selected in S3.1 is used for iterative selection to retain the feature wavelengths and corresponding spectral data with an absolute correlation coefficient ≥ 0.9 and the smallest root mean square error, thus obtaining the feature spectral dataset.

[0029] Step S4: Based on the feature spectral dataset obtained in step S3 and the true values ​​of coal quality parameters obtained in step S1, construct a machine learning prediction model and train it to obtain the optimal prediction model.

[0030] Preferably, step S4 specifically includes the following steps:

[0031] S4.1, Dataset Partitioning: The feature spectral dataset and the corresponding true values ​​of coal quality parameters are combined to form the modeling dataset, and the training set and test set are randomly divided in a 7:3 ratio; the training set is used for model training, and the test set is used for model performance verification.

[0032] S4.2, Model Selection and Construction: Select one of the following as the machine learning prediction model: Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Artificial Neural Network (ANN), or Convolutional Neural Network (CNN). Build the model based on Python's Scikit-learn library or TensorFlow framework.

[0033] S4.3, Model Hyperparameter Optimization: Hyperparameters of the model are optimized using grid search or Bayesian optimization. Different hyperparameter search ranges are set for different models: For Support Vector Machines, hyperparameters include the penalty coefficient C (search range 1–100) and kernel parameter γ (search range 0.001–0.1); for Random Forests, hyperparameters include the number of decision trees (search range 100–1000) and maximum tree depth (search range 5–50); for Gradient Boosting Decision Trees, hyperparameters include the learning rate (search range 0.01–0.3), the number of decision trees (search range 100–1000), and the maximum tree depth (search range 3–20); for Artificial Neural Networks, hyperparameters include the number of hidden layer neurons (search range 10–100), the learning rate (search range 0.001–0.1), and the number of iterations (search range 100–1000).

[0034] S4.4, Model Training and Validation: Train the optimized model using the training set, input the test set into the trained model, and obtain the predicted values ​​of coal quality parameters; evaluate the model performance using the coefficient of determination R², root mean square error RMSE, and mean absolute error MAE; when the model's R² ≥ 0.95 and RMSE ≤ 0.5, training can be stopped and the model can be determined as the optimal prediction model; if the model performance does not meet the requirements, return to step S3 to adjust the feature selection algorithm or return to steps S4.2 and S4.3 to adjust the model type and hyperparameters, and remodel until the performance requirements are met.

[0035] Step S5: Based on the optimal prediction model obtained in step S4, predict the coal quality parameters of the coal sample to be tested.

[0036] Preferably, step S5 includes:

[0037] S5.1, Coal sample processing: Select the coal sample to be tested, and process it by crushing, grinding, sieving and collecting the original Raman spectrum data of the coal sample according to the method in step S1;

[0038] S5.2, Preprocessing of the spectrum to be measured: The original Raman spectrum data of the coal sample to be predicted is subjected to baseline correction, spectral noise removal and normalization according to the method in step S2 to obtain the preprocessed spectrum data to be predicted.

[0039] S5.3, Feature wavelength extraction: According to the feature filtering algorithm in step S3, extract the corresponding feature spectral data from the preprocessed spectral data to be predicted;

[0040] S5.4, Coal quality parameter prediction: Input the extracted characteristic spectral data to be predicted into the optimal prediction model obtained in step S4, and quickly output the coal quality parameter prediction results of the coal sample to be predicted, thus completing the rapid prediction of coal quality parameters.

[0041] The beneficial effects obtained by adopting the above technical solution are as follows:

[0042] (1) This invention constructs a three-stage systematic preprocessing process for raw Raman spectra: the adaptive iterative reweighted least squares method (Air-PLS) is used to accurately separate the spectral signal from the baseline drift; Savitzky-Golay filtering is applied to effectively suppress random noise while preserving peak characteristics; and the maximum-minimum normalization or Z-score standardization options are provided to eliminate intensity differences between samples. This systematic three-stage spectral preprocessing process significantly improves the spectral signal-to-noise ratio and consistency, provides high-quality and comparable input data for subsequent modeling, enhances the stability and transferability of the model under different instruments and different batches of coal samples, and improves the model's generalization ability.

[0043] (2) This invention introduces a data-driven intelligent feature selection mechanism for Raman spectroscopy, which integrates a variety of advanced algorithms (CARS, SPA, random forest importance, GA, RFE) to automatically iteratively select the optimal feature wavelength subset that is highly correlated with coal quality parameters and has low redundancy. This significantly reduces model complexity, improves model generalization ability and prediction accuracy, and avoids human experience bias. It makes feature selection adaptive to coal sample characteristics and is suitable for multi-coal scenarios. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating a method for rapidly predicting coal quality parameters based on Raman spectroscopy combined with machine learning, according to the present invention.

[0045] Figure 2 Box plot showing the actual values ​​of coal quality parameters for 64 types of coal.

[0046] Figure 3 The original Raman spectra of 64 types of coal are shown.

[0047] Figure 4 Example of spectral feature wavelength extraction using the Competitive Adaptive Reweighted Sampling (CARS) method.

[0048] Figure 5 This is an example of spectral feature wavelength extraction using the Continuous Projection Algorithm (SPA).

[0049] Figure 6 To improve the performance index (R) of coal quality parameter content prediction using a partial least squares (PLS) model that combines feature wavelength extraction algorithm and random forest (RF) model. 2 and R MSE ).

[0050] Figure 7 The graph shows the results of predicting coal quality parameter content using the optimal model, Random Forest (RF). Detailed Implementation

[0051] The technical solution of the present invention will now be described more clearly and completely with reference to the accompanying drawings.

[0052] like Figure 1 As shown, the implementation of the method for predicting coal quality parameters by combining Raman spectroscopy with machine learning includes the following steps:

[0053] 1. Data Acquisition

[0054] (1) Coal quality parameter acquisition: 64 different coal types (bituminous coal, anthracite, lignite) and different degrees of metamorphism were selected. Each coal sample was crushed and passed through a 75μm standard sieve. The coal quality parameters (moisture, ash, volatile matter, and fixed carbon) of the coal samples were measured using national standard methods. Each coal sample was tested three times, and the average value was taken as the true value of the coal quality parameters to establish a coal quality parameter dataset. The national standard methods are GB / T 211-2017 (moisture determination) and GB / T 212-2008 (ash, volatile matter, and fixed carbon determination). The distribution of the true values ​​of the 64 coal quality parameters is as follows: Figure 2 As shown;

[0055] (2) Raman spectroscopy acquisition: The sieved coal samples were evenly spread on the sample stage of the Raman spectrometer (model: ThermoScientific DXR3xi). The acquisition parameters were set as follows: excitation wavelength 532nm, spectral scanning range 100~4000cm⁻¹, scanning time 20s, number of scans 3, laser power 100mW; the original Raman spectral data of each coal sample were acquired to establish the original Raman spectral dataset. The Raman spectra of 64 coal samples are as follows: Figure 3 As shown.

[0056] 2. Spectral data preprocessing

[0057] (1) Baseline correction: The original Raman spectrum was baseline corrected using the adaptive iterative reweighted least squares method (Air-PLS). The number of iterations was set to 100, and the initial value of the weight coefficient was set to 0.01 to obtain the baseline corrected spectral data.

[0058] (2) Spectral noise removal: The Savitzky-Golay smoothing method was used to smooth the baseline-corrected spectral data. The smoothing window width was set to 10 points and the polynomial order was set to 2 to obtain the smoothed spectral data.

[0059] (3) Normalization: The smoothed spectral data is processed by standard normalization. The mean μ and standard deviation σ of the spectral data at each wavelength are calculated and converted according to the formula x'=(x-μ) / σ to obtain the preprocessed spectral data.

[0060] 3. Selection of characteristic wavelengths

[0061] Competitive Adaptive Reweighted Sampling (CARS) and Continuous Projection Algorithm (SPA) were selected as feature selection algorithms. The preprocessed spectral data was correlated with the true gray values. The search range for the number of feature wavelengths was set to 10–50. Iterative selection was performed with the goal of minimizing the root mean square error, ultimately selecting the feature wavelengths and their corresponding spectral data to obtain the feature spectral dataset. A schematic diagram of CARS and SPA feature wavelength selection is shown below. Figure 4 and Figure 5 As shown.

[0062] 4. Build machine learning prediction models

[0063] (1) Data set division: The feature spectrum dataset and the corresponding coal quality parameter real values ​​are combined to form the modeling dataset, and the training set (45 samples) and the test set (19 samples) are randomly divided in a 7:3 ratio.

[0064] (2) Model selection and construction: Partial least squares (PLS) and random forest (RF) models were selected as machine learning prediction models.

[0065] (3) Model parameter optimization: The PLS model uses cross-validation (5-fold cross-validation) to optimize the number of principal components. The search range for the number of principal components is set to 1-15. With the goal of minimizing the root mean square error of cross-validation, the optimal number of principal components is determined to be 8. The RF model uses grid search to optimize hyperparameters. The optimal hyperparameters are 800 decision trees and a maximum tree depth of 30.

[0066] The grid search method was used to optimize the hyperparameters of the SVM. The search range of the penalty coefficient C was 1 to 50, with a step size of 5; the search range of the kernel function parameter γ was 0.001 to 0.05, with a step size of 0.005; the cross-validation fold was set to 5, and the optimal hyperparameters were finally determined to be C=30 and γ=0.02.

[0067] (4) Model training and validation: The optimized RF model is trained using the training set, and the test set is input into the trained model to obtain the predicted values ​​of coal quality parameters; the model performance indices R² and R₂ are calculated. MSE .

[0068] 5. Obtain coal quality parameter information based on the constructed model.

[0069] (1) Processing of coal samples to be tested: Select 20 coal samples to be tested, crush and screen them according to the method in step 1, and collect the original Raman spectral data;

[0070] (2) Preprocessing of the spectra to be predicted: Baseline correction, spectral noise removal and standard normalization are performed on the original Raman spectral data according to the method in step 2;

[0071] (3) Feature wavelength extraction: According to the CARS and SPA algorithms in step 3, extract the spectral data corresponding to the feature wavelength from the preprocessed spectral data to be tested;

[0072] (4) Coal quality parameter prediction: The characteristic spectral data are input into the optimal RF prediction model, and the model outputs the predicted values ​​of coal quality parameters (moisture, ash, volatile matter, and fixed carbon) of 20 coal samples within 10 seconds; the specific results are as follows. Figure 7 As shown. Comparing the predicted values ​​with the measured values ​​using the traditional gravimetric method, the average relative error is within an acceptable range, and the prediction accuracy meets the actual detection requirements. The model performs best when predicting fixed carbon content, with an R² value of 0.9811 and an RMSE of 0.9573; when predicting moisture content, the R² value is 0.9657. MSE The value is 1.6854. When predicting ash content, R² = 0.9715. MSE =1.7514; R²=0.9562, RMSE=2.1035 when predicting volatile matter content.

Claims

1. A method for rapidly predicting coal quality parameters based on Raman spectroscopy combined with machine learning, characterized in that, Includes the following steps: Step S1: Obtain the coal quality parameters and Raman spectrum of the coal sample to obtain the true values ​​of the coal quality parameters and the original Raman spectrum data; Step S2: Baseline correction, spectral noise removal and data normalization are performed on the raw Raman spectral data obtained in step S1 to obtain preprocessed spectral data. Step S3: Select characteristic wavelengths from the preprocessed spectral data obtained in step S2 to obtain a characteristic spectral dataset; Step S4: Based on the feature spectral dataset obtained in step S3 and the true values ​​of coal quality parameters obtained in step S1, construct a machine learning prediction model and train it to obtain the optimal prediction model. Step S5: Based on the optimal prediction model obtained in step S4, predict the coal quality parameters of the coal sample to be tested.

2. The method for rapidly predicting coal quality parameters according to claim 1, characterized in that, Step S1 includes: S1.1 Obtain the true values ​​of coal quality parameters of coal samples: Select coal samples of different coal types and different degrees of metamorphism, crush, grind and sieve the coal samples to a particle size of no more than 75μm, and measure the coal quality parameters of the coal samples, including moisture, ash, volatile matter and fixed carbon. Each coal sample is tested three times, and the average value is taken as the true value of coal quality parameters to establish a coal quality parameter dataset. S1.2, Obtain the raw Raman spectral data of the coal sample: Spread the coal sample processed in step S1.1 evenly on the sample stage of the Raman spectrometer, set the acquisition parameters of the Raman spectrometer, start the spectral acquisition, obtain the raw Raman spectral data of each coal sample, and establish the raw Raman spectral dataset.

3. The method for rapidly predicting coal quality parameters according to claim 2, characterized in that, In step S1.2, the acquisition parameters of the Raman spectrometer are set as follows: the excitation wavelength is 532nm or 785nm, the spectral scanning range is 100~4000cm⁻¹, the scanning time is 10~30s, the number of scans is 3~5, and the laser power is 50~200mW.

4. The method for rapidly predicting coal quality parameters according to claim 3, characterized in that, Step S2 includes: S2.1, Baseline Correction: The original Raman spectrum is corrected using the adaptive iterative reweighted least squares method Air-PLS. By iteratively optimizing the weighting coefficients, the spectral signal and baseline drift are accurately separated to obtain the baseline-corrected spectral data. S2.2, Spectral noise removal: The Savitzky-Golay smoothing method is used to smooth the baseline-corrected spectral data obtained in S2.

1. The smoothing window width is set to 5 to 15 points. Random noise is eliminated by polynomial fitting, while spectral feature information is preserved, and smoothed spectral data is obtained. S2.3, Data Normalization Processing: The smoothed spectral data obtained in S2.2 is processed using max-min normalization or standard normalization to obtain preprocessed spectral data. If maximum-minimum normalization is used to map the spectral data to the range [0,1], the calculation formula is as follows: x'=(x-x_min) / (x_max-x_min), Where x is the original spectral data, x_min is the minimum value of the spectral data, x_max is the maximum value of the spectral data, and x' is the normalized data; If standard normalization is used to convert the spectral data into standard normal distribution data with a mean of 0 and a variance of 1, the calculation formula is as follows: x'=(x-μ) / σ, Where μ is the mean of the spectral data, σ is the standard deviation of the spectral data, and x' is the normalized data.

5. The method for rapidly predicting coal quality parameters according to claim 4, characterized in that, In step S3, the preprocessed spectral data obtained in step S2 is subjected to feature screening to remove redundant information and irrelevant wavelengths, and feature wavelengths with strong correlation to coal quality parameters are extracted. The specific process is as follows: S3.1 Feature selection algorithm: Select one or more of the following as the feature selection algorithm: Competitive Adaptive Reweighted Sampling (CARS), Continuous Projection Algorithm (SPA), Random Forest Feature Importance Ranking, Genetic Algorithm (GA), or Recursive Feature Emission (RFE). S3.2, Feature Wavelength Extraction: The preprocessed spectral data obtained in step S2 is correlated with the true values ​​of coal quality parameters obtained in step S1. Using the correlation coefficient and root mean square error as evaluation indicators, the feature selection algorithm selected in S3.1 is used for iterative screening to retain the feature wavelengths and their corresponding spectral data with an absolute correlation coefficient ≥ 0.9 and the smallest root mean square error. The characteristic spectral dataset is obtained.

6. The method for rapidly predicting coal quality parameters according to claim 5, characterized in that, In step S4, the specific steps are as follows: S4.1, Dataset partitioning: The feature spectral dataset and the corresponding true values ​​of coal quality parameters are combined to form the modeling dataset, and the training set and test set are randomly divided in a 7:3 ratio; S4.2, Model Selection and Construction: Select one of the following as the machine learning prediction model: Support Vector Machine (SVM), Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Artificial Neural Network (ANN), or Convolutional Neural Network (CNN). Build the model based on Python's Scikit-learn library or TensorFlow framework. S4.3, Model Hyperparameter Optimization: Hyperparameters of the model are optimized using grid search or Bayesian optimization. Different hyperparameter search ranges are set for different models: For Support Vector Machines, hyperparameters include the penalty coefficient C (search range 1–100) and kernel parameter γ (search range 0.001–0.1); for Random Forests, hyperparameters include the number of decision trees (search range 100–1000) and maximum tree depth (search range 5–50); for Gradient Boosting Decision Trees, hyperparameters include the learning rate (search range 0.01–0.3), the number of decision trees (search range 100–1000), and the maximum tree depth (search range 3–20); for Artificial Neural Networks, hyperparameters include the number of hidden layer neurons (search range 10–100), the learning rate (search range 0.001–0.1), and the number of iterations (search range 100–1000). S4.4, Model Training and Validation: Train the optimized model using the training set, input the test set into the trained model, and obtain the predicted values ​​of coal quality parameters; evaluate the model performance using the coefficient of determination R², root mean square error RMSE, and mean absolute error MAE; when the model's R² ≥ 0.95 and RMSE ≤ 0.5, training can be stopped and the model can be determined as the optimal prediction model; if the model performance does not meet the requirements, return to step S3 to adjust the feature selection algorithm or return to steps S4.2 and S4.3 to adjust the model type and hyperparameters, and remodel until the performance requirements are met.

7. The method for rapidly predicting coal quality parameters according to claim 6, characterized in that, Step S5 includes: S5.1, Coal sample processing: Select the coal sample to be tested, and process it by crushing, grinding, sieving and collecting the original Raman spectrum data of the coal sample according to the method in step S1; S5.2, Preprocessing of the spectrum to be measured: According to the method in step S2, the original Raman spectrum data of the coal sample to be predicted is subjected to baseline correction, spectral noise removal and normalization to obtain the preprocessed spectrum data to be predicted. S5.3, Feature wavelength extraction: According to the feature filtering algorithm in step S3, extract the corresponding feature spectral data from the preprocessed spectral data to be predicted; S5.4, Coal quality parameter prediction: Input the extracted characteristic spectral data to be predicted into the optimal prediction model obtained in step S4, and quickly output the coal quality parameter prediction results of the coal sample to be predicted, thus completing the rapid prediction of coal quality parameters.

Citation Information

Patent Citations

  • A rapid coal quality detection method based on Raman spectroscopy analysis

    CN106198488B

  • Multi-dimensional fusion coal quality detection method and system

    CN119046684A

Cited By

  • Pipe network sludge component detection method, device, medium and equipment

    CN121558669A