Coal quality parameter prediction method and device, electronic equipment and storage medium

By using LIBS technology and GA-optimized XGBoost model, rapid detection of coal quality parameters G and Y values ​​was achieved, solving the problems of time-consuming and cumbersome traditional detection methods, providing real-time data support, and improving detection efficiency.

CN121922237APending Publication Date: 2026-04-24HEFEI GOLD STAR INTELLIGENT CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI GOLD STAR INTELLIGENT CONTROL TECH CO LTD
Filing Date
2025-12-16
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, the detection of coal quality parameters G and Y values ​​relies on traditional laboratory analysis methods, which are cumbersome and time-consuming, and cannot meet the needs of modern coal industry for rapid and real-time detection.

Method used

Laser-induced breakdown spectroscopy (LIBS) technology combined with genetic algorithm (GA) optimization of the XGBoost model is used to replace the traditional detection process through spectral acquisition and characteristic wavelength vector analysis, enabling rapid prediction of coal quality parameters.

Benefits of technology

The ability to complete coal quality parameter testing in a short time improves testing efficiency, provides real-time data support for industrial production, and meets the rapid feedback needs of the modern coal industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922237A_ABST
    Figure CN121922237A_ABST
Patent Text Reader

Abstract

The invention discloses a coal quality parameter prediction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a plurality of sampling spectrums obtained by carrying out the spectrum collection of a coal sample through an LIBS detector, and calculating the proportion of effective spectrums in the plurality of sampling spectrums; in response to the condition that the proportion is greater than a preset proportion threshold value, extracting a characteristic wavelength vector from the effective spectrum; and inputting the characteristic wavelength vector into a pre-constructed GA-XGBoost prediction model, and predicting to obtain the coal quality parameters of the coal sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of coal quality testing technology, specifically relating to a method for predicting coal quality parameters, a device for predicting coal quality parameters, an electronic device, and a computer-readable storage medium. Background Technology

[0002] In the field of coal quality analysis, the caking index (G value) and the maximum thickness of the plastic layer (Y value) of coking coal are the core indicators for evaluating its coking process performance and gasification reactivity, and are of decisive significance for guiding the efficient conversion and utilization of coal.

[0003] Currently, the detection of G-values ​​and Y-values ​​mainly relies on traditional offline laboratory analysis methods. Y-value determination typically uses a caking index analyzer, requiring standardized heating of the coal sample and real-time observation and recording of dynamic changes in the caking mass thickness by operators. This process is highly dependent on manual labor, prone to subjective errors, and inefficient. G-value determination requires mixing the coal sample with specialized anthracite in a precise ratio, heating it according to a specific procedure, and then performing a series of tedious sieving and weighing operations to calculate its caking index. These traditional methods, starting with the collection of representative samples, involve complex pretreatment steps such as crushing and reduction, followed by formal determination relying on manual labor and specialized equipment. The entire process is extremely cumbersome and time-consuming, typically requiring hours or even days to obtain the final result. This significant inefficiency is incompatible with the fast-paced, intelligent, and continuous production model pursued by the modern coal industry, and cannot meet the urgent need for rapid feedback and real-time guidance of coal quality data during production. It has become a technical bottleneck restricting the refined processing and efficient utilization of coal. Therefore, there is an urgent need to provide an efficient, accurate, and automated method for predicting coal quality parameters. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method, device, electronic device, and storage medium for predicting coal quality parameters. By introducing laser-induced breakdown spectroscopy (LIBS) technology and using a genetic algorithm (GA) to optimize the hyperparameters of the XGBoost model, it replaces the traditional cumbersome offline laboratory analysis process, enabling spectral acquisition and analysis to be completed in a short time, greatly improving the detection efficiency of coal quality parameters and providing real-time data support for industrial production.

[0005] In a first aspect, embodiments of this application provide a method for predicting coal quality parameters, including: Multiple sampling spectra obtained by spectral acquisition of coal samples using a LIBS detector are acquired, and the proportion of effective spectra in the multiple sampling spectra is calculated; In response to a proportion exceeding a preset proportion threshold, feature wavelength vectors are extracted from the effective spectrum; The characteristic wavelength vector is input into the pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

[0006] In some embodiments, calculating the percentage of valid spectra among multiple sampled spectra includes: Extract the maximum intensity of each spectrum; The sampled spectrum whose maximum intensity is within the preset intensity range is taken as the effective spectrum; Calculate the ratio of the number of effective spectra to the number of sampled spectra to obtain the proportion of effective spectra among multiple sampled spectra.

[0007] In some embodiments, extracting a characteristic wavelength vector from the effective spectrum includes: For each valid spectrum, the first spectral intensity vector of the valid spectrum is obtained, and the maximum-minimum value normalization process is performed on each spectral intensity in the first spectral intensity vector to obtain the second spectral intensity vector. The SHAP value is evaluated based on the SHAP value analysis method for each spectral intensity in the second spectral intensity vector; where the SHAP value represents the weight of the characteristic wavelength vector corresponding to the spectral intensity on the predicted coal quality parameters. Sort each spectral intensity in the second spectral intensity vector from highest to lowest according to the SHAP value, and select the target spectral intensity whose SHAP value is higher than a preset threshold from the sorted list. Obtain the characteristic wavelength vector corresponding to the target spectral intensity.

[0008] In some embodiments, a first spectral intensity vector of the effective spectrum is obtained, and a second spectral intensity vector is obtained by performing maximum-minimum normalization on each spectral intensity in the first spectral intensity vector, including: Determine the maximum and minimum spectral intensities in the first spectral intensity vector; Calculate the difference between the maximum and minimum spectral intensities; The second spectral intensity vector is obtained by calculating the ratio of the difference between each spectral intensity and the minimum spectral intensity in the first spectral intensity vector to the difference value.

[0009] In some embodiments, the method further includes: An XGBoost prediction model is constructed, and multiple preset hyperparameters of the XGBoost prediction model are binary encoded to generate the initial population of the genetic algorithm; among them, the hyperparameters include at least the learning rate, regularization coefficient and the maximum tree depth. The spectral data after feature screening is input into the XGBoost prediction model for training. The average absolute error between the predicted and actual values ​​of coal quality parameters is used as the fitness function to calculate the fitness value corresponding to the chromosome of each individual in the initial population. Based on the fitness value, perform selection, crossover, and mutation operations on the current initial population to generate a new generation of population; In response to the optimal individual in the new generation population having a fitness value greater than the preset fitness threshold or reaching the maximum number of iterations, the corresponding hyperparameters are taken as the optimal solution, and the GA-XGBoost prediction model is trained based on the optimal solution.

[0010] In some embodiments, the feature-filtered spectral data includes a feature wavelength vector.

[0011] In some embodiments, coal quality parameters include the caking index and the maximum thickness of the plastic layer.

[0012] Secondly, embodiments of this application provide a device for predicting coal quality parameters, comprising: The acquisition module is configured to acquire multiple sampling spectra obtained by spectral acquisition of coal samples through a LIBS detector, and calculate the proportion of effective spectra among the multiple sampling spectra; The response module is configured to extract a feature wavelength vector from the effective spectrum in response to a proportion greater than a preset proportion threshold. The prediction module is configured to input the characteristic wavelength vector into a pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

[0013] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the coal quality parameter prediction method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the coal quality parameter prediction method as described in the first aspect.

[0015] The method, apparatus, electronic equipment, and storage medium for predicting coal quality parameters provided in this application first acquire multiple sampled spectra of a coal sample using a LIBS detector, and calculate the proportion of effective spectra among these spectra. Further, in response to a proportion exceeding a preset threshold, a feature wavelength vector is extracted from the effective spectra. Finally, the feature wavelength vector is input into a pre-constructed GA-XGBoost prediction model to predict the coal quality parameters of the coal sample. This application, by introducing laser-induced breakdown spectroscopy (LIBS) technology and employing a genetic algorithm (GA) to optimize the hyperparameters of the XGBoost model, replaces the traditional, cumbersome offline laboratory analysis process. It enables rapid spectral acquisition and analysis, significantly improving the efficiency of coal quality parameter detection and providing real-time data support for industrial production.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 A flowchart illustrating a method for predicting coal quality parameters provided in this application embodiment; Figure 2 This is a schematic diagram of the GA-XGBoost prediction model construction process provided in the embodiments of this application; Figure 3 A schematic diagram illustrating the overall process of constructing the GA-XGBoost prediction model provided in this application embodiment; Figure 4 A schematic diagram of a coal quality parameter prediction device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0018] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0019] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0020] As described in the background section, the G value represents the caking index, an indicator of the caking property of bituminous coal, referring to the coal's ability to bind itself or added inert substances during dry distillation; the Y value represents the maximum thickness of the plastic layer, which is the maximum value of the difference between the upper and lower layers of the plastic body measured by a probe in the determination of the plastic layer index of bituminous coal. Both are key coal quality parameters for evaluating the quality of prime coking coal.

[0021] In the field of coal quality analysis, the G-value and Y-value, as core quality indicators of coking coal, directly determine the performance of coal in industrial processes such as coking and gasification. However, the current laboratory testing methods have significant limitations and are no longer suitable for the modern coal industry's demand for efficient and real-time testing. The detection of G-value and Y-value mainly relies on traditional laboratory testing methods. Taking Y-value detection as an example, the commonly used plastic layer analyzer requires strict adherence to specific specifications for heating the coal sample. During heating, operators must closely observe and accurately record the thickness changes of the plastic layer formed by the coal sample to determine the Y-value. The determination of the G-value is equally complex, requiring the coal sample to be mixed with special anthracite in a specific ratio, heated under specified conditions, and then subjected to a series of precise operations such as sieving to finally calculate the caking index. However, these traditional testing methods are extremely cumbersome and time-consuming. From the collection of coal samples, strict sampling standards must be followed to ensure the samples are representative. Complex pretreatment steps, including crushing and reduction, are required before the formal experimental testing stage begins. The entire process often takes several hours or even days to complete, which seriously affects the testing efficiency of the modern, fast-paced coal industry and fails to meet the urgent need to obtain real-time coal quality information to guide production.

[0022] Based on this, this application proposes a method, apparatus, electronic device, and storage medium for predicting coal quality parameters. First, multiple sampled spectra are acquired from a coal sample using a LIBS detector, and the proportion of effective spectra among these spectra is calculated. Further, in response to a proportion exceeding a preset threshold, a feature wavelength vector is extracted from the effective spectra. Finally, the feature wavelength vector is input into a pre-constructed GA-XGBoost prediction model to predict the coal quality parameters of the coal sample. This application, by introducing laser-induced breakdown spectroscopy (LIBS) technology and employing a genetic algorithm (GA) to optimize the hyperparameters of the XGBoost model, replaces the traditional, cumbersome offline laboratory analysis process. It enables spectral acquisition and analysis to be completed in a short time, significantly improving the efficiency of coal quality parameter detection and providing real-time data support for industrial production.

[0023] refer to Figure 1 This is a flowchart illustrating a method for predicting coal quality parameters provided in an embodiment of this application.

[0024] Step S101: Obtain multiple sampling spectra obtained by spectral acquisition of coal samples using a LIBS detector, and calculate the proportion of effective spectra among the multiple sampling spectra.

[0025] Specifically, in this step, the same coal sample is struck multiple times (M times) using a LIBS analyzer, collecting M spectral data points. The maximum intensity value of each spectrum is extracted to form a maximum value dataset. Furthermore, these maximum values ​​fall within a preset intensity range. The effective spectrum (i.e., the number of effective spectra) is counted. Calculate the proportion of the effective spectrum. , .

[0026] Step S102: In response to the proportion being greater than a preset proportion threshold, extract the feature wavelength vector from the effective spectrum.

[0027] Specifically, if the proportion of effective spectrum is greater than the preset proportion threshold, a small number of characteristic wavelengths that are most critical for predicting coal quality parameters are selected from the full-band spectral data to achieve data dimensionality reduction and improve model efficiency and generalization ability.

[0028] Step S103: Input the characteristic wavelength vector into the pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

[0029] Specifically, the feature wavelength vector obtained in step S102 is used as input and fed into the GA-optimized XGBoost model to directly obtain the prediction results of the coal quality parameters of the coal sample.

[0030] As an optional embodiment, coal quality parameters include the caking index and the maximum thickness of the plastic layer.

[0031] The adhesion index is represented by the G value, and the maximum thickness of the adhesive layer is represented by the Y value.

[0032] As an optional embodiment, calculating the proportion of effective spectra among multiple sampled spectra includes: extracting the maximum intensity value of each spectrum; taking the sampled spectra whose maximum intensity value is within a preset intensity range as effective spectra; and calculating the ratio of the number of effective spectra to the number of sampled spectra to obtain the proportion of effective spectra among multiple sampled spectra.

[0033] First, extract the maximum intensity of each sampled spectrum to form a maximum intensity dataset. ,in, This represents the total number of sampled spectra. Further, it represents the maximum intensity value in the maximum value dataset. With a preset intensity range Compare and count those that meet the criteria. Spectral quantity These spectra are then identified as valid spectra. Finally, the number of valid spectra is calculated. With the total number of sampled spectra ratio ,Right now The ratio As a percentage of the effective spectrum.

[0034] It should be noted that the preset intensity range It can be based on specific operating parameters of the LIBS detector (such as laser energy and detector gain) and determined through prior experiments. The aim is to screen out spectral data with stable signal intensity and high signal-to-noise ratio, and to exclude invalid or abnormal spectra caused by unstable excitation, weak signal or oversaturation.

[0035] As an optional embodiment, extracting feature wavelength vectors from the effective spectrum includes: for each effective spectrum, obtaining a first spectral intensity vector of the effective spectrum; performing maximum-minimum normalization on each spectral intensity in the first spectral intensity vector to obtain a second spectral intensity vector; evaluating the SHAP value corresponding to each spectral intensity in the second spectral intensity vector based on the SHAP value analysis method; wherein, the SHAP value characterizes the weight of the feature wavelength vector corresponding to the spectral intensity on the predicted coal quality parameters; sorting each spectral intensity in the second spectral intensity vector from high to low according to the SHAP value, selecting the target spectral intensity with a SHAP value higher than a preset threshold from the sort; and obtaining the feature wavelength vector corresponding to the target spectral intensity.

[0036] The core of this step lies in three key aspects: data standardization, feature importance assessment, and feature selection. The ultimate goal is to construct an optimal, low-dimensional feature wavelength vector as the input to the model.

[0037] As an optional embodiment, a first spectral intensity vector of effective spectra is obtained, and a maximum-minimum normalization process is performed on each spectral intensity in the first spectral intensity vector to obtain a second spectral intensity vector, including: determining the maximum spectral intensity and the minimum spectral intensity in the first spectral intensity vector; calculating the difference between the maximum spectral intensity and the minimum spectral intensity; and calculating the ratio of the difference between each spectral intensity and the minimum spectral intensity in the first spectral intensity vector to the difference, thereby obtaining the second spectral intensity vector.

[0038] Specifically, the original first spectral intensity vector of each effective spectrum is first obtained. The second spectral intensity vector is obtained by performing maximum-minimum normalization on the vector. Where N represents the Nth effective spectrum, the maximum-minimum normalization formula is as follows:

[0039] in, This represents the maximum spectral intensity of the Nth effective spectrum. This represents the minimum spectral intensity of the Nth effective spectrum.

[0040] Furthermore, based on an initial prediction model pre-trained on the training set and a SHAP value analysis method, the SHAP value of the wavelength corresponding to each spectral intensity value in the second spectral intensity vector is calculated. The SHAP value quantifies the contribution weight of the characteristic wavelength to the prediction result of coal quality parameters (G or Y values). According to the calculated SHAP values, all characteristic wavelengths in the second spectral intensity vector are sorted from highest to lowest importance. A SHAP value threshold is set, and all target characteristic wavelengths with SHAP values ​​higher than the preset threshold are selected from the sorted list, or the target characteristic wavelengths corresponding to the top-ranked (e.g., the top 15) SHAP values ​​are directly selected from the sorted list. The spectral intensity values ​​corresponding to these target characteristic wavelengths are extracted from the second spectral intensity vector and arranged in the original wavelength order to form a new, dimension-reduced characteristic wavelength vector, which serves as the input to the subsequent GA-XGBoost prediction model.

[0041] It should be noted that the SHAP value (SHapley Additive exPlanations) is a model interpretability method based on game theory Shapley values, used to quantify the contribution of each feature to the prediction results of the machine learning model. In the embodiments of this application, the SHAP value represents the weight (contribution) of the feature wavelength vector corresponding to the spectral intensity to the predicted coal quality parameters.

[0042] Based on this, the most critical characteristic wavelength combinations for predicting coal quality parameters can be automatically and interpretably extracted from high-dimensional raw spectral data, laying a solid foundation for building high-performance prediction models.

[0043] As an optional embodiment, the method further includes: constructing an XGBoost prediction model, binary encoding multiple preset hyperparameters of the XGBoost prediction model to generate an initial population for the genetic algorithm; wherein the hyperparameters include at least the learning rate, regularization coefficient, and maximum tree depth; inputting feature-selected spectral data into the XGBoost prediction model for training, using the mean absolute error between the predicted and true values ​​of coal quality parameters as the fitness function, and calculating the fitness value corresponding to the chromosome of each individual in the initial population; performing selection, crossover, and mutation operations on the current initial population based on the fitness value to generate a new generation population; in response to the fitness value of the best individual in the new generation population being greater than a preset fitness threshold or reaching the maximum number of iterations, taking the corresponding hyperparameters as the optimal solution, and training the GA-XGBoost prediction model based on the optimal solution.

[0044] As an optional embodiment, the feature-filtered spectral data includes a feature wavelength vector.

[0045] Specifically, in the embodiments of this application, the training set and validation set (spectral data after feature selection) for constructing the GA-XGBoost prediction model can be feature wavelength vectors. Based on this, the acquisition method is the same as the acquisition method of feature wavelength vectors.

[0046] refer to Figure 2 This is a schematic diagram of the GA-XGBoost prediction model construction process provided in the embodiments of this application.

[0047] Step S201: Construct an XGBoost prediction model and encode multiple preset hyperparameters of the XGBoost prediction model in binary to generate the initial population of the genetic algorithm; wherein, the hyperparameters include at least the learning rate, regularization coefficient and the maximum tree depth.

[0048] Step S202: Input the spectral data after feature screening into the XGBoost prediction model for training. Use the average absolute error between the predicted and actual values ​​of coal quality parameters as the fitness function to calculate the fitness value corresponding to the chromosome of each individual in the initial population. Step S203: Based on the fitness value, perform selection, crossover, and mutation operations on the current initial population to generate a new generation of population; Step S204: In response to the fitness value of the best individual in the new generation population being greater than the preset fitness threshold or reaching the maximum number of iterations, the corresponding hyperparameters are taken as the optimal solution, and the GA-XGBoost prediction model is trained based on the optimal solution.

[0049] XGBoost is an ensemble learning algorithm based on tree models. It iteratively trains multiple decision trees and uses gradient descent to optimize the loss function, making the model's predictions more accurate. The algorithm first constructs numerous decision tree models. The first decision tree predicts coal quality parameters based on feature wavelength vectors, while the other decision trees predict the difference between the predicted and actual values. Finally, the results from each decision tree are summed. It's important to note that while theoretically increasing the number of decision trees indefinitely could guarantee high accuracy on the training set, overfitting requires limiting the model complexity in the objective function.

[0050] refer to Figure 3 This is a schematic diagram of the overall process of constructing the GA-XGBoost prediction model provided in the embodiments of this application.

[0051] Specifically, in step S201, the hyperparameters of the XGBoost prediction model are binary encoded to generate an initial population. The XGBoost hyperparameter set is defined. ,in, For learning rate, The regularization coefficient is . Let be the tree depth. The hyperparameters are binary-encoded and converted into chromosomes according to the following rules: Learning rate : 5-ary, with a precision of 0.01. Controls the contribution of each tree to the final model. The smaller the value, the more robust the model, but more trees are needed.

[0052] Regularization coefficient : Quadratic, with a precision of 0.1. L2 regularization term, directly penalizes leaf weights to prevent overfitting.

[0053] Tree depth 4-bit binary. Controls the complexity of the tree. The deeper the tree, the stronger the model, but also the more prone it is to overfitting.

[0054] Genetic algorithms process gene strings (chromosomes). Therefore, it is necessary to encode continuous hyperparameters of different types into fixed-length binary strings.

[0055] In detail, learning rate Represented using 5 bits, 5 bits can represent There are 32 states, and the interval is divided into 32 states. The binary representation is divided into 31 parts, with each part having a precision of (0.3-0.01) / 31≈0.00935, which meets the precision requirements. For example, binary 00000 is mapped to 0.01, and 11111 is mapped to 0.30.

[0056] The regularization coefficient is represented by 4 bits of binary code. 4 bits can represent It has 16 states, with a precision of 1 / 15≈0.0667.

[0057] Tree depth There are 13 possible values. They are represented by 4 bits (which can represent 16 values), directly mapped to 3~15, and redundant codes can be discarded or reused cyclically.

[0058] A chromosome is simply the concatenation of these three binary strings. For example, 01011 ( )1010 ( 1100 ).

[0059] Each tree "evaluates" the coal quality from a different perspective. Finally, the "evaluations" from all trees are summed to obtain the final predicted coal quality parameters (G value or Y value). The XGBoost prediction model output consists of a weighted sum of K decision trees, expressed as:

[0060] in, Indicates the input number Each feature spectral vector outputs a predicted score, and the predicted scores of all trees are summed to obtain the final predicted G or Y value. In the t-th iteration, the model's predicted value is the sum of the predictions from the first t-1 trees and the new number. Superposition:

[0061] in, This represents the model's prediction of the current G or Y value for the i-th coal sample in iteration t.

[0062] Furthermore, in order to find this best Define an objective function that simultaneously measures prediction accuracy and model complexity to prevent overfitting. The objective function is derived from the loss function. and regularization term Together, they aim to balance prediction accuracy and model complexity, and are represented as:

[0063] in, The loss function is the coal quality parameter (G value or Y value) predicted by the i-th characteristic wavelength parameter. Compared with actual coal quality parameters (G value or Y value) The gap between them This is a regularization term used to control the complexity of the tree. The regularization term is defined as:

[0064] Where T is the number of leaf nodes in the tree; the fewer the leaves, the simpler the tree. Let be the weight value of the j-th leaf node. Specifically, for any feature wavelength vector assigned to this node by this tree, its coal quality parameter prediction needs to be added to this weight value. value, and The penalty coefficients for the number of leaves and their weights are important hyperparameters that control the severity of the penalty applied to the number of leaves and their weights.

[0065] In the t-th iteration, the objective function can be simplified to optimizing the current tree. :

[0066] To efficiently optimize this objective, XGBoost approximates it using a second-order Taylor expansion and, through a series of derivations (differentiation, setting the derivative to zero), obtains two key results (optimal leaf weights and structure scores): Furthermore, regarding the loss function Approximation using a second-order Taylor expansion:

[0067] in, and Represent the loss function respectively The first and second gradients.

[0068] set up Let j be the set of feature wavelength vectors belonging to leaf node j. Then the objective function can be further decomposed into leaf node weights. The function (structure score) indicates the structure of the tree; the smaller the score, the better the tree structure.

[0069] Furthermore, on By taking the derivative and setting it to zero, we can obtain the optimal leaf weights:

[0070] Furthermore, Substituting into the objective function, we obtain the simplified objective value:

[0071] When generating the tree structure, the optimal split point is selected by maximizing the split gain. That is, during tree construction, it's necessary to decide at which value of which feature (wavelength) to split. The decision is based on the split gain, which measures how much the objective function is reduced by a single split.

[0072] in, , This is the sum of the first and second gradients of the left child node after the split. , It is the sum of the first and second gradients of the right child node after the split. This is a metric to evaluate the effectiveness of splitting a tree node based on a specific wavelength threshold. It iterates through all possible split points and selects the one with the highest gain. If the maximum gain is negative, the splitting process stops.

[0073] In step S202, the binary-encoded feature wavelength vector is input into the XGBoost prediction model for training, and the node parameter matrix of each decision tree is recorded simultaneously to form a quantitative basis for individual performance.

[0074] Regarding step S203, traditional XGBoost prediction models typically rely on experience or grid search to pre-define fixed hyperparameter combinations. While simple, this approach struggles to adapt to the characteristics of different datasets and is prone to getting trapped in local optima, limiting further performance improvements. In contrast, XGBoost prediction models optimized using genetic algorithms treat hyperparameters as searchable variables within a certain range, leveraging the global search capability of genetic algorithms to iteratively generate, evaluate, and select optimal hyperparameter combinations. This process utilizes a fitness function (such as MAE or...) This method measures the performance of each set of hyperparameters to ultimately find the configuration best suited for the current dataset. It not only significantly improves the model's accuracy, stability, and generalization ability, enabling more precise predictions within a smaller error range, but also effectively reduces the risk of overfitting and enhances the model's reliability on unknown data.

[0075] Furthermore, the prediction error of each candidate individual on the training set is calculated, and the individual fitness value is quantified using a preset fitness function. A threshold is set. As a standard for judging the merits of individuals, among which:

[0076] Wherein, R and P are the actual and predicted values ​​of coal quality parameters (G value or Y value), respectively.

[0077] Based on the fitness values, selection, crossover, and mutation operations are performed on the initial population to generate a new generation of candidate solutions (i.e., the new generation population). The highest fitness values ​​in the new generation are then compared. The fitness value of the optimal individual is greater than the preset fitness threshold. This determines whether to terminate the iteration.

[0078] Genetic operations (selection, crossover, mutation): Selection: Similar to natural selection, individuals with high fitness (i.e., low error) are preferentially selected to reproduce and propagate the next generation. This ensures the transmission of superior genes.

[0079] Crossover: Two "parent" individuals are randomly selected, and a portion of their chromosomes are exchanged to produce "offspring." This is equivalent to combining the advantages of two excellent configurations.

[0080] Mutation: Randomly changing a bit on a chromosome with a very small probability (0 becomes 1, 1 becomes 0). This introduces new possibilities and helps to escape local optima.

[0081] The process of evaluation, selection, crossover, and mutation is repeated to generate a new generation of the population. Each generation may produce a better configuration than its parent.

[0082] like If the condition is met, the parameter solidification stage will begin; otherwise, steps S202 and S203 will be repeated until the condition is met.

[0083] For step S204, the training set is input into the model for final node splitting calculation and objective function optimization to obtain the final GA-XGBoost prediction model. The feature wavelength vectors of the test set are then input into the final GA-XGBoost prediction model to obtain the prediction results of the final coal quality parameters (G or Y values), and the model's performance index on the test set is calculated. , MAE, MSE).

[0084] In summary, the coal quality parameter prediction method provided in this application first acquires multiple sampled spectra of a coal sample using a LIBS analyzer, and calculates the proportion of effective spectra among these spectra. Further, in response to a proportion exceeding a preset threshold, a feature wavelength vector is extracted from the effective spectra. Finally, the feature wavelength vector is input into a pre-constructed GA-XGBoost prediction model to predict the coal quality parameters of the coal sample. This application, by introducing laser-induced breakdown spectroscopy (LIBS) technology and employing a genetic algorithm (GA) to optimize the hyperparameters of the XGBoost model, replaces the traditional, cumbersome offline laboratory analysis process. It enables rapid spectral acquisition and analysis, significantly improving the efficiency of coal quality parameter detection and providing real-time data support for industrial production.

[0085] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and the multiple devices will interact with each other to complete the above method.

[0086] It should be noted that the above description describes some embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in a different order than that shown in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] Corresponding to the above embodiments, the present invention also proposes a device for predicting coal quality parameters.

[0088] like Figure 4The diagram shown is a schematic of a coal quality parameter prediction device provided in an embodiment of this application. The device includes: an acquisition module 401, a response module 402, and a prediction module 403.

[0089] The acquisition module 401 is configured to acquire multiple sampling spectra obtained by spectral acquisition of coal samples through a LIBS detector, and calculate the proportion of effective spectra in the multiple sampling spectra; The response module 402 is configured to extract a feature wavelength vector from the effective spectrum in response to a proportion greater than a preset proportion threshold. The prediction module 403 is configured to input the characteristic wavelength vector into a pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

[0090] Optionally, module 401 is also configured as follows: Extract the maximum intensity of each spectrum; The sampled spectrum whose maximum intensity is within the preset intensity range is taken as the effective spectrum; Calculate the ratio of the number of effective spectra to the number of sampled spectra to obtain the proportion of effective spectra among multiple sampled spectra.

[0091] Optionally, response module 402 is also configured as follows: For each valid spectrum, the first spectral intensity vector of the valid spectrum is obtained, and the maximum-minimum value normalization process is performed on each spectral intensity in the first spectral intensity vector to obtain the second spectral intensity vector. The SHAP value is evaluated based on the SHAP value analysis method for each spectral intensity in the second spectral intensity vector; where the SHAP value represents the weight of the characteristic wavelength vector corresponding to the spectral intensity on the predicted coal quality parameters. Sort each spectral intensity in the second spectral intensity vector from highest to lowest according to the SHAP value, and select the target spectral intensity whose SHAP value is higher than a preset threshold from the sorted list. Obtain the characteristic wavelength vector corresponding to the target spectral intensity.

[0092] Optionally, response module 402 is also configured as follows: Determine the maximum and minimum spectral intensities in the first spectral intensity vector; Calculate the difference between the maximum and minimum spectral intensities; The second spectral intensity vector is obtained by calculating the ratio of the difference between each spectral intensity and the minimum spectral intensity in the first spectral intensity vector to the difference value.

[0093] Optionally, module 401 is also configured as follows: An XGBoost prediction model is constructed, and multiple preset hyperparameters of the XGBoost prediction model are binary encoded to generate the initial population of the genetic algorithm; among them, the hyperparameters include at least the learning rate, regularization coefficient and the maximum tree depth. The spectral data after feature screening is input into the XGBoost prediction model for training. The average absolute error between the predicted and actual values ​​of coal quality parameters is used as the fitness function to calculate the fitness value corresponding to the chromosome of each individual in the initial population. Based on the fitness value, perform selection, crossover, and mutation operations on the current initial population to generate a new generation of population; In response to the optimal individual in the new generation population having a fitness value greater than the preset fitness threshold or reaching the maximum number of iterations, the corresponding hyperparameters are taken as the optimal solution, and the GA-XGBoost prediction model is trained based on the optimal solution.

[0094] Optionally, the feature-filtered spectral data may include feature wavelength vectors.

[0095] Optional parameters for coal quality include the caking index and the maximum thickness of the plastic layer.

[0096] For ease of description, the above system is described by dividing it into various modules based on their functions. Of course, in implementing this invention, the functions of each module can be implemented in one or more software and / or hardware components.

[0097] The apparatus of the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0098] Corresponding to the above embodiments, the present invention also proposes an electronic device.

[0099] refer to Figure 5 The diagram below is a block diagram of an electronic device according to some embodiments of the present invention. It illustrates a more specific hardware structure of the electronic device provided in this embodiment. The device may include: a processor 510, a memory 520, an input / output interface 530, a communication interface 540, and a bus 550. The processor 510, memory 520, input / output interface 530, and communication interface 540 are interconnected internally via the bus 550.

[0100] The processor 510 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0101] The memory 520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 520 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 520 and is called and executed by the processor 510.

[0102] Input / output interface 530 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.

[0103] The communication interface 540 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0104] Bus 550 includes a pathway for transmitting information between various components of the device, such as processor 510, memory 520, input / output interface 530, and communication interface 540.

[0105] It should be noted that although the above-described device only shows the processor 510, memory 520, input / output interface 530, communication interface 540, and bus 550, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0106] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0107] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the present invention also provides a computer-readable storage medium storing computer instructions for causing a computer to perform the methods of any of the above embodiments.

[0108] The aforementioned computer-readable storage medium can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).

[0109] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods of any of the above exemplary method sections, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0110] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Rather, the steps depicted in the flowchart may be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0111] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0112] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this invention should have the ordinary meaning understood by those skilled in the art. The terms "first," "second," and similar terms used in the embodiments of this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0113] While the spirit and principles of the invention have been described with reference to several specific embodiments, it should be understood that the invention is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for ease of description. The invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be interpreted in the broadest sense, thereby encompassing all such modifications and equivalent structures and functions.

Claims

1. A method for predicting coal quality parameters, characterized in that, include: Multiple sampling spectra obtained by spectral acquisition of coal samples using a LIBS detector are acquired, and the proportion of effective spectra in the multiple sampling spectra is calculated; In response to the proportion being greater than a preset proportion threshold, a feature wavelength vector is extracted from the effective spectrum; The characteristic wavelength vector is input into a pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

2. The method for predicting coal quality parameters according to claim 1, characterized in that, The calculation of the percentage of effective spectra among the multiple sampled spectra includes: Extract the maximum intensity of each spectrum; The sampled spectrum whose maximum intensity is within a preset intensity range is taken as the effective spectrum; Calculate the ratio of the number of effective spectra to the number of sampled spectra to obtain the proportion of effective spectra among the multiple sampled spectra.

3. The method for predicting coal quality parameters according to claim 1, characterized in that, The step of extracting the feature wavelength vector from the effective spectrum includes: For each of the effective spectra, a first spectral intensity vector of the effective spectrum is obtained, and a maximum-minimum normalization process is performed on each spectral intensity in the first spectral intensity vector to obtain a second spectral intensity vector; The SHAP value is evaluated based on the SHAP value analysis method for each spectral intensity in the second spectral intensity vector; wherein the SHAP value represents the weight of the characteristic wavelength vector corresponding to the spectral intensity in predicting the coal quality parameter; Based on the SHAP value, each spectral intensity in the second spectral intensity vector is sorted from high to low, and a target spectral intensity with a SHAP value higher than a preset threshold is selected from the sorted values. Obtain the characteristic wavelength vector corresponding to the target spectral intensity.

4. The method for predicting coal quality parameters according to claim 3, characterized in that, The process of obtaining the first spectral intensity vector of the effective spectrum, and performing maximum-minimum normalization on each spectral intensity in the first spectral intensity vector to obtain the second spectral intensity vector, includes: Determine the maximum and minimum spectral intensities in the first spectral intensity vector; Calculate the difference between the maximum spectral intensity and the minimum spectral intensity; The second spectral intensity vector is obtained by calculating the ratio of the difference between each spectral intensity in the first spectral intensity vector and the minimum spectral intensity to the difference value.

5. The method for predicting coal quality parameters according to claim 1, characterized in that, The method further includes: An XGBoost prediction model is constructed, and multiple preset hyperparameters of the XGBoost prediction model are binary encoded to generate the initial population of the genetic algorithm; wherein, the hyperparameters include at least the learning rate, regularization coefficient, and maximum tree depth; The spectral data after feature screening is input into the XGBoost prediction model for training. The average absolute error between the predicted and actual values ​​of coal quality parameters is used as the fitness function to calculate the fitness value corresponding to the chromosome of each individual in the initial population. Based on the fitness value, selection, crossover, and mutation operations are performed on the current initial population to generate a new generation population; In response to the fitness value of the best individual in the new generation population being greater than the preset fitness threshold or reaching the maximum number of iterations, the corresponding hyperparameters are taken as the optimal solution, and the GA-XGBoost prediction model is trained based on the optimal solution.

6. The method for predicting coal quality parameters according to claim 5, characterized in that, The feature-filtered spectral data includes the feature wavelength vector.

7. The method for predicting coal quality parameters according to claim 1, characterized in that, The coal quality parameters include the caking index and the maximum thickness of the plastic layer.

8. A device for predicting coal quality parameters, characterized in that, include: The acquisition module is configured to acquire multiple sampling spectra obtained by spectral acquisition of coal samples using a LIBS detector, and calculate the proportion of effective spectra in the multiple sampling spectra; The response module is configured to extract a feature wavelength vector from the effective spectrum in response to the proportion being greater than a preset proportion threshold. The prediction module is configured to input the characteristic wavelength vector into a pre-built GA-XGBoost prediction model to predict the coal quality parameters of the coal sample.

9. An electronic device, characterized in that, include: A processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method for predicting coal quality parameters as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the coal quality parameter prediction method as described in any one of claims 1 to 7.