A method, device, electronic device and storage medium for training a coal quality analysis model

By performing sample set division and cross-validation optimization on the coal quality analysis model, the problem of insufficient generalization ability of the model is solved and higher analysis accuracy is achieved.

CN117216558BActive Publication Date: 2025-07-29GUANGDONG ENERGY GROUP SCIENCE & TECHNOLOGY RESEARCH INSTITUTE CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311120852.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-07-29
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

The existing coal quality analysis model established based on LIBS is insufficient in generalization ability, and it is prone to overfitting, which affects the accuracy of coal quality analysis.

Method used

By dividing the sample set of the initial coal quality analysis model into a training set, a verification set and a test set, the ash score and volatile score are used for clustering and grouping, and using one-left cross-validation for optimization, combining the verification set to optimize the candidate model, the final coal quality analysis model is obtained.

Benefits of technology

It enhances the generalization ability of the coal quality analysis model, avoids overfitting, and improves the accuracy of coal quality analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216558B_ABST
    Figure CN117216558B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, device, electronic device, and storage medium for training a coal quality analysis model. The method includes: if the initial coal quality analysis model does not meet the pre-set convergence condition, sorting out a training set, a validation set, and a test set of the initial coal quality analysis model from all coal samples in the pre-obtained coal sample library; wherein, the coal sample library at least includes spectral data, coal quality indexes to be analyzed, ash content values, and volatile matter content values of each coal sample; optimizing through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition; and optimizing the candidate coal quality analysis model based on the validation set to obtain a final coal quality analysis model. The method of the embodiment of the present invention enhances the generalization ability of the coal quality analysis model, avoids overfitting of the model, and improves the accuracy of the coal quality analysis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to coal quality detection technologies, and in particular, to a method, device, electronic device, and storage medium for training a coal quality analysis model. Background Art

[0002] In coal-fired power plants, the inspection of coal quality entering the furnace mainly uses off-line analysis technologies, that is, after on-site sampling, it is sent to the laboratory for preparation and analysis. However, the off-line analysis technology has low detection efficiency and is difficult to meet the optimized operation requirements of on-line / rapid coal quality inspection in power plants. In recent years, Laser-Induced Breakdown Spectroscopy (LIBS) has attracted much attention in the field of coal quality inspection due to its unique advantages such as simple and easy operation, no need for complex sample pretreatment, and multi-element synchronous on-line analysis. The working principle of using LIBS for coal quality inspection is to focus a pulsed laser beam on the surface of the sample and generate plasma. A high-resolution spectrometer collects the optical radiation signals emitted during the cooling process of the plasma. By analyzing the spectra with specific wavelengths and intensities, the elemental types and concentration information of the coal sample are obtained. Then, a coal quality analysis model is established using the spectra corresponding to each coal sample.

[0003] Currently, the generalization ability of the coal quality analysis model established based on LIBS is insufficient, which easily causes the model to overfit, seriously affecting the accuracy of coal quality analysis. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device, electronic device, and storage medium for training a coal quality analysis model, which can divide the coal sample set according to the ash content and volatile content of the coal samples to obtain a training set, a validation set, and a test set. The training set is input into the algorithm to establish models for the indexes to be measured (calorific value, carbon content), and optimization is carried out through leave-one-out cross-validation. The validation set is used to further optimize the model to obtain the final coal quality analysis model. This method enhances the generalization ability of the coal quality analysis model, improves the reliability of the coal quality analysis model, and avoids overfitting of the model.

[0005] In a first aspect, the embodiments of the present invention provide a method for training a coal quality analysis model, including:

[0006] If the initial coal quality analysis model does not meet the pre-set convergence condition, then a training set, a validation set, and a test set of the initial coal quality analysis model are sorted out from all the coal samples in the pre-obtained coal sample library; wherein, the coal sample library at least includes the spectral data, coal quality indexes to be analyzed, ash content, and volatile content of each coal sample; wherein, the coal quality indexes to be analyzed include carbon content and calorific value;

[0007] Based on the training set, optimization is performed through leave-one-out cross-validation to obtain a candidate coal quality analysis model that meets the convergence condition;

[0008] Based on the validation set, the candidate coal quality analysis model is optimized to obtain the final coal quality analysis model.

[0009] In a second aspect, an embodiment of the present invention provides a device for training a coal quality analysis model, the device including:

[0010] A data sorting module, configured to sort out the training set, validation set, and test set of the initial coal quality analysis model from all the coal samples in the pre-acquired coal sample library if the initial coal quality analysis model does not meet the pre-set convergence condition; wherein, the coal sample library at least includes the spectral data, coal quality indexes to be analyzed, ash content, and volatile content of each coal sample; wherein, the coal quality indexes to be analyzed include carbon content and calorific value;

[0011] A model training module, configured to perform optimization through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition;

[0012] A model optimization module, configured to optimize the candidate coal quality analysis model based on the validation set to obtain the final coal quality analysis model.

[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the coal quality analysis model training method as described in any one of the embodiments of the present invention.

[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the coal quality analysis model training method as described in any one of the embodiments of the present invention.

[0015] In the embodiments of the present invention, if the initial coal quality analysis model does not meet the pre-set convergence condition, a training set, a validation set, and a test set of the initial coal quality analysis model are sorted out from all the coal samples in the pre-obtained coal quality sample library; wherein, the coal quality sample library at least includes the spectral data, the coal quality indexes to be analyzed, the ash content value, and the volatile content value of each coal sample; wherein, the coal quality indexes to be analyzed include the carbon content and the calorific value; according to the training set, optimization is carried out through leave-one-out cross-validation to obtain a candidate coal quality analysis model that meets the convergence condition; the candidate coal quality analysis model is optimized based on the validation set to obtain the final coal quality analysis model. That is, in the embodiments of the present invention, the coal samples can be divided into a training set, a validation set, and a test set, the training set is used to train the coal quality analysis model, and the validation set is used to further optimize the coal quality analysis model, which enhances the generalization ability of the coal quality analysis model, avoids the phenomenon of overfitting of the model, and improves the accuracy of the coal quality analysis model. Brief Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a method for training a coal quality analysis model provided by an embodiment of the present invention;

[0018] Figure 2 It is a flowchart of obtaining a candidate coal quality analysis model that meets the convergence condition provided by an embodiment of the present invention;

[0019] Figure 3 It is a schematic structural diagram of a device for training a coal quality analysis model provided by an embodiment of the present invention;

[0020] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0021] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that, for the sake of description, only the parts related to the present invention are shown in the drawings rather than all the structures.

[0022] Figure 1This is a flowchart of the coal quality analysis model training method provided by the embodiments of the present invention. The method of the embodiments of the present invention can enhance the generalization ability of the coal quality analysis model, avoid overfitting of the model, and improve the accuracy of the coal quality analysis model. This method can be executed by the coal quality analysis model training device provided by the embodiments of the present invention, and the device can be implemented in software and / or hardware. The electronic device can be a controller of an electric vehicle, etc. The following embodiments will take the integration of the device in the electronic device as an example for description. Refer to Figure 1 , and the method specifically may include the following steps:

[0023] Step 101: If the initial coal quality analysis model does not meet the pre-set convergence condition, then sort out the training set, validation set, and test set of the initial coal quality analysis model from all the coal samples in the pre-obtained coal sample library.

[0024] Among them, the coal sample library at least includes the spectral data, coal quality indexes to be analyzed, ash content, and volatile content of each coal sample. The coal quality indexes to be analyzed include carbon content and calorific value. The initial coal quality analysis model is a model that has not been trained well. The training set is a sample set for training the initial coal quality analysis model, and the validation set is a sample set for optimizing the candidate coal quality analysis model. Coal spectral data refers to the data obtained by spectral analysis of coal, including visible spectrum, infrared spectrum, ultraviolet spectrum, etc. These data can be used for qualitative and quantitative analysis of coal, such as determining parameters such as the type of coal, carbon content, and sulfur content. In addition, the spectral data of coal can also be used to study the chemical composition, structural characteristics, etc. of coal to better understand the nature and application value of coal.

[0025] In an alternative embodiment, after obtaining the coal sample library, cluster and group all the coal samples according to the ash content and volatile content of all the coal samples to obtain multiple groups of coal samples; screen out the training set samples, validation set samples, and test set samples from each group of coal samples according to a preset ratio; combine the training set samples of all groups of coal samples to obtain the training set; combine the validation set samples of all groups of coal samples to obtain the validation set; combine the test set samples of all groups of coal samples to obtain the test set.

[0026] Step 102: Optimize through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition.

[0027] Among them, the candidate coal quality analysis model is the model obtained by training the initial coal quality analysis model with a training set. Specifically, the training set includes a large number of coal combustion samples, as well as the coal quality index values to be analyzed corresponding to each coal combustion sample. The coal quality index values to be analyzed are the index values of the coal combustion samples that the user needs to obtain, such as calorific value, carbon content and other indexes. Cross-validation is a commonly used model evaluation method for evaluating the performance and generalization ability of machine learning models. Leave-one-out cross-validation is a special form of cross-validation. In leave-one-out cross-validation, for a data set containing N samples, N times of training and evaluation will be carried out. Each time, N-1 samples are used for training, and then the excluded one sample is used for testing. Finally, the results of the N times of evaluation are averaged as the performance evaluation index of the model.

[0028] In an alternative embodiment, training the initial coal quality analysis model with the current sample and performing optimization through leave-one-out cross-validation includes: selecting a sample from the training set as the current sample, and establishing a spectral matrix and an index matrix of the initial coal quality analysis model based on the spectral data of the current sample and the index values of the coal quality index to be analyzed corresponding to the current sample; calculating the variances of the principal components of the spectral matrix and the principal components of the index matrix, and determining a spectral coefficient matrix, a spectral residual matrix, an index coefficient matrix and an index residual matrix based on the variances; adding the product of the principal components of the spectral matrix and the transpose matrix of the spectral coefficient matrix to the spectral residual matrix to obtain a spectral regression matrix; adding the product of the principal components of the index matrix and the transpose matrix of the index coefficient matrix to the index residual matrix to obtain an index regression matrix. Obtaining the correlation between the principal components of the spectral matrix and the principal components of the index matrix, and when the correlation is greater than a preset value, determining a target residual matrix based on the spectral residual matrix and the index residual matrix; determining a target coefficient matrix based on the spectral coefficient matrix and the index coefficient matrix; adding the product of the principal components of the spectral matrix and the transpose matrix of the target coefficient matrix to the target residual matrix to obtain a target regression matrix. When the accuracy of the target residual matrix is less than the preset accuracy, or when the number of the principal components of the index matrix is less than the preset number, updating the spectral residual matrix to the spectral matrix, updating the target residual matrix to the index matrix, adjusting the model parameters of the initial coal quality analysis model based on the updated index matrix, and repeating the step of selecting a sample from the training set as the current sample until a candidate coal quality analysis model that meets the convergence condition is obtained.

[0029] In this solution, the index matrix of the initial coal quality analysis model is established based on the coal quality indexes to be analyzed, that is, based on calorific value and carbon content. Therefore, the candidate coal quality analysis model trained according to the training set can be used to analyze the calorific value and carbon content in coal quality. The division of the sample set is obtained by performing cluster analysis according to the ash content value and volatile content value of the coal samples. That is, in this solution, instead of directly classifying the coal samples using carbon content and calorific value, the ash content value and volatile content value with high correlation with calorific value and carbon content are selected to classify the coal samples. In this way, the generalization ability of the coal quality analysis model is enhanced, and a better prediction and analysis result is obtained than directly grouping the samples using carbon content and calorific value.

[0030] Step 103: Optimize the candidate coal quality analysis model based on the validation set to obtain the final coal quality analysis model.

[0031] Among them, the validation set is used to further optimize the candidate coal quality analysis model. Specifically, after training the initial coal quality analysis model with the training set to obtain the candidate coal quality analysis model, the validation set can be used to evaluate the candidate coal quality model to find the best model parameters (that is, the model parameters when obtaining the best index matrix). Further, the prediction result of the model is verified through the test set.

[0032] In an embodiment of this solution, optionally, optimizing the candidate coal quality analysis model based on the validation set to obtain the final coal quality analysis model includes: inputting the validation set into the candidate coal quality analysis model to obtain the evaluation index value corresponding to the validation set; optimizing the candidate coal quality analysis model based on the evaluation index value corresponding to the validation set to obtain the final coal quality analysis model.

[0033] In this solution, the evaluation indexes include coefficient of determination, root mean square error, and mean absolute error. The coefficient of determination is a statistical index used to evaluate the goodness of fit of a regression model. It represents the proportion of the variability of the dependent variable that can be explained by the model, that is, the fitting degree of the model to the data. The root mean square error is a commonly used index to measure the difference between the predicted value and the actual observed value of the model, and it is used to evaluate the fitting degree of the model on the given data. The root mean square error can be obtained by calculating the mean of the squares of the differences between the predicted value and the actual observed value and taking its square root. The mean absolute error is a commonly used index to measure the difference between the predicted value and the actual observed value of the model, and it is used to evaluate the fitting degree of the model on the given data. The mean absolute error can be obtained by calculating the average of the absolute values of the differences between the predicted value and the actual observed value.

[0034] Specifically, all samples in the validation set are input into the candidate coal quality analysis model, and various evaluation index values corresponding to all samples in the validation set output by the candidate coal quality analysis model are obtained. If the various evaluation index values do not reach the preset evaluation index values, the model parameters of the candidate coal quality model are adjusted until the best model parameters are found.

[0035] For the technical solution of this embodiment, if the initial coal quality analysis model does not meet the preset convergence condition, the training set, validation set, and test set of the initial coal quality analysis model are sorted out from all coal combustion samples in the pre-obtained coal quality sample library; among them, the coal quality sample library at least includes the spectral data, coal quality indexes to be analyzed, ash content values, and volatile content values of each coal combustion sample. According to the training set, optimization is performed through leave-one-out cross-validation to obtain a candidate coal quality analysis model that meets the convergence condition; the candidate coal quality analysis model is optimized based on the validation set to obtain the final coal quality analysis model. The technical solution of this embodiment can divide the coal combustion samples into a training set, a validation set, and a test set, train the coal quality analysis model with the training set, and further optimize the coal quality analysis model with the validation set, enhancing the generalization ability of the coal quality analysis model, avoiding the phenomenon of overfitting of the model, and improving the accuracy of the coal quality analysis model.

[0036] Figure 2 This is a flowchart for obtaining a candidate coal quality analysis model that meets the convergence condition provided by an embodiment of the present invention, and this embodiment is a refinement based on the above embodiment. The specific method can be as Figure 2 shown, and the method may include the following steps:

[0037] Step 201: Cluster and group all coal combustion samples according to the ash content values and volatile content values of all coal combustion samples to obtain multiple groups of coal combustion samples.

[0038] Among them, the clustering grouping is the grouping obtained according to the analysis results after clustering analysis of all coal samples. Clustering analysis is a process of classifying data into different classes or clusters. Objects in the same cluster have great similarities, while objects in different clusters have great dissimilarities. In this solution, the Statistical Package for the Social Sciences (SPSS) can be used to perform clustering analysis on coal samples. Specifically, during the classification process, ash content and volatile content are used as variables for clustering analysis. The proximity degree between all samples is measured according to the ward method, and the closest ones are first clustered into a small class. Then, the proximity degree between the remaining samples and this small class is measured, and the currently closest sample or small class is clustered into a class again. This process is repeated until all samples are clustered into one class. The interval distance for determining the proximity degree between samples is the Euclidean distance, and the standardized value is 0 - 1. The ash content value refers to the amount of residue after the coal is completely burned at a certain temperature. In practical applications, the lower the ash content of coal for various uses, the better. An increase in ash content will cause a decrease in the calorific value of the coal. In addition, the higher the ash content of the coal, a large amount of heat will also be carried away during the emission protection process. Therefore, the higher the ash content of the coal, the easier it is to slag during the combustion and gasification processes, which will affect the normal operation.

[0039] In an alternative implementation, the coal sample library includes the numbers of each coal sample, as well as the ash content values and volatile content values of each coal sample. The ash content values and volatile content values of all coal samples are input into SPSS, and SPSS is used to perform clustering information on all coal samples based on the ash content values and volatile content values, that is, coal samples with close ash content values and volatile content values are grouped into one group. Further, multiple groups of coal samples output by SPSS and the numbers of each coal sample in each group of coal samples are obtained.

[0040] Step 202: Screen out training set samples, validation set samples, and test set samples from each group of coal samples according to a preset ratio.

[0041] Among them, the training set is a sample set used to train the initialized coal quality analysis model, and the validation set is a sample set used to optimize the candidate coal quality analysis model. The preset ratio can be 8:1:1. Specifically, after obtaining multiple groups of coal samples, training set samples, validation set samples, and test set samples can be screened out from each group of coal samples according to the preset ratio.

[0042] Exemplarily, assume that there are a total of 3 groups of coal samples, with 60, 30, and 10 coal samples in the 3 groups of coal samples respectively. Then, according to the ratio of 8:1:1, 48 training set samples, 6 validation set samples, and 6 test set samples can be selected from the first group of coal samples. 24 training set samples, 3 validation set samples, and 3 test set samples can be selected from the second group of coal samples. 8 training set samples, 1 validation set sample, and 1 test set sample can be selected from the third group of coal samples.

[0043] Step 203: Combine the training set samples of all groups of coal samples to obtain a training set; combine the validation set samples of all groups of coal samples to obtain a validation set; combine the test set samples of all groups of coal samples to obtain a test set.

[0044] Specifically, after obtaining the training set samples, validation set samples, and test set samples, the training set samples, validation set samples, and test set samples of each group of coal samples are combined respectively to obtain a training set, a validation set, and a test set. Exemplarily, 48 training set samples, 6 validation set samples, and 6 test set samples are selected from the first group of coal samples. 24 training set samples, 3 validation set samples, and 3 test set samples are selected from the second group of coal samples. 8 training set samples, 1 validation set sample, and 1 test set sample are selected from the third group of coal samples. Then the training set includes: 48 training set samples from the first group + 24 training set samples from the second group + 8 training set samples from the third group = 80 training set samples. The validation set includes: 6 validation set samples from the first group + 3 validation set samples from the second group + 1 validation set sample from the third group = 10 validation set samples. The test set includes: 6 test set samples from the first group + 3 test set samples from the second group + 1 test set sample from the third group = 10 test set samples.

[0045] Step 204: Extract a coal sample from the training set as the current sample, and establish a spectral matrix and an index matrix of the initial coal quality analysis model based on the spectral data of the current sample and the index values of the coal quality indexes to be analyzed corresponding to the current sample.

[0046] Specifically, the coal quality indexes to be analyzed include calorific value and carbon content, and the index matrix includes carbon content values and calorific value values. After extracting the current sample, a spectral matrix of the initial coal quality analysis model is established based on the spectral data of the current sample, and an index matrix of the initial coal quality analysis model is established based on the index values of the coal quality indexes to be analyzed corresponding to the current sample. Among them, the spectral matrix and the index matrix are component matrices, and the component matrix is a type of matrix. The component matrix is the factor loading matrix obtained by the principal component method. When comparing with the same group of subjects, it is necessary to ensure that there is no mutual influence between the two experimental treatments, and at the same time, the position order should be balanced. The component matrix includes principal components and axis vectors, etc.

[0047] In steps 201 - 204, cluster analysis was performed on the coal samples according to the ash content and volatile content with high correlations with calorific value and carbon content. The samples were grouped at a certain spatial distance, and a training set, a validation set, and a test set were obtained according to the grouping results. An index matrix of the grouped sample set and the indexes to be measured (carbon content, calorific value) was constructed, and an initial model was obtained through the training set, and the model was optimized using the validation set. Such a modeling method enhances the generalization ability of the coal quality analysis model and effectively avoids the occurrence of overfitting of the model.

[0048] Step 205: Based on the spectral matrix and the index matrix, determine the spectral regression matrix and the index regression matrix of the initial coal quality analysis model.

[0049] Among them, the spectral matrix of the current sample is a matrix obtained from the spectral data of the current sample. The index matrix of the current sample is a matrix obtained from the index values, ash content, and volatile content of the coal quality indexes corresponding to the current sample. In this solution, optionally, based on the spectral matrix and the index matrix, determining the spectral regression matrix and the index regression matrix of the initial coal quality analysis model includes the following steps A1 - A3:

[0050] Step A1: Calculate the variances of the principal components of the spectral matrix and the principal components of the index matrix, and determine the spectral coefficient matrix, the spectral residual matrix, the index coefficient matrix, and the index residual matrix based on the variances.

[0051] Specifically, after establishing the spectral matrix and the index matrix, principal component analysis can be performed on the spectral matrix and the index matrix. During the principal component analysis process, the variances of the principal components of the spectral matrix and the principal components of the index matrix are obtained.

[0052] Exemplarily, based on the spectral data of the current sample, the spectral matrix is established as X, and the index matrix is established as Y. Let the first principal component of X be t1, and the first principal component of Y be u1. Let ω1 and c1 be the axis vectors of the first principal components of X and Y respectively, that is, t1 = X * ω1, u1 = Y * c1. During the principal component analysis process, the variances of t1 and u1 and the correlation between t1 and u1 are obtained using the axis vectors ω1 and c1. ω1 and c1 are adjusted and calculated using the correlation between t1 and u1 and their respective variances. When the variances and correlations of t1 and u1 are maximized, ω1 and c1 are determined.

[0053] Furthermore, the least - squares method is used to determine the spectral coefficient matrix, the spectral residual matrix, the spectral residual matrix, and the index residual matrix. The least - squares residual method is a commonly used linear regression analysis method. Its main idea is to determine the coefficients of the best - fitting line by finding the method that minimizes the sum of the squares of the distances between the data points and the fitting line.

[0054] Step A2: Add the product of the principal components of the spectral matrix and the transpose matrix of the spectral coefficient matrix to the spectral residual matrix to obtain the spectral regression matrix.

[0055] Specifically, after determining the spectral coefficient matrix and the spectral residual matrix, add the product of the principal components of the spectral matrix and the transpose matrix of the spectral coefficient matrix to the spectral residual matrix to obtain the spectral regression matrix.

[0056] Exemplarily, if the spectral coefficient matrix is t1, the spectral coefficient matrix is represented by p1, E represents the spectral residual matrix, and the spectral matrix of the previous step is regressed, and the spectral regression matrix is used to replace the spectral matrix, then the spectral regression matrix can be obtained as X = t1p1 T + E.

[0057] Step A3: Add the product of the principal components of the index matrix and the transpose matrix of the index coefficient matrix to the index residual matrix to obtain the index regression matrix.

[0058] Specifically, after determining the index coefficient matrix and the index residual matrix, add the product of the principal components of the index matrix and the transpose matrix of the index coefficient matrix to the index residual matrix to obtain the index regression matrix.

[0059] Exemplarily, if the index coefficient matrix is u1, the index coefficient matrix is represented by q1, G represents the index residual matrix, and the index matrix of the previous step is regressed, and the index regression matrix is used to replace the index matrix, then the index regression matrix can be obtained as

[0060] Through the above steps, based on the quantitative analysis method of the least squares method, it is possible to efficiently analyze complex samples, accurately obtain the spectral regression matrix and the index regression matrix, laying a foundation for subsequent adjustment and optimization of model parameters.

[0061] Step 206: Determine the target regression matrix based on the spectral regression matrix and the index regression matrix, and update the index matrix according to the spectral regression matrix and the target regression matrix.

[0062] Specifically, perform regression modeling on the index regression matrix based on the spectral regression matrix and the index regression matrix to obtain the target regression matrix. In this solution, optionally, determining the target regression matrix based on the spectral regression matrix and the index regression matrix includes the following steps B1 - Step B2:

[0063] Step B1: Obtain the correlation between the principal components of the spectral matrix and the principal components of the index matrix. When the correlation is greater than a preset value, determine the target residual matrix based on the spectral residual matrix and the index residual matrix, and determine the target coefficient matrix based on the spectral coefficient matrix and the index coefficient matrix.

[0064] Perform principal component analysis on the principal components of the spectral matrix and the principal components of the index matrix to obtain the correlation between the principal components of the spectral matrix and the principal components of the index matrix. When the correlation is greater than the preset value, it indicates that there is no abnormality in the calculation process. Further, based on the spectral residual matrix and the index residual matrix, the target residual matrix is calculated using the least squares method. Based on the spectral coefficient matrix and the index coefficient matrix, the target coefficient matrix is calculated using the least squares method.

[0065] Step B2: Add the product of the principal components of the spectral matrix and the transposed matrix of the target coefficient matrix to the target residual matrix to obtain the target regression matrix.

[0066] Specifically, after obtaining the target residual matrix and the target coefficient matrix, perform regression on the index regression matrix of the previous step again, and replace the index regression matrix with the target regression matrix. Exemplarily, based on the correlation between the principal components of the spectral matrix and the principal components of the index matrix, perform regression modeling on the index regression matrix based on the principal component t1 to obtain the target regression matrix Y = t1r1 T +F. Where r1 is the target coefficient matrix and F is the target residual matrix.

[0067] Further, after obtaining the target regression matrix, update the index matrix according to the spectral regression matrix and the target regression matrix. In this solution, optionally, updating the index matrix includes: when the accuracy of the target residual matrix is less than the preset accuracy, or when the number of principal components of the index matrix is less than the preset number, update the spectral residual matrix to the spectral matrix and update the target residual matrix to the index matrix.

[0068] Specifically, when the accuracy of the target residual matrix F does not meet the preset accuracy requirement, or when the number of principal components (i.e., latent variables) does not reach the upper limit, take the residual part E that the principal component t1 in X cannot explain as the new X, and take the residual part F that the principal component t1 in Y cannot explain as the new Y, and repeat the above steps until the residual F meets the accuracy requirement or the number of principal components reaches the upper limit, then stop updating to obtain the final index matrix. Suppose there are k principal components in the end, then the original X and Y can be expressed as:

[0069]

[0070] Further, represent X and Y in matrix form to obtain:

[0071] X = TP T +E; Y = TR T +F = XWR T +F; where W can be calculated from the data of the training set, and R represents the establishment of the quantitative analysis model.

[0072] After the training is completed, for the spectral data x corresponding to a coal sample in the validation set and the test set, first, the principal components can be calculated using W, i.e., t1 = x T ω1, t2 = x T ω2, …, t k = x T ω k , and then substitute it into to obtain the Y corresponding to the validation set and the Y corresponding to the test set.

[0073] Step 207: Adjust the model parameters of the initial coal quality analysis model based on the updated index matrix, and repeat the step of selecting a sample from the training set as the current sample until a candidate coal quality analysis model that meets the convergence condition is obtained.

[0074] Among them, the model parameters can determine the performance of the model. Specifically, after obtaining the index matrix (i.e., the predicted index value of the current sample output by the initial coal quality model), use a pre-set loss function to calculate the gap between the index matrix and the true index value of the current sample, and continuously adjust the model parameters according to this gap to make the value predicted by the model closer to the true value, thereby improving the accuracy of the model.

[0075] In the above steps, the target regression matrix can be accurately determined, and the index matrix can be updated according to the spectral regression matrix and the target regression matrix, and the predicted index value corresponding to the current sample output by the initial coal quality analysis model can be accurately obtained, laying a foundation for accurately adjusting the model parameters in the subsequent steps.

[0076] In the embodiment of the present invention, all coal samples are clustered and grouped according to the ash content and volatile content of all coal samples to obtain multiple groups of coal samples. Training set samples, validation set samples, and test set samples are respectively selected from each group of coal samples according to a preset ratio. The training set samples of all groups of coal samples are combined to obtain a training set; the validation set samples of all groups of coal samples are combined to obtain a validation set; the test set samples of all groups of coal samples are combined to obtain a test set. A coal sample is extracted from the training set as the current sample, and a spectral matrix and an index matrix of an initial coal quality analysis model are established based on the spectral data of the current sample and the index value of the coal quality index to be analyzed corresponding to the current sample. Based on the spectral matrix and the index matrix, a spectral regression matrix and an index regression matrix of the initial coal quality analysis model are determined. A target regression matrix is determined based on the spectral regression matrix and the index regression matrix, and the index matrix is updated according to the spectral regression matrix and the target regression matrix. The model parameters of the initial coal quality analysis model are adjusted based on the updated index matrix. The step of selecting a sample from the training set as the current sample is repeatedly executed until a candidate coal quality analysis model that meets the convergence condition is obtained. The technical solution of this embodiment uses the ash content and volatile content as clustering analysis variables to perform clustering analysis on all coal samples, and then a small number of samples are extracted from each category of coal samples as the validation set and the test set respectively, and the remaining samples are used as the training set to train the coal quality analysis model, so that the model pays attention to each category of coal samples during the training process. In this way, the generalization of the model can be enhanced, the phenomenon of overfitting of the model can be avoided, and the accuracy of the coal quality analysis model can be further improved.

[0077] Figure 3 FIG. is a schematic structural diagram of a coal quality analysis model training device provided by an embodiment of the present invention. This device is applicable to execute the coal quality analysis model training method provided by the embodiment of the present invention. As Figure 3 shown, this device may specifically include:

[0078] A data sorting module 301, configured to sort out a training set, a validation set, and a test set of the initial coal quality analysis model from all coal samples in a pre-acquired coal sample library if the initial coal quality analysis model does not meet a pre-set convergence condition; wherein, the coal sample library at least includes the spectral data, the coal quality index to be analyzed, the ash content, and the volatile content of each coal sample;

[0079] A model training module 302, configured to perform optimization through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition;

[0080] A model optimization module 303, configured to optimize the candidate coal quality analysis model based on the validation set to obtain a final coal quality analysis model.

[0081] Optionally, the model optimization module 303 is specifically configured to: input the validation set into the candidate coal quality analysis model to obtain the evaluation index value corresponding to the validation set;

[0082] Optimize the candidate coal quality analysis model based on the evaluation index value corresponding to the validation set to obtain the final coal quality analysis model.

[0083] Optionally, the data sorting module 301 is specifically configured to: perform clustering and grouping on all coal samples according to the ash content value and volatile content value of all coal samples to obtain multiple groups of coal samples;

[0084] Select training set samples, validation set samples, and test set samples from each group of coal samples according to a preset ratio;

[0085] Combine the training set samples of all groups of coal samples to obtain the training set; combine the validation set samples of all groups of coal samples to obtain the validation set; combine the test set samples of all groups of coal samples to obtain the test set.

[0086] Optionally, the coal sample library further includes the index values of the coal quality indexes to be analyzed obtained in advance for each coal sample. The model training module 302 is specifically configured to: select a sample from the training set as the current sample, and establish the spectral matrix and index matrix of the initial coal quality analysis model based on the spectral data of the current sample and the index values of the coal quality indexes to be analyzed corresponding to the current sample;

[0087] Based on the spectral matrix and the index matrix, determine the spectral regression matrix and index regression matrix of the initial coal quality analysis model, and determine the target regression matrix based on the spectral regression matrix and the index regression matrix;

[0088] Update the index matrix according to the spectral regression matrix and the target regression matrix, and optimize the model parameters of the initial coal quality analysis model based on the updated index matrix. Repeat the step of selecting a sample from the training set as the current sample until a candidate coal quality analysis model that meets the convergence condition is obtained.

[0089] Optionally, the model training module 302 is further configured to: calculate the variances of the principal components of the spectral matrix and the principal components of the index matrix, and determine the spectral coefficient matrix, spectral residual matrix, index coefficient matrix, and index residual matrix based on the variances;

[0090] Add the product of the principal components of the spectral matrix and the transposed matrix of the spectral coefficient matrix to the spectral residual matrix to obtain the spectral regression matrix;

[0091] Add the product of the principal components of the index matrix and the transpose matrix of the index coefficient matrix to the index residual matrix to obtain the index regression matrix.

[0092] Optionally, the model training module 302 is further configured to: obtain the correlation between the principal components of the spectral matrix and the principal components of the index matrix, and when the correlation is greater than a preset value, determine a target residual matrix based on the spectral residual matrix and the index residual matrix;

[0093] Determine a target coefficient matrix based on the spectral coefficient matrix and the index coefficient matrix;

[0094] Add the product of the principal components of the spectral matrix and the transpose matrix of the target coefficient matrix to the target residual matrix to obtain the target regression matrix.

[0095] Optionally, the model training module 302 is further configured to: when the accuracy of the target residual matrix is less than a preset accuracy, or when the number of principal components of the index matrix reaches a preset number, update the spectral residual matrix to the spectral matrix and update the target residual matrix to the index matrix.

[0096] The coal quality analysis device provided by the embodiments of the present invention can execute the coal quality analysis model training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. The content not described in detail in this embodiment can be referred to the description in any method embodiment of the present invention.

[0097] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Refer to Figure 4 which shows a schematic structural diagram of a computer system 12 of an electronic device suitable for implementing the embodiments of the present invention. Figure 4 The illustrated electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0098] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0099] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0100] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 4 not shown, commonly referred to as a "hard disk drive"). Although Figure 4 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data media interfaces. The memory 28 can include at least one program product having a set (such as at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.

[0101] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in the memory 28. Such program modules 42 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present invention.

[0102] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Additionally, in this embodiment of the electronic device 12, the display 24 does not exist as an independent entity but is embedded in the mirror. When the display surface of the display 24 is not displaying, the display surface of the display 24 visually merges with the mirror surface. And, the electronic device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As Figure 4As shown, network adapter 20 communicates with other modules of electronic device 12 via bus 18. It should be understood that although Figure 4 not shown in Figure 4 , other hardware and / or software modules may be used in conjunction with electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0103] Processing unit 16 executes various functional applications and coal quality analysis by running programs stored in system memory 28. For example, it implements a coal quality analysis model training method provided by an embodiment of the present invention: If the initial coal quality analysis model does not meet the pre-set convergence condition, then a training set, a validation set, and a test set of the initial coal quality analysis model are sorted out from all the coal samples in the pre-acquired coal sample library; wherein, the coal sample library at least includes spectral data, coal quality indexes to be analyzed, ash content values, and volatile content values of each coal sample; according to the training set, optimization is performed through leave-one-out cross-validation to obtain a candidate coal quality analysis model that meets the convergence condition; based on the validation set, the candidate coal quality analysis model is optimized to obtain a final coal quality analysis model.

[0104] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a coal quality analysis model training method provided by all embodiments of the present invention: If the initial coal quality analysis model does not meet the pre-set convergence condition, then a training set, a validation set, and a test set of the initial coal quality analysis model are sorted out from all the coal samples in the pre-acquired coal sample library; wherein, the coal sample library at least includes spectral data, coal quality indexes to be analyzed, ash content values, and volatile content values of each coal sample; according to the training set, optimization is performed through leave-one-out cross-validation to obtain a candidate coal quality analysis model that meets the convergence condition; based on the validation set, the candidate coal quality analysis model is optimized to obtain a final coal quality analysis model. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0105] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0106] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0107] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0108] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for training a coal quality analysis model, characterized in that, Including: If the initial coal quality analysis model does not meet the preset convergence condition, a training set, a validation set, and a test set of the initial coal quality analysis model are sorted out from all the coal samples in the pre-obtained coal quality sample library; wherein, the coal quality sample library at least includes spectral data, coal quality indexes to be analyzed, ash content values, and volatile content values of each coal sample; wherein, the coal quality indexes to be analyzed include carbon content and calorific value. Sorting out a training set, a validation set, and a test set of the initial coal quality analysis model from all the coal samples in the pre-obtained coal quality sample library includes: Clustering and grouping all the coal samples according to the ash content values and volatile content values of all the coal samples to obtain multiple groups of coal samples; screening out training set samples, validation set samples, and test set samples from each group of coal samples according to a preset ratio; combining the training set samples of all groups of coal samples to obtain the training set; combining the validation set samples of all groups of coal samples to obtain the validation set; combining the test set samples of all groups of coal samples to obtain the test set. Optimizing through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition. Optimizing the candidate coal quality analysis model based on the validation set to obtain a final coal quality analysis model; optimizing the candidate coal quality analysis model based on the validation set to obtain a final coal quality analysis model, including: inputting the validation set into the candidate coal quality analysis model to obtain the evaluation index value corresponding to the validation set; optimizing the candidate coal quality analysis model based on the evaluation index value corresponding to the validation set to obtain a final coal quality analysis model.

2. The method according to claim 1, wherein The coal quality sample library further includes the index values of the coal quality indexes to be analyzed pre-obtained corresponding to each coal sample. Optimizing through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition includes: Selecting one sample from the training set as the current sample, and establishing a spectral matrix and an index matrix of the initial coal quality analysis model based on the spectral data of the current sample and the index values of the coal quality indexes to be analyzed corresponding to the current sample; wherein, the index matrix includes the carbon content value and the calorific value of the current sample. Based on the spectral matrix and the index matrix, determining a spectral regression matrix and an index regression matrix of the initial coal quality analysis model, and determining a target regression matrix based on the spectral regression matrix and the index regression matrix. Updating the index matrix according to the spectral regression matrix and the target regression matrix, and optimizing the model parameters of the initial coal quality analysis model based on the updated index matrix, and repeating the step of selecting one sample from the training set as the current sample until a candidate coal quality analysis model that meets the convergence condition is obtained.

3. The method according to claim 2, wherein Based on the spectral matrix and the index matrix, determining a spectral regression matrix and an index regression matrix of the initial coal quality analysis model, including: Calculate the variances of the principal components of the spectral matrix and the principal components of the index matrix, and determine the spectral coefficient matrix, spectral residual matrix, index coefficient matrix, and index residual matrix based on the variances; Add the product of the principal components of the spectral matrix and the transposed matrix of the spectral coefficient matrix to the spectral residual matrix to obtain the spectral regression matrix; Add the product of the principal components of the index matrix and the transposed matrix of the index coefficient matrix to the index residual matrix to obtain the index regression matrix.

4. The method according to claim 3, wherein Determine the target regression matrix based on the spectral regression matrix and the index regression matrix, including: Obtain the correlation between the principal components of the spectral matrix and the principal components of the index matrix. When the correlation is greater than a preset value, determine the target residual matrix based on the spectral residual matrix and the index residual matrix; Determine the target coefficient matrix based on the spectral coefficient matrix and the index coefficient matrix; Add the product of the principal components of the spectral matrix and the transposed matrix of the target coefficient matrix to the target residual matrix to obtain the target regression matrix.

5. The method according to claim 4, characterized in that, Update the index matrix according to the spectral regression matrix and the target regression matrix, including: When the accuracy of the target residual matrix is less than the preset accuracy, or when the number of principal components of the index matrix is less than the preset number, update the spectral residual matrix to the spectral matrix and update the target residual matrix to the index matrix.

6. A device for training a coal quality analysis model, characterized in that, Include: A data sorting module, configured to sort out the training set, validation set, and test set of the initial coal quality analysis model from all the coal samples in the pre-acquired coal quality sample library if the initial coal quality analysis model does not meet the pre-set convergence condition; wherein, the coal quality sample library at least includes the spectral data, coal quality indexes to be analyzed, ash content, and volatile content of each coal sample; wherein, the coal quality indexes to be analyzed include carbon content and calorific value; Sort out the training set, validation set, and test set of the initial coal quality analysis model from all the coal samples in the pre-acquired coal quality sample library, including: Cluster and group all the coal samples according to the ash content and volatile content of all the coal samples to obtain multiple groups of coal samples; screen out the training set samples, validation set samples, and test set samples from each group of coal samples according to a preset ratio; combine the training set samples of all groups of coal samples to obtain the training set; combine the validation set samples of all groups of coal samples to obtain the validation set; combine the test set samples of all groups of coal samples to obtain the test set; A model training module, configured to perform optimization through leave-one-out cross-validation according to the training set to obtain a candidate coal quality analysis model that meets the convergence condition; A model optimization module, configured to optimize the candidate coal quality analysis model based on the validation set to obtain the final coal quality analysis model.

7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the coal quality analysis model training method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the coal quality analysis model training method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Coal quality measurement method based on laser-induced breakdown spectroscopy and near infrared spectroscopy information fusion

    CN111044503A

  • New method for quickly establishing coal component quantitative prediction model based on small sample size

    CN116380875A