A method for determining the content of major and minor elements in iron ore using VI-BP-ANN assisted LIBS.

By combining random forest and backpropagation artificial neural network, the input variables of the BP-ANN model were optimized, which solved the problems of high uncertainty and large error in the determination of elemental content in iron ore by laser-induced breakdown spectroscopy. This resulted in higher determination accuracy and lower overfitting risk, improving the stability and efficiency of elemental analysis in iron ore.

CN115541561BActive Publication Date: 2026-04-07SHANGHAI ENTRY-EXIT INSPECTION & QUARANTINE BUREAU IND PROD & RAW MATERIALS TESTING TECH CENT
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing laser-induced breakdown spectroscopy techniques suffer from high uncertainty and large errors when determining the elemental content of iron ore. In particular, the BP-ANN model is prone to overfitting under high-dimensional data, resulting in low accuracy of quantitative analysis.

Method used

By combining random forest and backpropagation artificial neural network, and through variable importance measurement and threshold control, the input variables of the BP-ANN model are optimized, redundant information is reduced, and the model stability and accuracy are improved.

Benefits of technology

While reducing analysis time, it significantly improved the accuracy of determining the content of total iron, calcium, magnesium, aluminum and silicon in iron ore, reduced the risk of overfitting, and enhanced the quantitative analysis capability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115541561B_ABST
    Figure CN115541561B_ABST
Patent Text Reader

Abstract

This invention discloses a method for determining the elemental content of iron ore using variable importance-artificial neural network-assisted laser-induced breakdown spectroscopy. The method includes the following steps: S1: Collect LIBS spectral data and elemental content data from at least 13 batches of iron ore in at least four categories. Use radiometric analysis (RF) to measure the importance of LIBS spectral features. A variable importance threshold is used to optimize the input variables of the BP-ANN model. The variable importance threshold and the number of neurons in the BP-ANN model are optimized using the determination coefficients and root mean square error of five-fold cross-validation to establish a VI-BP-ANN model. S2: Using the VI-BP-ANN model from S1, input the LIBS spectral data. The model ranks the features according to variable importance and calculates the elemental content in the iron ore. This invention's detection method can rapidly detect the total iron content, calcium content, magnesium content, aluminum content, and silicon content in iron ore.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for determining the content of major and minor elements in iron ore using VI-BP-ANN assisted LIBS. Background Technology

[0002] The contents of total iron (TFe), CaO, MgO, Al2O3, and SiO2 are important indicators for evaluating the quality of iron ore, affecting its trade price and blast furnace ironmaking process. Traditional analytical methods include titration, gravimetric analysis, spectrophotometry, atomic absorption spectrometry, and wavelength dispersive X-ray fluorescence spectrometry (WD-XRF), which typically require complex laboratory pretreatment, are energy-intensive, and have long testing cycles. Laser-induced breakdown spectroscopy (LIBS) is an atomic emission spectrometry technique that uses high-energy laser pulses to bombard the surface of a material to obtain the elemental composition and content of the analyte. It offers advantages such as in-situ accuracy, speed, and no need for complex sample preparation, and has attracted widespread attention in the field of iron ore composition analysis. However, due to the matrix effects of different types of iron ore, laser energy fluctuations, and uncertainties in plasma spatiotemporal evolution, quantitative analysis of elemental content in iron ore using LIBS faces severe challenges of high measurement uncertainty and large errors.

[0003] Due to spectral interference, self-absorption, and matrix effects, calibration curves based on the intensity of a single spectral line cannot accurately reflect the elemental content of iron ore, resulting in low analytical accuracy. Combining multivariate regression analysis with LIBS is a solution to overcome matrix effects in the quantitative analysis of iron ore using LIBS. For the analysis of elemental content in iron ore, some multivariate regression analysis methods have been used in conjunction with LIBS, such as random forest (RF), partial least squares (PLS), support vector machine (SVM), and principal component regression (PCR). These studies generally use the full spectrum as the input variable, and the number of iron ore samples studied is relatively small, which easily leads to model overfitting and poor quantitative accuracy, limiting the practical application of the models.

[0004] Backpropagation artificial neural networks (BP-ANN), as an emerging multivariate analysis method, offer significant advantages for high-dimensional data. However, using full-spectral data from iron ore LIBS spectroscopy as input to a BP-ANN model can easily lead to the curse of dimensionality, causing overfitting and impacting the predictive ability and analysis time of quantitative regression models. Variable importance methods measure the impact of each input variable on the overall predictive performance of the model through data permutation. This method not only measures the importance score of each variable but also assesses the relationships between variables. By calculating the variable importance of input spectral features and selecting a small subset of feature variables that cover the original spectral information as model input, the interference of redundant variables on the model can be effectively reduced. Currently, there are no reports on combining variable importance with BP-ANN for quantitative elemental analysis of iron ore LIBS spectroscopy. Summary of the Invention

[0005] The technical problem this invention aims to solve is to overcome the application difficulties of existing laser-induced breakdown spectroscopy (LIBS) techniques in determining the elemental content of iron ore, which suffer from high uncertainty and large errors. This invention provides a variable importance-backpropagation artificial neural network (VI-BP-ANN) assisted method for determining the content of major and minor elements in iron ore using LIBS spectroscopy. This method can rapidly detect the total iron content, calcium content (calculated as CaO), magnesium content (calculated as MgO), aluminum content (calculated as Al2O3), and silicon content (calculated as SiO2) in iron ore with high accuracy.

[0006] The present invention solves the above-mentioned technical problems through the following technical solutions.

[0007] Backpropagation artificial neural networks (BP-ANN) and random forests (RF) are two representative models in machine learning, often used independently in related research. However, in practical applications, due to spectral interference, self-absorption, and matrix effects, BP-ANN suffers from severe overfitting and low accuracy in quantitative analysis when analyzing small-sample, high-dimensional laser-induced breakdown spectroscopy (LIBS) data.

[0008] In this invention, the applicant creatively combines Random Forest (RF) with Backpropagation Artificial Neural Network (BP-ANN) and applies it to the quantitative analysis of elemental content in imported iron ore. This application fully utilizes the different advantages of Random Forest (high tolerance to noise) and BP-ANN (powerful self-learning ability), cleverly combining them. The Random Forest measures the importance of features and ranks them according to importance. An importance threshold is used to control the variables input to the BP-ANN, effectively filtering redundant information while preserving the original features. Accuracy is greatly improved by using fewer original LIBS spectral features as input, and analysis time is effectively reduced. Compared to other machine learning methods, it effectively reduces overfitting and improves the stability of model analysis.

[0009] This invention provides a method for determining the content of major and minor elements in iron ore by laser-induced breakdown LIBS spectroscopy assisted by a variable importance-backpropagation artificial neural network (VI-BP-ANN), comprising the following steps:

[0010] S1: Collect LIBS spectral data and elemental content data of iron ore from at least 13 batches in at least 4 categories, with the elemental content measured in wt%. Use Random Forest (RF) to measure the importance of LIBS spectral features. Variable importance thresholds are used to optimize the input variables of the backpropagation-artificial neural network (BP-ANN) model. The variable importance thresholds and the number of neurons in the BP-ANN model are determined using the coefficients R0 through 5-fold cross-validation (5-CV). 2 To optimize the VI-BP-ANN model, the following steps are included: (The steps are described below.)

[0011] 1) Preprocess the LIBS spectral data; 2) Optimize ntree (number of trees in the forest) and mtry (number of feature variables when the regression tree branches at a node) based on the OOB error rate, and establish an RF (Random Forest) model; 3) Use the RF model to score and rank the feature variables of the LIBS spectrum based on their importance; 4) Optimize the threshold for the importance of the variables and the number of neurons in the BP-ANN model, and use the threshold to control the input variables of the BP-ANN model; Use 5-CV (five-fold crossover) to examine the impact of the threshold and the number of neurons on the VI-BP-ANN model, and obtain the evaluation index, which is RMSECV (root mean square error of cross-validation) and R. 2 (Goodness of fit), establish a VI-BP-ANN (Variable Importance-Backpropagation Artificial Neural Network) model; 5) Select the threshold and the number of neurons to train the VI-BP-ANN model;

[0012] S2: Using the VI-BP-ANN model described in S1, input LIBS spectral data. The model sorts the features according to the importance of variables and reads the spectral features according to the threshold to calculate the element content in the iron ore.

[0013] The elements include a primary element and / or a secondary element; the primary element is iron; the secondary elements include one or more of calcium, magnesium, aluminum and silicon.

[0014] In this invention, preferably, the iron content is the total iron content.

[0015] In this invention, preferably, the calcium content is expressed as CaO content.

[0016] In this invention, preferably, the magnesium content is expressed as MgO content.

[0017] In this invention, preferably, the aluminum content is expressed as Al2O3 content.

[0018] In this invention, preferably, the silicon content is expressed as SiO2 content.

[0019] In a preferred embodiment, in S1, when the element content is the total iron content, the number of iron ore batches for each category is at least 20.

[0020] In S1, preferably, the iron ore originates from one or more of Australia, South Africa, Kazakhstan, and Chile.

[0021] In S1, preferably, the iron ore is one or more of the following: FMG mixed powder, Mac powder, PB powder, Newman lump, PB lump, Yandi powder, Newman powder, South African concentrate, Australian concentrate, mixed powder, South African powder, pellets, 52 iron lump ore, 71.52 iron lump ore, iron ore, and Atacama pelletized iron concentrate.

[0022] In S1, the elemental content of the iron ore can generally be obtained through chemical analysis.

[0023] In S1, the total iron content of the iron ore can generally range from 25 to 72 wt%.

[0024] In S1, the silicon dioxide content of the iron ore is generally in the range of 1.03 to 15.66 wt%, calculated as SiO2.

[0025] In S1, the alumina content of the iron ore is generally in the range of 0.2 to 3.06 wt%, calculated as Al2O3.

[0026] In S1, the calcium oxide content of the iron ore is generally in the range of 0.016 to 1.768 wt%, calculated as CaO.

[0027] In S1, the magnesium oxide content of the iron ore is generally in the range of 0.034 to 9.9 wt%, calculated as MgO.

[0028] In step S1, preferably, before measuring the LIBS spectral data and the elemental content data, a sample pretreatment step is performed on the iron ore; the sample pretreatment step includes: sequentially aggregating and pressing the iron ore to form a cake-shaped iron ore, and purging the cake-shaped iron ore.

[0029] Preferably, the pressure of the tablet is 30t.

[0030] Preferably, the tableting time is 30 seconds.

[0031] Preferably, the thickness of the cake-shaped iron ore is 5 mm.

[0032] In S1, preferably, the LIBS spectral data is obtained by spectral acquisition using a LIBS system.

[0033] Preferably, the laser ablation source of the LIBS system is a laser; more preferably, the laser is a Q-switched pulsed Nd:YAG laser; even more preferably, the maximum emission wavelength of the laser is 1064 nm, the repetition frequency of the laser is 5 Hz, and the laser energy of the laser is 30 MJ.

[0034] Preferably, the spectral acquisition step is as follows: a laser beam is vertically focused onto the surface of the iron ore through a lens, and the generated plasma radiation is coupled to a spectrometer through an optical fiber; more preferably, the focal length of the lens is 50 mm; more preferably, the measurement range of the spectrometer is 190-950 nm; more preferably, the radiation delay time is 1 μs to reduce the strong continuous background caused by bremsstrahlung in the early stage of plasma formation and improve the spectral intensity and spectral signal-to-background ratio.

[0035] Preferably, the spectral acquisition is performed in the form of a 5×5 matrix; the 5×5 matrix form means that 5 laser pulses are accumulated at each position of the matrix to reduce the fluctuation of laser energy and reduce the influence of sample heterogeneity on LIBS spectral measurement.

[0036] In step 1) of S1, preferably, the preprocessing method includes one or more of smoothing, multivariate scattering correction, and normalization; more preferably, it is a combination of smoothing, multivariate scattering correction, normalization, normalization and multivariate scattering correction, or normalization and smoothing; even more preferably, it is normalization. The normalization is used to reduce the impact of pulse fluctuations and unstable ablation of the sample on the spectral data. The smoothing is mainly based on K-order polynomial fitting of data points within a certain length window, which can improve the smoothness of the spectrum and reduce noise interference. The multivariate scattering correction (MSC) can effectively eliminate spectral differences caused by different scattering levels and enhance the correlation between the spectrum and the data.

[0037] In step 2) of S1, ntree is the number of trees in the forest; mtry is the number of feature variables when the regression tree branches at a node.

[0038] In step 2) of S1, the LIBS spectral number M of the iron ore is equal to 12814; M is both the number of variables and the dimension of the input data, and the optimal branch is selected according to the optimal principle.

[0039] Preferably, the mtry is smaller than the M.

[0040] In step 2) of S1, preferably, when the element is the main element, the mtry is M / 3, M / 4, M / 5, M / 6, M / 7 or M / 8; more preferably, when the mtry is M / 4 and the ntree is 100, the OOB error rate is 0.089.

[0041] In step 2) of S1, preferably, when the element is the sub-element, the mtry is 0.5√M to 8√M; the ntree is 100 to 800; wherein, more preferably, the mtry is 0.5√M·x, 1≤x≤16, where x is an integer; more preferably, the ntree is 100·y, 1≤y≤8, where y is an integer; even more preferably, when the mtry is 5.5√M and the ntree is 200, the OOB error rate of the silicon element is 0.1358; More preferably, when mtry is 1.5√M and ntree is 200, the OOB error rate of aluminum is 0.1030; More preferably, when mtry is 1.5√M and ntree is 200, the OOB error rate of calcium is 0.0886; More preferably, when mtry is 2.5√M and ntree is 600, the OOB error rate of magnesium is 0.0263.

[0042] In step 3) of S1, variable importance refers to the influence of the input variable on the quantitative result. By measuring the importance of variables, variables with low importance can be removed, thereby obtaining a simple and effective model.

[0043] In step 3) of S1, preferably, when the element is the principal element, the importance of the variable is 0 to 0.03 when the systematic wavelength range of the LIBS spectrum is 200-500 nm; the importance of the variable is 0 when the systematic wavelength range of the LIBS spectrum is 501-759 nm; and the importance of the variable is 0 to 0.1 when the systematic wavelength range of the LIBS spectrum is 760-770 nm.

[0044] In step 4) of S1, a larger threshold results in fewer input variables, effectively reducing the impact of redundant or irrelevant variables as features on the construction of the BP-ANN model. However, too few input variables may lead to the loss of useful feature information. Choosing an appropriate variable importance threshold and determining the optimal input variables has a crucial impact on the quality of model construction.

[0045] In step 4) of S1, if the number of neurons is too small, the model will not fit properly and will not be able to establish an effective connection between the input spectral features and the output content. If the number of neurons is too large, it will be easy to fall into overfitting, and the quantitative accuracy will be affected when predicting the TFe content of unknown samples.

[0046] In step 4) of S1, in a preferred embodiment, when the element is the principal element, the threshold is 0.001, the number of neurons is 42, the number of variables is 58, and the importance of the variables is 0.9238.

[0047] In step 4) of S1, in another preferred embodiment, when the element is the principal element, the threshold is 0.0005, the number of neurons is 48, the number of variables is 80, and the importance of the variables is 0.9382.

[0048] In step 4) of S1, in another preferred embodiment, when the element is the principal element, the threshold is 0.0002, the number of neurons is 50, the number of variables is 170, and the importance of the variables is 0.9689.

[0049] In step 4) of S1, in another preferred embodiment, when the element is the principal element, the threshold is 0.0001, the number of neurons is 48, the number of variables is 262, and the importance of the variables is 0.9820.

[0050] In step 4) of S1, in another preferred embodiment, when the element is the principal element, the threshold is 0, the number of neurons is 36, the number of variables is 3517, and the importance of the variables is 1.

[0051] In step 4) of S1, the LIBS spectrum of the iron ore contains too much noise and redundant variables. By controlling the threshold input to the BP-ANN model, the redundant variables in the spectrum can be reduced.

[0052] In step 4) of S1, preferably, the transfer function of the neuron is the Tanh activation function.

[0053] In step 4) of S1, preferably, the construction of the VI-BP-ANN model includes the optimization of the number of neurons and input variables, and the variables input to the BP-ANN model are controlled using variable importance thresholds.

[0054] In step 4) of S1, the five-fold cross-validation (5-CV) involves randomly dividing the calibration set into five parts, with one part serving as the validation set and the remaining parts serving as the training set for each iteration, in order to make full use of the dataset to examine the model's performance.

[0055] In step 4) of S1, the R 2 R represents the goodness of fit; RMSECV is the root mean square error of cross-validation. 2 When the maximum value and the minimum value of RMSECV are achieved, the model with the optimal number of neurons and variable importance thresholds is saved, which is the VI-BP-ANN model.

[0056] In step 4) of S1, preferably, the R 2 The calculation formulas for RMSECV are as follows:

[0057]

[0058]

[0059] In this invention, preferably, when the element is the main element, for the iron content, when the threshold is 0.0002, the number of neurons is 50, and the R... 2 The value is 0.9674, and the RMSECV is 0.3089 wt%.

[0060] In this invention, preferably, when the element is the sub-element, for the silicon element, the variable importance is 0 to 0.001 or the full spectrum; preferably, the variable importance is 0.00001, the number of variables is 1683, the number of neurons is 32, and the RMSECV is 0.7639 wt%.

[0061] In this invention, preferably, when the element is the sub-element, for the aluminum element, the variable importance is 0 to 0.001 or the full spectrum; preferably, the variable importance is the full spectrum, the number of variables is 12814, the number of neurons is 50, and the RMSECV is 0.1763wt%.

[0062] In this invention, preferably, when the element is the sub-element, for the calcium element, the variable importance is 0 to 0.001 or the full spectrum; preferably, the variable importance is 0.000001, the number of variables is 3762, the number of neurons is 42, and the RMSECV is 0.1125 wt%.

[0063] In this invention, preferably, when the element is the sub-element, for the magnesium element, the variable importance is 0 to 0.001 or the full spectrum; preferably, the variable importance is 0.000002, the number of variables is 1616, the number of neurons is 46, and the RMSECV is 0.2679 wt%.

[0064] In this invention, the random forest can calculate and rank the importance of each variable. It uses data permutation to evaluate the impact of each variable on the overall predictive performance of the quantitative model, and finally gives an importance score for each variable. Since true variables are more important than noise, the higher the variable importance score, the greater its contribution to the model, and vice versa. Therefore, variable importance can be used as a feature selection criterion, allowing a small subset of variables that cover most of the input information to be used for modeling, thereby reducing overfitting.

[0065] In this invention, the importance of variables is calculated through the random forest, and a small subset of characteristically important variables are used as input to the BP-ANN model, which can improve the quantitative analysis results of total iron in iron ore by BP-ANN.

[0066] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0067] The reagents and raw materials used in this invention are all commercially available.

[0068] The positive and progressive effects of this invention are as follows:

[0069] 1. This application fully utilizes the different advantages of Random Forest (high tolerance to noise) and BP-ANN (powerful self-learning ability), and cleverly combines them. RF has a good tolerance to noise. It is used to measure the importance of spectral features and rank them according to their importance without changing the original features. Then, the variables input to BP-ANN are controlled by variable importance thresholds.

[0070] 2. The method proposed in this application can identify the importance of each feature by relying on the inherent relationship between data and labels while preserving the original features and variables. Then, different weights are reassigned to the input variables within the BP-ANN network, thereby achieving effective linkage between RF and BP-ANN and dual feature variable importance assessment, avoiding the loss of effective information. While reducing analysis time and overfitting, it effectively filters out interference and redundant information, improving the accuracy of the model.

[0071] 3. The method in this application achieves significantly improved accuracy and effectively reduces analysis time while using fewer raw LIBS spectral features as input. Compared with other machine learning methods, it can effectively reduce overfitting and improve the stability of model analysis. Attached Figure Description

[0072] Figure 1 This is a flowchart of establishing the VI-BP-ANN model in Example 1.

[0073] Figure 2 This is a graph showing the OOB error rate under different ntree and mtry values ​​in Example 1.

[0074] Figure 3 This is a variable importance diagram of the LIBS spectrum of iron ore in Example 1.

[0075] Figure 4 The figures show the prediction results of Example 1 based on the VI-BP-ANN model and Comparative Example 1 based on the BP-ANN model.

[0076] Figure 5 This is a comparison chart of the prediction performance of different models in Example 1 and Comparative Example 1; where, Figure 5 a is the 5-CV internal validation diagram; Figure 5 b is the external verification graph.

[0077] Figure 6 A comparison diagram of oxides treated with different pretreatment methods; where, Figure 6 a is a comparison chart of SiO2 using different pretreatment methods; Figure 6 b is a comparison diagram of Al2O3 using different pretreatment methods; Figure 6c is a comparison graph of CaO using different pretreatment methods; Figure 6 d is a comparison chart of MgO using different pretreatment methods.

[0078] Figure 7 The relationship between different ntree and mtry values ​​of different oxides and the OOB error rate; where, Figure 7 a represents the relationship between different ntree and mtry values ​​of SiO2 and the OOB error rate; Figure 7 b represents the relationship between different ntree and mtry values ​​of Al2O3 and the OOB error rate; Figure 7 c represents the relationship between different ntree and mtry values ​​of CaO and the OOB error rate; Figure 7 d represents the relationship between different ntree and mtry values ​​of MgO and the OOB error rate.

[0079] Figure 8 The results show the predictions for oxides using the VI-BP-ANN model; where, Figure 8 a represents the prediction result of the VI-BP-ANN model for SiO2; Figure 8 b represents the prediction result of the VI-BP-ANN model for Al2O3; Figure 8 c represents the prediction result of the VI-BP-ANN model for CaO; Figure 8 d represents the prediction result of MgO by the VI-BP-ANN model.

[0080] Figure 9 This is a comparison chart showing the performance of different models in predicting oxides in iron ore based on 5-CV validation; among them, Figure 9 a is a comparison of the performance of different models in predicting SiO2 in iron ore based on 5-CV validation; Figure 9 b is a comparison chart of the performance of different models in predicting Al2O3 in iron ore based on 5-CV validation; Figure 9 c is a comparison chart of the performance of different models in predicting CaO in iron ore based on 5-CV validation; Figure 9 d is a comparison of the performance of different models in predicting MgO in iron ore based on 5-CV validation. Detailed Implementation

[0081] The present invention is further illustrated below by way of embodiments, but the invention is not limited to the scope of the embodiments described herein. Experimental methods in the following embodiments that do not specify specific conditions were performed according to conventional methods and conditions, or as selected according to the product instructions.

[0082] Example 1

[0083] Figure 1 This is a flowchart of the process for establishing the VI-BP-ANN model in this embodiment.

[0084] S1:

[0085] Eighty iron ore samples from four Australian brands were selected. Samples 1-20 were FMG blend powder (FMG); samples 21-32 were MAC powder (MAC); samples 33-56 were PB powder (PB); and samples 57-80 were Newman ore. The total iron content of these four types of iron ore ranged from 40% to 65%. The iron ore was processed into powder, placed in plastic rings, and compressed using a tablet press at a pressure of 30t for 30s to obtain iron ore flakes approximately 5mm thick.

[0086] The total iron content of iron ore samples was determined using the chemical analysis method according to national standard GB / T 6730.5-2007. The results are shown in Table 1 and are used as reference values. The determination range of total iron content (TFe, wt%) is 25–72%.

[0087] Table 1 shows the total iron content of different types of iron ore samples. α is the sample used for testing.

[0088] Table 1

[0089]

[0090]

[0091] LIBS spectral data were acquired using a commercial LIBS system (Chemreveal 3764, TSI). The laser ablation source was a Q-switched pulsed Nd:YAG laser with a maximum emission wavelength of 1064 nm, a repetition rate of 5 Hz, and a laser energy of 30 MJ. The prepared iron ore samples were placed on a manually adjustable XYZ platform, the position of which could be observed by a camera. The laser beam was vertically focused onto the sample surface through a 50 mm focal length lens to generate plasma. The generated plasma radiation was coupled to the spectrometer via optical fiber. The spectrometer's measurement range was 190-950 nm, and the radiation delay time was 1 μs. LIBS spectral acquisition was performed in a 5×5 matrix, with 5 laser pulses accumulated at each position of the matrix to obtain the spectrum. The average of the 12 matrices formed for each iron ore sample constituted one analytical spectrum. Eighty iron ore samples yielded eighty analytical spectra, which were randomly divided into a calibration set (64 iron ore samples) and a test set (16 iron ore samples) at a ratio of 4:1. The calibration set samples were used to calibrate the model, and the test set samples were used to test the model. The entire quantitative analysis of total iron content in iron ore was implemented in Python 3.7 (sklearn 2.3.1).

[0092] Iron ore samples are composed of various chemical elements, such as Fe, Ca, Mg, Si, Al, P, S, Na, and K. The related LIBS spectra consist of hundreds of atomic lines, with a data volume as high as 12,814. Due to the different elemental compositions in iron ore samples from different brands, the intensity of the emission lines also varies significantly.

[0093] Table 2 shows the range of major elemental contents of iron ore from different brands. Besides TFe content, the contents of SiO2, Al2O3, CaO, and MgO also differ among the four brands of iron ore. SiO2 content ranges from 2.54% to 6.19%, Al2O3 content from 1.228% to 2.85%, CaO content from 0.001% to 0.196%, and MgO content from 0.041% to 0.162%.

[0094] Table 2

[0095] Brand TFe <![CDATA[SiO2]]> <![CDATA[Al2O3]]> CaO MgO FMG 57.84-58.69 4.880-6.190 2.560-2.780 0.041-0.171 0.057-0.116 MAC 60.28-61.38 2.540-5.460 2.140-2.540 0.009-0.054 0.073-0.149 PB 61.33-62.26 3.220-4.110 2.120-2.850 0.001-0.196 0.067-0.162 Newman 62.41-64.18 3.043-4.240 1.228-1.660 0.001-0.137 0.041-0.131

[0096] Principal Component Analysis (PCA), a powerful tool for multivariate analysis and searching for differences between spectra, was used to assess matrix variations in iron ore. When the number of factors was set to 2, the cumulative contribution rate reached 99.988%. Based on the two-dimensional PC1 and PC2 plots of the iron ore samples, it was found that the two-dimensional scatter plots of iron ore samples from different brands were distributed in different regions with varying degrees of dispersion. Furthermore, there was partial overlap between Newman, PB, and MAC. Quantitative analysis of TFe content in iron ore samples using LIBS is easily affected by matrix effects and interference from surrounding emission lines, leading to reduced accuracy.

[0097] In this embodiment, Random Forest (RF) is used to measure the importance of each variable in the LIBS spectrum of iron ore.

[0098] Based on the above spectral data and total iron content data, a VI-BP-ANN model is established, which includes the following steps:

[0099] 1) Optimize the parameters ntree and mtry of the random forest based on the OOB error rate; Figure 2 The table shows the OOB error rate for different ntree and mtry values, displaying the trend of OOB error rate with ntree variation when mtry is M / 3, M / 4, M / 5, M / 6, M / 7, and M / 8 (M is the LIBS spectrum number of iron ore, 12814). Figure 2As can be seen, the OOB error rate is lowest when ntree equals 40 and mtry equals M / 5. However, the model's performance does not show stable performance because the OOB error rate increases, then decreases, and then increases again as ntree increases. In this respect, ntree equals 100, which is more suitable for searching for local minima of different mtry values. Therefore, the OOB error rate is minimized to 0.089 when ntree equals 100 and mtry equals M / 4, and the RF model is built based on this.

[0100] 2) Use the Random Forest (RF) model to score and rank the importance of LIBS feature variables for iron ore. Although the LIBS spectrum of iron ore has 12,184 features, there are too many noisy and redundant variables in the LIBS spectrum, and it is not necessary to score the importance of all features in the Random Forest variable importance ranking. The input of feature variables into the BP-ANN model can be controlled by setting a threshold for variable importance, thereby reducing redundant features in the spectrum.

[0101] The higher the importance of a variable, the more important the feature, and vice versa. Figure 3 The graph shows the importance of variables in the LIBS spectrum of iron ore. It indicates that the importance of all variables is in the range of 0-0.10, but not all variables are important. Most variables have an importance of 0. The variables with higher importance are mainly concentrated in the wavelength range of 200-500 nm, which is where iron emission lines are dense. The wavelength range of 760-770 nm shows a feature with higher variable importance, which is mainly related to the characteristic emission lines of K. The peak shape here is relatively independent and is not affected by other spectral lines.

[0102] Table 3 shows the number of variables and cumulative variable importance at different thresholds. When the threshold is 0.0001, there are 262 variables and the variable importance is 0.98204; when the threshold is set to 0, there are 3517 variables and the variable importance is 1.

[0103] Table 3

[0104]

[0105]

[0106] 3) To determine the thresholds for variable importance and the number of neurons, the impact on the model was examined using five-fold cross-validation (5-CV) at different thresholds (0.001, 0.0005, 0.0002, 0.0001, and 0) and the number of VI-BP-ANN neurons (30–60). The results were measured using RMSECV and R0. 2 As an evaluation metric, the goodness of fit R is calculated. 2 The root mean square error (RMSECV) of cross-validation; R2 The calculation formulas for RMSECV are as follows:

[0107]

[0108]

[0109] Table 4 shows the prediction results under different thresholds and numbers of neurons. When the threshold for variable importance is 0, the number of variable features is 3517, the optimal number of neurons is 36, the RMSECV is 0.3906wt%, and the R² value is [missing value]. 2 The R² value was 0.9474, and the model optimization time was 26 min 53 s. As the variable importance threshold continued to increase, the number of input variables decreased, and the model optimization time was significantly shortened. When the variable importance threshold was 0.0002 and the optimal number of neurons was 50, the calibrated model achieved the best performance, with an R² value of 0.3089 wt% and an R² value of [missing value]. 2 The value was 0.9674, and the required optimization time was only 24 seconds. Compared with using the full spectrum as the input variable, RMSECV decreased from 0.3427 wt% to 0.3089 wt%, and R0 was significantly lower. 2 Increasing the importance value from 0.9625 to 0.9674 significantly improved the performance of the calibration model, while simultaneously reducing the optimization time from 184 min 25 s to 24 s, exhibiting an exponential decrease. However, when the variable importance threshold exceeded 0.0002, the RMSECV continued to increase, and R... 2 The importance of variables continues to decline. This means that as the number of variable features decreases, some important information may be lost. Therefore, a variable importance threshold of 0.0002 and 50 neurons were used to construct the VI-BP-ANN model.

[0110] Table 4. Model prediction results under different variable importance thresholds.

[0111] Threshold for variable importance Optimal number of neurons RMSECV (wt%) <![CDATA[R 2 ]]> Modeling time Full spectrum 38 0.3427 0.9625 184min25s 0.001 42 0.3814 0.9513 10s 0.0005 48 0.3364 0.9641 12s 0.0002 50 0.3089 0.9674 24s 0.0001 48 0.3319 0.9615 39s 0 36 0.3906 0.9474 26min53s

[0112] 4) Compared to the original spectrum, a large amount of spectral noise and redundant information was effectively removed. After selection using the variable importance threshold (0.0002), only 170 features covering most of the effective information in the original spectrum were retained across the entire spectrum. The selected features, in addition to the target element Fe for quantitative analysis, also included emission lines of other elements such as Si, Al, Ca, Mg, Mn, Ti, Zr, Li, Na, and K, indicating that the quantitative calibration results for Fe were contributed by a large number of other elements. These selected elemental features are the peaks and the wavelengths surrounding the peaks, meaning that the emission peak regions in the LIBS data contain more valuable spectral information. Meanwhile, partial spectral baselines in the 540-610 nm and 800-970 nm ranges were selected. Since there are almost no characteristic emission lines in these wavelength ranges, these selected features may be more representative of background information related to the instrument and environment. The above analysis shows that variable importance can fully consider the interactions between elements and the spectral baseline. By using fewer variables, a more efficient and simpler model can be established for the quantitative analysis of TFe content in branded iron ore using LIBS.

[0113] Step S2:

[0114] Using the VI-BP-ANN model described in S1, input the LIBS spectral data for model prediction, and calculate the total iron content in the iron ore.

[0115] Comparative Example 1

[0116] Using the full spectrum as a threshold for variable importance, a BP-ANN calibration model was constructed to calculate the total iron content in iron ore.

[0117] For the PLS model, the optimal latent variables were optimized using five-fold cross-validation, and the optimal number of latent variables optimized by 5-CV was 13.

[0118] For the SVM model, its parameters were optimized by combining grid search and five-fold cross-validation. The optimal parameters of the model were a linear kernel function and a penalty parameter C = 0.01.

[0119] The optimization method is the same for the RF model and the SVM model, with the optimal parameters being ntree = 100 and mtry = M / 4;

[0120] For the VI-RF model, the optimal parameters are ntree = 100 and mtry = M / 6.

[0121] Example 1

[0122] LIBS spectral data were input into the VI-BP-ANN model and BP-ANN calibration model constructed in Example 1 and Comparative Example 1, respectively, to predict the TFe content of the test set.

[0123] Figure 4 The graph shows the prediction results of Example 1 based on the VI-BP-ANN model and the BP-ANN model in Comparative Example 1. Figure 4 The dashed line at the midpoint represents Y = X, signifying that the reference value equals the predicted value. The dashed line segments on either side of the midpoint represent Y = X ± 0.275, where 0.275 is the acceptable error according to traditional chemical analysis standard GBT 6730.65. Square points represent predictions based on the VI-BP-ANN calibration model, and circular points represent predictions based on the BP-ANN calibration model. For the 16 test samples, the square points are closer to Y = X than the circular points, meaning the predictions based on the VI-BP-ANN calibration model are closer to Y = X than those based on the BP-ANN calibration model. This indicates that the VI-BP-ANN-based predictions are closer to the reference value and more accurate. Most of the VI-BP-ANN-based predictions fall within the dashed line Y = X ± 0.275. A few samples, while not within the dashed line, are close to either side of the black dashed line, indicating that the VI-BP-ANN-based predictions are comparable to those based on traditional chemical analysis methods. In contrast, the prediction results based on the BP-ANN model differed significantly from the dashed line Y = X ± 0.275. Except for a few samples that were within the black dashed line, most were outside the dashed line, far deviating from the dashed line Y = X ± 0.275. This indicates that the prediction results of the test samples based on the BP-ANN calibration model differed significantly from the reference value and could not effectively reflect the true value of TFe content in iron ore.

[0124] Example 2

[0125] Figure 5 This is a comparison chart of the prediction performance of different models in Example 1 and Comparative Example 1; where, Figure 5 a is the 5-CV internal validation diagram; Figure 5 b is the external validation plot. For the same iron ore LIBS spectrum, the VI-BP-ANN model outperforms other methods on the calibration set, with the lowest RMSEP and RMSECV, and R... 2 The highest RMSEP was achieved by PLS, followed by BP-ANN, SVM, and VI-RF, while RF exhibited the worst performance in the calibration set, as shown in Table 5. The predicted (externally validated) RMSEP was calculated using a formula similar to RMSECV.

[0126] Table 5

[0127]

[0128] As shown in Table 5, compared to the prediction results of the other five calibration models, the VI-BP-ANN calibration model has a higher R-value on the test set. 2 Its R is 0.9450. 2 The highest RMSEP was observed, while the lowest was observed RMSEP. The BP-ANN calibration model, constructed using the full spectrum, exhibited the worst predictive performance when predicting total iron content in iron ore. Compared to the calibration set results, the BP-ANN model clearly overfitted due to its use of the full spectrum as an input variable, resulting in a significantly lower RMSEP compared to the calibration set. 2 Predicted R 2 The RMSEP was 0.7417, and the RMSEP was 0.7251 wt%. Furthermore, both PLS and SVM exhibited some degree of overfitting; the PLS calibration model showed a low RMSEP on the test set. 2 The RMSEP values ​​were 0.8727 and 0.3800 wt%, respectively; the RMSEP of the SVR calibration model was... 2 The RMSEP values ​​were 0.8812 and 0.3746 wt%, respectively; the RF calibration model achieved RMSEP values ​​of 0.8812 and 0.3746 wt% on the test set. 2 The RMSEP values ​​are 0.8586 and 0.4007 wt%, respectively. This is a slight improvement over the calibration set, likely due to the fact that RF itself is a tree model. For the VI-RF model, the RMSEP of the prediction set is... 2 The results obtained from RMSEP and RF models are very similar.

[0129] like Figure 5 As shown, for the PLS calibration model, the optimal latent variables were optimized through five-fold cross-validation, and the optimal number of latent variables was 13. When training the PLS model on the calibration set, the result was: R 2 =0.9628, RMSECV = 0.3534wt%. For the SVM model, its parameters were optimized using a combination of grid search and five-fold cross-validation. The optimal parameters for the model were a linear kernel function and a penalty parameter C = 0.01. When training the SVR model on the calibration set, the results were: R 2 =0.9329, RMSECV = 0.4682wt%. Random Forest Regression (RF) was optimized in the same way, with ntree = 100 and mtry = M / 4. The results based on RF on the calibration set are: R 2 =0.8082, RMSECV = 0.6344wt%. For the VI-RF model, the optimal parameters are ntree = 100, mtry = M / 6, and the prediction results on the calibration set are R. 2 =0.8932, RMSECV=0.5239wt%.

[0130] The results above demonstrate that, when analyzing LIBS spectral data of the same iron ore, the VI-BP-ANN model outperforms other methods on the calibration set, with R...2 The highest performance was achieved by RMSECV, followed by PLS, BP-ANN, and SVR, while RF performed the worst on the calibration set. Introducing a variable importance threshold to control the input variables of the BP-ANN model makes it possible to remove spectral noise and redundant variables, thereby reducing modeling time, improving prediction accuracy, and reducing the risk of overfitting.

[0131] Example 2

[0132] S1:

[0133] Representative iron ore powder samples from 13 brands in four countries—Australia, South Africa, Kazakhstan, and Chile—were collected. Table 2.1 shows the brands, sample numbers, and concentration ranges of major elements, totaling 244 samples. Among them, 6 brands—PB lumps, Newman lumps, Yandi powder, Newman powder, Australian concentrate, and mixed powder—were from Australia; 2 brands—South African concentrate and South African powder—were from South Africa; 4 brands—pelleted ore, 52 lumps, 71.52 lumps, and iron ore (powder)—were from Kazakhstan; and the Atacama pelletized iron concentrate brand was from Chile. X-ray fluorescence (XRF) values ​​were used as reference values, with TFe ranging from 53.26 to 65.40 wt%, SiO2 from 1.03 to 15.66 wt%, Al2O3 from 0.20 to 3.06 wt%, CaO from 0.016 to 1.768 wt%, and MgO from 0.034 to 9.900 wt%. Before LIBS measurements, to reduce the randomness and variability of the measurements, the iron ore samples need to undergo simple pretreatment. Powdered iron ore samples are gathered using polyethylene plastic rings and placed under a tablet press, pressed at 30t for 30s to form a cake shape. After pressing, the surface is cleaned with a syringe. To prevent cross-contamination, the mold is wiped with ethanol before pressing each representative sample.

[0134] The measurement results of LIBS are shown in Table 6, which shows the quantity of branded iron ore and the concentration range of major elements (wt%).

[0135] Table 6

[0136]

[0137] LIBS spectral data were acquired using a commercial LIBS system (Chemreveal 3764, TSI). The laser ablation source was a Q-switched pulsed Nd:YAG laser with a maximum emission wavelength of 1064 nm, a laser energy of 30 MJ, a delay time of 1 μs, and a frequency of 5 Hz. To minimize the influence of matrix effects caused by uneven elemental concentration distribution and differences in physical properties, LIBS spectral acquisition was performed in a 5×5 matrix. Five consecutive excitations were performed at each location, and the results were accumulated into one spectrum. The LIBS spectra collected from six different locations on the sample surface were then averaged into a single spectrum. A total of 244 analytical LIBS spectra were obtained from 244 iron ore samples. These spectra were randomly divided into a calibration set and a test set at a ratio of 80% and 20%, respectively. The calibration set was used to calibrate the model, and the test set was used to verify the model's performance.

[0138] Based on the above spectral data and total iron content data, a VI-BP-ANN model is established, which includes the following steps:

[0139] (1) Spectral preprocessing;

[0140] Five different preprocessing methods are used to preprocess the original LIBS spectra to improve the accuracy of the model: smoothing, multiplicative scattering correction (MSC), normalization, normalization + MSC, and normalization + smoothing.

[0141] The performance of different pretreatment methods was verified by five-fold cross-validation and the coefficient R was determined. 2 The model was evaluated using two metrics: root mean square error (RMSECV). Meanwhile, considering the potential impact of changes in input variables on the model, the number of neurons was optimized for each preprocessing method to obtain a good model.

[0142] Figure 6 A comparison diagram of oxides treated with different pretreatment methods; where, Figure 6 a is a comparison chart of SiO2 using different pretreatment methods; Figure 6 b is a comparison diagram of Al2O3 using different pretreatment methods; Figure 6 c is a comparison graph of CaO using different pretreatment methods; Figure 6 d is a comparison chart of MgO treated with different pretreatment methods. From Figure 6 As can be seen from a, 6b, 6c, and 6d, for SiO2, Al2O3, CaO, and MgO, normalization exhibits better performance compared to other pretreatment methods and the original spectra, with the RMSECV being the lowest at this point. 2 maximum.

[0143] Among the smoothing, MSC, and normalization preprocessing methods, normalization performed best for SiO2, Al2O3, and CaO. Smoothing preprocessing slightly improved the model's performance, while MSC preprocessing resulted in worse performance. For MgO, smoothing the spectrum led to poor model performance, while MSC preprocessing slightly improved it. Combining smoothing and MSC preprocessing with normalization for spectral preprocessing yielded better results. Figure 6 As can be seen from a, 6b, 6c, and 6d, compared to normalization, the performance of the model for quantitative analysis of SiO2, Al2O3, CaO, and MgO decreased in both preprocessing methods. 2 The decrease in normalization leads to an increase in RMSECV. Compared to smoothing or MSC preprocessing only, the model's performance is improved, indicating that normalization plays a dominant role in improving model performance. Further adding smoothing or MSC preprocessing on top of normalization distorts the original spectral information, failing to reflect the true relationships between spectral lines and thus reducing quantitative accuracy. Therefore, in the experiment, normalization was used for the analysis of SiO2, Al2O3, CaO, and MgO in iron ore, with optimal neuron counts of 40, 50, 50, and 40, respectively.

[0144] (2) RF parameter optimization: The two parameters ntree and mtry of the RF model are optimized using the OOB error rate to obtain the RF model;

[0145] For the analysis of SiO2, Al2O3, CaO, and MgO, the error rate of Out-of-Body (OOB) was studied under different values ​​of ntree and mtry. Specifically, ntree was set to 100, 200, 300, 400, 500, 600, 700, or 800, and mtry ranged from... arrive Every The value is taken (M is the LIBS spectral number of iron ore, 12814). Figure 7 The relationship between different ntree and mtry values ​​of different oxides and the OOB error rate; where, Figure 7 a represents the relationship between different ntree and mtry values ​​of SiO2 and the OOB error rate; Figure 7 b represents the relationship between different ntree and mtry values ​​of Al2O3 and the OOB error rate; Figure 7 c represents the relationship between different ntree and mtry values ​​of CaO and the OOB error rate; Figure 7 d represents the relationship between different ntree and mtry values ​​of MgO and the OOB error rate. For example... Figure 7As shown in a, 7b, 7c, and 7d, the OOB error rate exhibits the same trend under different ntree values ​​as mtry changes. For SiO2 and MgO analysis, the OOB error rate first increases and then decreases with increasing mtry. For Al2O3 analysis, the OOB error rate continuously increases with increasing mtry. However, for CaO analysis, the OOB error rate fluctuates with changes in mtry, which may be related to the characteristics of the iron ore sample itself and the range of elemental contents. For SiO2 analysis, when ntree = 200, At this time, the OOB error rate is the lowest at 0.1358. For Al2O3 analysis, when ntree = 200, At this time, the OOB error rate is the lowest, at 0.1030. For CaO analysis, when ntree = 200, At this time, the OOB error rate is the lowest at 0.0886. For MgO analysis, when ntree = 600, At this time, the minimum OOB error rate is 0.0263.

[0146] (3) Variable importance measurement: When the RF model is optimal, the RF is used to score the variable importance of the LIBS features of iron ore, and the feature variables are reordered according to the importance of the variables.

[0147] (4) Optimization of variable importance threshold and number of neurons: The number of variables input to the BP-ANN model is controlled by the variable importance threshold. Under the corresponding variable importance threshold, the number of hidden neurons, which is the most important parameter of the BP-ANN, is optimized by five-fold cross-validation, using RMSECV and R. 2 As an evaluation metric; calculate the goodness-of-fit R0. 2 The root mean square error (RMSECV) of cross-validation; R 2 The calculation formulas for RMSECV are as follows:

[0148]

[0149]

[0150] To reduce the number of variables input to the BP-ANN model, the importance of features in the LIBS spectrum of iron ore was scored using a Randomized Randomized Framework (RF) model, and the features were reordered based on their importance. A variable importance threshold was used to control the variables input to the BP-ANN model. Considering the variation in input variables, five-fold cross-validation and two evaluation metrics (R², R₀, and V²) were used in the experiment to obtain a better calibrated model. 2The impact of the number of neurons (30-50) in the BP-ANN calibrated model at each variable threshold was examined using RMSECV. The minimum RMSECV resulted in the highest R² value. 2 The variable importance threshold is optimal when it is maximized. Tables 7-10 show the BP-ANN models for analyzing oxides with different variable importance.

[0151] Table 7

[0152] Variable importance Number of variables Number of neurons RMSECV (wt%) <![CDATA[R 2 ]]> time Full spectrum 12814 40 0.8612 0.9027 37min33s 0.001 175 48 1.1137 0.8399 5s94 0.0005 344 34 1.0564 0.8554 11s16 0.0002 642 42 1.0166 0.8685 24s61 0.0001 771 50 1.003 0.8701 40s24 0.00005 970 30 0.9358 0.8859 31s20 0.00002 1345 36 0.9291 0.8881 1min09s 0.00001 1683 32 0.7639 0.9199 1 minute 21 seconds 0.000005 2174 34 0.8138 0.9103 2min25s 0.000002 2720 36 0.7675 0.9189 2min47s 0 11045 34 0.8746 0.9002 26min08s

[0153] Table 7 shows the BP-ANN models used for SiO2 analysis with different variable importance. For SiO2 analysis, when using the raw spectrum (12814) as the input variable, the optimal number of neurons is 40, the RMSECV is 0.8612wt%, and the R0 is 0.5%. 2 The value is 0.9027. Within the variable importance threshold range of 0 to 0.001, as the variable importance threshold decreases and the number of variables increases, RMSECV shows a trend of first decreasing and then increasing. 2 The modeling time initially increases and then decreases. In terms of time, the modeling time increases with the number of input variables. The shortest modeling time (5.94 seconds) is achieved when the variable importance threshold is 0.001; however, when the threshold is set to 0, the modeling time increases by 26.08 seconds. The optimal calibration model is achieved when the variable importance threshold is 0.00001 and the number of neurons is 32, resulting in the lowest RMSECV of 0.7639 and R0. 2 The maximum value is 0.9199. Compared to using the original spectrum as input, the performance of the SiO2 calibration model is significantly improved, with R... 2 The value increased from 0.9027 to 0.9199, while the RMSECV decreased from 0.8612 wt% to 0.7639 wt%. The number of variables decreased from 12814 to 2720, and the modeling time was also greatly shortened from 37 min 33 s to 1 min 21 s.

[0154] Table 8

[0155] Variable importance Number of variables Number of neurons RMSECV (wt%) <![CDATA[R 2 ]]> time Full spectrum 12814 50 0.1763 0.9149 7min42s 0.001 306 32 0.2112 0.8792 2s64 0.0005 446 36 0.2002 0.8914 5s22 0.0002 560 34 0.1984 0.8930 4s89 0.0001 792 32 0.1905 0.9012 6s89 0.00005 1248 50 0.1899 0.9017 16s96 0.00002 1877 50 0.1901 0.9013 25s54 0.00001 2486 48 0.1855 0.9061 38s87 0.000005 3384 32 0.1836 0.9082 47s22 0.000002 4673 42 0.1833 0.9082 1 minute 37 seconds 0.000001 5712 50 0.1815 0.9089 2min09s 0 10868 48 0.1804 0.9110 5min55s

[0156] Table 8 shows the BP-ANN models used for Al2O3 analysis with different variable importance. For Al2O3 analysis, the calibrated model performed best when using the raw spectrum as input to the BP-ANN model; introducing variable importance thresholds did not improve model performance. When the variable importance threshold was 0, the optimal number of neurons was 48, resulting in a relatively low RMSECV and a high Ri. 2Compared to the original spectrum, although the modeling time was reduced from 7 min 42 s to 5 min 55 s, the RMSECV increased from 0.1763 wt% to 0.1804 wt%, and R... 2 The value decreased from 0.9149 to 0.9110. This indicates that for Al2O3 analysis, introducing variable importance thresholds to control input variables during BP-ANN modeling may result in the loss of important information, suggesting that variable importance is not suitable for the quantitative analysis of all oxides in iron ore. Therefore, for SiO2, CaO, and MgO analysis, the variable importance thresholds and the number of neurons were set to 0.000002 and 36, 0.000001 and 42, and 0.000002 and 46, respectively, to construct the BP-ANN calibration model. For Al2O3 analysis, the full spectrum was used as input, and the number of neurons was set to 50 to construct the Al2O3 quantitative calibration model for subsequent analysis.

[0157] Table 9

[0158]

[0159]

[0160] Table 9 shows the BP-ANN models used for CaO analysis with different variable importance thresholds. For CaO analysis, when the variable importance threshold is set to 0.000001, the optimal number of neurons is 42, resulting in the best performance of the calibrated model, the lowest RMSECV, and the lowest R0. 2 Maximum. Compared to using the original spectrum as input to the BP-ANN model, the R-value of the five-fold cross-validation is the highest. 2 The value was increased from 0.9421 to 0.9423, RMSECV decreased from 0.1128wt% to 0.1125wt%, the number of input variables decreased from 12814 to 3762, and the modeling time was shortened from 3min37s to 37s53.

[0161] Table 10

[0162] Variable importance Number of variables Number of neurons RMSECV (wt%) <![CDATA[R 2 ]]> time Full spectrum 12814 40 0.2748 0.9841 12min44s 0.001 189 50 0.2719 0.9845 3s20 0.0005 225 44 0.2728 0.9845 2s81 0.0003 649 48 0.2727 0.9846 12s28 0.0002 719 44 0.2721 0.9847 10s61 0.0001 728 50 0.2718 0.9847 15s08 0.00005 768 34 0.2710 0.9848 10s48 0.00002 895 34 0.2712 0.9848 13s76 0.00001 1081 50 0.2700 0.9850 31s31 0.000005 1329 42 0.2698 0.9849 35s72 0.000002 1616 46 0.2679 0.9851 42s93 0.000001 2076 32 0.2690 0.9850 51s07 0 12751 40 0.2740 0.9843 10min32s

[0163] Table 10 shows the performance of BP-ANN models for MgO analysis with different variable importance thresholds. For MgO analysis, when the variable importance threshold is set to 0.000002, the optimal number of neurons is 46, at which point the calibrated model performs best, with the lowest RMSECV and R0. 2 Maximum. Compared to using the original spectrum as input, RMSECV decreased from 0.2748 wt% to 0.2679 wt%, R 2 The value was increased from 0.9841 to 0.9851, the number of input variables decreased from 12814 to 1616, and the modeling time decreased from 12 min 44 s to 42 s 93.

[0164] The above analysis shows that for the analysis of SiO2, CaO and MgO, combining variable importance with the BP-ANN model can improve the performance of the calibration model while reducing the number of input variables and shortening the modeling time. This indicates that variable importance can effectively screen out useful information that represents the original spectrum and discard information or noise that is harmful to the model.

[0165] The model takes LIBS spectral data of the test samples as input, sorts the features according to variable importance, reads the spectral features according to the optimal variable importance threshold, and returns the relevant element prediction results. Data preprocessing is performed using Pirouette (Infometrix, Inc.), and variable importance measurement and artificial neural network modeling are both completed using Python 3.8.3 (Sklearn 0.23.1).

[0166] Comparative Example 2

[0167] To verify the ability of the VI-BP-ANN model to quantitatively analyze SiO2, Al2O3, CaO, and MgO in iron ore, normalization was used to preprocess the LIBS spectral data. Using full-spectrum data as input variables, PLS, RF, and SVM models were constructed to calculate the contents of SiO2, Al2O3, CaO, and MgO in iron ore.

[0168] For the PLS calibration model, the optimal number of latent variables was optimized using 5-CV, and was 16, 17, 22 and 20 respectively;

[0169] For the SVM calibration model, the kernel type and hyperparameter C were optimized through grid search and 5-CV. The two parameters (kernel, C) were set to (linear, 0.31), (rbf, 9.91), (linear, 0.01), and (linear, 0.11), respectively.

[0170] For the RF calibration model, the parameters ntree and mtry were optimized in the same way as PLS and SVM. The optimal parameter setting for RF was: ntree = 100. Used for SiO2; ntree = 100, Used for Al2O3; ntree = 300, Used for CaO; ntree = 200. Used for MgO, where M is the characteristic number 12814 of the original LIBS spectrum of iron ore.

[0171] Example 3

[0172] To verify the predictive ability of the VI-BP-ANN model for the analysis of SiO2, Al2O3, CaO and MgO in iron ore, the contents of SiO2, Al2O3, CaO and MgO in iron ore test samples were predicted using the VI-BP-ANN calibration model.

[0173] Figure 8 The results show the predictions for oxides using the VI-BP-ANN model; where, Figure 8 a represents the prediction result of the VI-BP-ANN model for SiO2; Figure 8 b represents the prediction result of the VI-BP-ANN model for Al2O3; Figure 8 c represents the prediction result of the VI-BP-ANN model for CaO; Figure 8 d represents the prediction result of MgO by the VI-BP-ANN model.

[0174] from Figure 8 As can be seen from samples a, 8b, 8c, and 8d, the VI-BP-ANN model has good predictive ability for SiO2, Al2O3, CaO, and MgO in iron ore. The R-values ​​for the four oxide test sets are... 2 All are greater than 0.95, even for MgO, R 2 The highest value can reach 0.9977. For SiO2 prediction in the test set, the root mean square error (RMSEP) is 0.3723 wt%, R 2 The RMSEP was 0.9771. For Al2O3 prediction in the test set, the RMSEP was 0.1298 wt%, and R 2 The RMSEP for CaO prediction in the test set was 0.9504; for CaO prediction in the test set, the RMSEP was 0.0524 wt%, and the R... 2 The RMSEP for MgO prediction on the test set is 0.9878; for MgO prediction on the test set, the RMSEP is 0.1490, and the R... 2 It is 0.9971.

[0175] Example 4

[0176] In this embodiment, the LIBS spectral data that has been predicted by the model is input into the model of Example 2 and Comparative Example 2.

[0177] For the PLS calibration model, the number of variables was optimized using five-fold cross-validation for SiO2, Al2O3, CaO, and MgO analyses, resulting in 16, 17, 22, and 20 variables, respectively. For the SVM calibration model, the kernel function type and hyperparameter C were optimized using grid search and five-fold cross-validation. For SiO2, Al2O3, CaO, and MgO analyses, the two parameters (kernel, C) were set to (linear, 0.31), (rbf, 9.91), (linear, 0.01), or (linear, 0.11), respectively. For the RF calibration model, the two important parameters ntree and mtry were optimized in the same way. The optimal RF parameter was set to: ntree = 100. Used for SiO2; ntree = 100, Used for Al2O3; ntree = 300, Used for CaO; ntree = 200, For MgO, where M is the characteristic number 12814 of the original LIBS spectrum of iron ore. The coefficient of determination R was obtained using five-fold cross-validation and RMSECV. 2 The predictive capabilities of the optimized PLS, SVR, RF, and VI-BP-ANN calibration models were internally validated.

[0178] Figure 9 This is a comparison chart showing the performance of different models predicting oxides in iron ore based on five-fold cross-validation; where, Figure 9 a is a comparison chart of the performance of different models in predicting SiO2 in iron ore based on five-fold cross-validation; Figure 9 b is a comparison chart of the performance of different models in predicting Al2O3 in iron ore based on five-fold cross-validation; Figure 9 c is a comparison chart of the performance of different models in predicting CaO in iron ore based on five-fold cross-validation; Figure 9 d is a comparison of the performance of different models predicting MgO in iron ore based on five-fold cross-validation. The smaller the RMSECV, the better the R... 2 The larger the value, the better the model's predictive performance. From Figure 9 As can be seen from a, 9b, 9c, and 9d, all four models exhibited good predictive performance. Specific data are shown in Table 11.

[0179] Table 11

[0180]

[0181] Table 11 shows the performance comparison of different models during internal validation. For SiO2 analysis, the RMSECV ranged from 0.7369 wt% to 1.2142 wt%, and R0... 2 The values ​​ranged from 0.8002 to 0.9209; for Al2O3 analysis, RMSECV and R...2 The concentrations were 0.1763 wt%–0.2074 wt% and 0.8856–0.9149, respectively; for CaO analysis, RMSECV and R... 2 The concentrations were 0.1125 wt%–0.1257 wt% and 0.9245–0.9423, respectively; for MgO analysis, RMSECV and R 2 The values ​​were 0.2679 wt% to 0.4747 wt% and 0.9466 to 0.9851, respectively.

[0182] Among them, the VI-BP-ANN model has the smallest RMSECV value, R 2 The value is the largest, indicating the best predictive ability.

[0183] External validation was performed using test samples to analyze and predict the contents of SiO2, Al2O3, CaO and MgO in iron ore using three methods. The prediction results were compared with those of the VI-BP-ANN model, as shown in Table 12.

[0184] Table 12

[0185]

[0186]

[0187] Table 12 compares the performance of different models in external validation. First, for the analysis of SiO2, Al2O3, CaO, and MgO, the results of all four models on the test set are significantly better than those based on five-fold cross-validation, indicating that the constructed models have good generalization ability on the test set. In the test set, the VI-BP-ANN model shows superior predictive performance compared to the PLS, SVR, and RF calibration models. It is worth noting that among the four models, for the analysis of SiO2, Al2O3, CaO, and MgO, in terms of RMSEP, SiO2 has the highest RMSEP, followed by MgO and Al2O3, while CaO has the lowest. Considering the elemental composition concentration ranges of the 13 brands of iron ore in Table 6, this is mainly due to the differences in the elemental content ranges among the iron ore samples. Furthermore, since the samples used are real iron ore, the elemental content has a certain discontinuity, which can lead to significant differences between the predicted results and the reference values ​​for some samples during actual prediction.

Claims

1. A method for determining the content of major and minor elements in iron ore by laser-induced breakdown LIBS spectroscopy assisted by a variable importance-backpropagation artificial neural network (VI-BP-ANN), characterized in that, It includes the following steps: S1: Collect LIBS spectral data and elemental content data of iron ore from at least 13 batches in at least 4 categories, with the elemental content measured in wt%. Use Random Forest (RF) to measure the importance of LIBS spectral features. Variable importance thresholds are used to optimize the input variables of the Backpropagation-Artificial Neural Network (BP-ANN) model. The variable importance thresholds and the number of neurons in the BP-ANN model are determined using a 5-fold cross-validation (5-CV) coefficient R0. 2 To optimize the VI-BP-ANN model, the following steps are included: (The steps are described below.) 1) Preprocess the LIBS spectral data; 2) Optimize ntree and mtry based on the OOB error rate to build a RF model; 3) Use the RF model to score and rank the feature variables of the LIBS spectrum based on their importance; 4) Optimize the threshold for the importance of the variables and the number of neurons in the BP-ANN model, using the threshold to control the input variables of the BP-ANN model; Use 5-CV to examine the impact of the threshold and the number of neurons on the VI-BP-ANN model, and obtain evaluation metrics, namely RMSECV and R. 2 5) Establish a VI-BP-ANN model; 6) Select the threshold and the number of neurons to train the VI-BP-ANN model; S2: Using the VI-BP-ANN model described in S1, input LIBS spectral data. The model sorts the features according to the importance of variables and reads the spectral features according to the threshold to calculate the elemental content of iron ore.

2. The method as described in claim 1, characterized in that, The elements include a primary element and / or a secondary element; the primary element is iron; the secondary elements include one or more of calcium, magnesium, aluminum and silicon.

3. The method as described in claim 2, characterized in that, The iron content is the total iron content; And / or, the calcium content is expressed as CaO content; And / or, the magnesium content is expressed as MgO content; And / or, the aluminum content is expressed as Al2O3 content; And / or, the silicon content is expressed as SiO2 content.

4. The method as described in claim 1, characterized in that, In S1, when the element content is the total iron content, the number of iron ore batches for each category is at least 20.

5. The method as described in claim 1, characterized in that, In S1, the elemental content of the iron ore is obtained by chemical analysis. And / or, the total iron content of the iron ore ranges from 25 to 72 wt%; And / or, the silicon content of the iron ore ranges from 1.03 to 15.66 wt%, calculated as SiO2; And / or, the aluminum content of the iron ore ranges from 0.2 to 3.06 wt%, calculated as Al2O3; And / or, the calcium content of the iron ore ranges from 0.016 to 1.768 wt%, calculated as CaO; And / or, the magnesium content of the iron ore ranges from 0.034 to 9.9 wt%, calculated as MgO.

6. The method as described in claim 1, characterized in that, In S1, before measuring the LIBS spectral data, a sample pretreatment step is performed on the iron ore. The sample pretreatment step includes: sequentially aggregating and pressing the iron ore to form a cake-shaped iron ore, and purging the cake-shaped iron ore.

7. The method as described in claim 6, characterized in that, In S1, the pressure of the tablet is 30t; And / or, the tableting time is 30 seconds; And / or, the thickness of the cake-shaped iron ore is 5 mm.

8. The method as described in claim 1, characterized in that, In S1, the LIBS spectral data is obtained by the LIBS system through spectral acquisition.

9. The method as described in claim 8, characterized in that, In S1, the laser ablation source of the LIBS system is a laser; And / or, the steps of the spectral acquisition are as follows: the laser beam is vertically focused onto the surface of the iron ore through a lens, and the generated plasma radiation is coupled to the spectrometer through an optical fiber.

10. The method as described in claim 9, characterized in that, In S1, the laser is a Q-switched pulsed Nd:YAG laser; And / or, the maximum emission wavelength of the laser is 1064 nm, the repetition frequency of the laser is 5 Hz, and the laser energy is 30 MJ; And / or, the focal length of the lens is 50mm; And / or, the measurement range of the spectrometer is 190-950 nm; And / or, the delay time of the radiation is 1 μs.

11. The method as described in claim 1, characterized in that, In step 1) of S1, the preprocessing method includes one or more of smoothing, multivariate scattering correction and normalization.

12. The method as described in claim 11, characterized in that, In step 1) of S1, the preprocessing method is smoothing, multivariate scattering correction, normalization, normalization combined with multivariate scattering correction, or normalization combined with smoothing.

13. The method as described in claim 12, characterized in that, In step 1) of S1, the preprocessing method is normalization.

14. The method as described in claim 2, characterized in that, In step 2) of S1, when the element is the main element, the mtry is M / 3, M / 4, M / 5, M / 6, M / 7 or M / 8; Alternatively, when the element is the second element, the mtry is The ntree is 100 to 800.

15. The method as described in claim 14, characterized in that, In step 2) of S1, when the element is the main element, the mtry is M / 4, and the ntree is 100, the OOB error rate is 0.

089.

16. The method as described in claim 14, characterized in that, In step 2) of S1, when the element is the next element, the mtry is 1 ≤ x ≤ 16, where x is an integer; And / or, the ntree is 100·y, 1≤y≤8, where y is an integer.

17. The method as described in claim 16, characterized in that, In step 2) of S1, when the element is the second element, the mtry is When ntree is 200, the OOB error rate of the silicon element is 0.1358; And / or, when the mtry is When ntree is 200, the OOB error rate of the aluminum element is 0.1030; And / or, when the mtry is When ntree is 200, the OOB error rate of the calcium element is 0.0886; And / or, when the mtry is When ntree is 600, the OOB error rate of magnesium is 0.0263.

18. The method as described in claim 2, characterized in that, In step 4) of S1, when the element is the principal element, the threshold is 0.001, the number of neurons is 42, the number of variables is 58, and the variable importance is 0.9238; or, the threshold is 0.0005, the number of neurons is 48, the number of variables is 80, and the variable importance is 0.9382; or, the threshold is 0.0002, the number of neurons is 50, the number of variables is 170, and the variable importance is 0.9689; or, the threshold is 0.0001, the number of neurons is 48, the number of variables is 262, and the variable importance is 0.9820; or, the threshold is 0, the number of neurons is 36, the number of variables is 3517, and the variable importance is 1. And / or, in step 4) of S1, if the element is the main element, then the R 2 The calculation formulas for RMSECV are as follows:

19. The method as described in claim 2, characterized in that, When the element is the main element, for the iron content, when the threshold is 0.0002, the number of neurons is 50, and the R... 2 The value is 0.9674, and the RMSECV is 0.3089 wt%.

20. The method as described in claim 2, characterized in that, When the element is the sub-element, for the silicon element, the variable importance is 0 to 0.001 or the full spectrum; And / or, for the aluminum element, the variable importance is 0 to 0.001 or the full spectrum; And / or, for the calcium element, the variable importance is 0 to 0.001 or the full spectrum; And / or, for the magnesium element, the variable importance is 0 to 0.001 or the full spectrum.

21. The method as described in claim 20, characterized in that, When the element is the sub-element, for the silicon element, the variable importance is 0.00001, the number of variables is 1683, the number of neurons is 32, and the RMSECV is 0.7639wt%. And / or, for the aluminum element, the variable importance is full spectrum, the number of variables is 12814, the number of neurons is 50, and the RMSECV is 0.1763wt%. And / or, for the calcium element, the variable importance is 0.000001, the number of variables is 3762, the number of neurons is 42, and the RMSECV is 0.1125 wt%. And / or, for the magnesium element, the variable importance is 0.000002, the number of variables is 1616, the number of neurons is 46, and the RMSECV is 0.2679wt%.

Citation Information

Patent Citations

  • Dynamic soft measurement modeling method based on input variable selection and LSTM neural network

    CN114547974A

  • Methods of predicting and monitoring tyrosine kinase inhibitor therapy

    US20070254295A1