Titanium dioxide grade prediction method based on integrated learning SVM-RM-GBM stacking model

Through integrated learning of the SVM-RM-GBM stacking model, the high cost, low efficiency and environmental pollution problems of titanium dioxide taste prediction in high titanium slag are solved, and high-precision, low-cost and environmentally friendly titanium dioxide taste prediction is achieved, providing intelligent metallurgical process optimization.

CN120256937APending Publication Date: 2025-07-04KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371104.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems of high cost, low efficiency and environmental pollution in improving the taste of titanium dioxide in high titanium slag, and traditional detection methods have defects in time and prediction lag.

Method used

Using an integrated learning SVM-RM-GBM stacking model, an integrated learning backbone network is constructed by pre-processing, feature derivation, PCA dimensionality reduction and feature importance analysis on the raw material data set, and combined with attention mechanism and Monte Carlo simulation, the raw material proportion and process parameters are optimized.

Benefits of technology

It improves the accuracy and robustness of titanium dioxide taste prediction, reduces redundant feature interference, provides an environmentally friendly and efficient metallurgical process optimization strategy, reduces cost and time, and improves prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AY0LUBY9GESBYN7RABKPAUNP99Z507ULDBM3PDPR
    Figure AY0LUBY9GESBYN7RABKPAUNP99Z507ULDBM3PDPR
  • Figure BFNV4IGI0BFVWNHQISEQHACO1XLZPGRAF5QUVX7A
    Figure BFNV4IGI0BFVWNHQISEQHACO1XLZPGRAF5QUVX7A
  • Figure CUBH3REHCNWRD8QDBTOXRSEI9Y2OWRHN4H5GZLRL
    Figure CUBH3REHCNWRD8QDBTOXRSEI9Y2OWRHN4H5GZLRL
Patent Text Reader

Abstract

The invention relates to the field of machine learning, in particular to a titanium dioxide grade prediction method based on an integrated learning SVM-RM-GBM stacking model. The method comprises the following steps: preprocessing an original material data set to ensure data quality and meet model input requirements; based on the preprocessed data, performing feature derivation to generate new feature variables capable of reflecting the influence of chemical components and physical properties of the raw materials on the grade of titanium dioxide; a PCA dimension reduction technology and feature importance analysis are adopted to screen out a feature set which is most critical to titanium dioxide grade prediction; constructing an integrated learning backbone network, integrating a random forest model, a gradient elevator model and a support vector machine model, and performing prediction weighting by taking linear regression as a meta-model; and carrying out weight distribution on the key features by utilizing a strategy similar to an attention mechanism, and integrating prediction results of the basic model so as to evaluate and determine the grade of titanium dioxide in the high titanium slag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-energy scheduling in integrated energy systems, and more specifically, to a method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model. Background Art

[0002] With the development of industrial production, as an important industrial raw material, the quality of high-titanium slag directly affects the performance and application of downstream products. Titanium dioxide, as the main component of high-titanium slag, has a wide range of applications in multiple fields such as coatings, plastics, paper, cosmetics, solar cells, and photocatalysts. Therefore, improving the grade and extraction efficiency of titanium dioxide in high-titanium slag has become a problem to be solved in industrial production.

[0003] Traditional methods for improving the titanium dioxide grade in high-titanium slag mainly rely on empirical formulas and process improvements, which have problems such as high cost, low efficiency, and environmental pollution. In addition, existing detection technologies such as chemical analysis, X-ray diffraction, scanning electron microscopy, and energy dispersive spectroscopy, although they can provide relatively accurate measurement results, have certain defects in terms of time, cost, and prediction lag. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model to solve the problems mentioned in the above background art.

[0005] To achieve the above purpose, the present invention provides a method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model, including the following steps: S1. Preprocess the factory raw material dataset; S2. Based on the preprocessed data, perform feature derivation to generate new feature variables; S3. Use PCA dimensionality reduction technology and feature importance analysis to screen out the most critical feature set for predicting the titanium dioxide grade from the derived features; S4. Build the backbone network of the integrated learning method, integrate random forest, gradient boosting machine, and support vector machine models, and use linear regression as the meta-model to form a stacking model; S5. Input the screened key features into the backbone network, use the weight allocation strategy of the attention mechanism to weight the prediction results of the base models through the meta-model, so as to integrate the predictions of each model, and evaluate and determine the grade of the titanium dioxide; S6. Use the Monte Carlo simulation method, combine model prediction and the optimal value range of feature variables, simulate different industrial scenarios, and optimize the raw material ratio and process parameters.

[0006] As a further improvement of this technical solution, in S1, the preprocessing of the raw material dataset sequentially includes missing data processing, outlier identification and correction, data standardization, and normalization; the raw material dataset includes chemical composition data and physical property data. The chemical composition data includes titanium dioxide content, main impurity components, and other trace elements; the physical property data includes particle size distribution, density, and specific surface area. As a further improvement of this technical solution, in S2, the steps involved in generating new feature variables based on the preprocessed data are as follows: Analyze the chemical composition of the raw materials based on the molar ratio feature. The specific molar ratio feature is as follows: ; In the formula, represents the molar ratio; represents the number of moles of titanium dioxide; represents the number of moles of element α; Considering the influence of particle size, density, and specific surface area on the chemical composition of the raw materials, therefore, the molar ratio feature is optimized by introducing particle size, density, and specific surface area. The optimized molar ratio feature is specifically: ; In the formula, represents the weighted function of the physical properties of the chemical composition titanium dioxide; represents the weighted function of the physical properties of the chemical composition α; represents the molar ratio extended feature; By performing polynomial transformation on the original features, high-dimensional polynomial ratio features are generated to help capture the complex non-linear relationships between chemical components and between chemical components and physical properties. The specific polynomial ratio features are: ; In the formula, represents the polynomial ratio feature; represents the combined feature of titanium dioxide; represents the combined feature of α.

[0007] As a further improvement of this technical solution, in S3, the PCA dimensionality reduction technique is used to reduce the dimensionality of the features while retaining the most important information in the dataset, and the features that affect the prediction of titanium dioxide grade are determined through feature importance analysis. The PCA dimensionality reduction technique is based on the covariance matrix algorithm, calculates the covariance matrix between features, and describes the correlation between features; As a further improvement of this technical solution, the PCA dimensionality reduction technique based on the covariance matrix algorithm is used to screen out the feature set for predicting titanium dioxide grade from the derived features. Specifically: ; In the formula, represents the covariance matrix; represents the data matrix; represents the sample mean vector; represents the number of samples; represents the transpose of the matrix.

[0008] As a further improvement of this technical solution, the feature importance analysis is based on the Lasso objective function to select the features that have the most influence on predicting the titanium dioxide grade, specifically: ; In the formula, represents that the goal is to minimize the value of the entire expression; represents the number of samples; represents the number of features; represents the true target value of the i-th sample; represents the value of the j-th feature in the i-th sample; represents the regression coefficient of the j-th feature; represents the regularization parameter; i represents the index coefficient.

[0009] As a further improvement of this technical solution, in step S4, the steps involved in constructing the new stacked model are: First, through the multi-decision tree integration structure of the random forest, high-dimensional data and non-linear relationships are processed; the random forest model training formula involved is: ; In the formula, represents the predicted value of the random forest for the sample xi; represents the total number of decision trees in the random forest; represents the predicted value of the j-th decision tree for the sample xi; represents the index coefficient; Secondly, the gradient boosting machine starts with a weak prediction model, and the formula derivation of the gradient boosting machine is: ; In the formula, represents the predicted value of the model for the i-th sample at the k-th iteration; represents the learning rate; represents the prediction residual of the weak learner for the i-th sample in the k-th iteration; represents the predicted value of the model for the i-th sample at the k-th iteration; represents the feature vector of the i-th sample; By gradually minimizing the residuals of the loss function, the model parameters are dynamically adjusted. Through iterative optimization, the bias and variance are effectively reduced, and the prediction accuracy of the model is improved. The optimization formula for minimizing the loss function is as follows: ; In the formula, represents the true target value of the i-th sample; represents the predicted value of the i-th sample after the k-th iteration; represents the total number of samples in the dataset; represents the index coefficient; represents the feature vector of the i-th sample; i represents the index coefficient; The support vector machine model maximizes the margin between data points and the decision boundary, and searches for an optimal hyperplane in the high-dimensional space to minimize the prediction error. Specifically: ; In the formula, represents the true target value of the i-th sample; represents the predicted value of the i-th sample; represents a predefined tolerance; represents the index coefficient; represents the loss function value of the support vector machine model; represents the total number of samples in the dataset; Finally, using linear regression as the meta-model, by weighting the prediction outputs of each base model and optimizing the weighting coefficients through the least squares method, the accurate prediction of the titanium dioxide content in high-titanium slag is achieved. Specifically: ; In the formula, represents the predicted value of the final stacked model; represents the predicted value of the i-th sample of the random forest model; represents the predicted value of the i-th sample of the gradient boosting machine model; The i-th sample represents the predicted value of the support vector machine model; is the weighting coefficient optimized through the least squares method and cross-validation.

[0010] As a further improvement of this technical solution, in S5, by calculating the gradient of each feature on the model output, using the backpropagation algorithm to identify the feature that has the greatest impact on the prediction result, and accordingly adjusting its weight; subsequently, using the meta-model to comprehensively analyze the weighted features and integrating the prediction results of each base model.

[0011] As a further improvement of this technical solution, the steps involved in integrating the prediction results of each base model are as follows: Calculate the gradient of each feature on the model output, specifically as follows: ; In the formula, represents the loss function; represents the partial derivative; represents the predicted output of the model for the input x at the k-th iteration; represents the partial derivative of the loss function L with respect to the feature xi; represents the i-th feature; represents the partial derivative of the loss function L with respect to the output Fk(x) of the current layer; represents the partial derivative of the output Fk(x) of the current layer with respect to the feature xi; Use the backpropagation algorithm to identify the feature that has the greatest impact on the prediction result and adjust its weight accordingly. The specific backpropagation algorithm is as follows: ; In the formula, represents the weight of the feature xi; represents the contribution degree of the i-th feature to the prediction result; represents the sum of the gradients of all features; Use the meta-model to comprehensively analyze the weighted features and integrate the prediction results of each basic model into the weighted prediction output of the meta-model, specifically as follows: ; In the formula, represents the optimal weight of each basic model in the meta-model; represents the prediction output of each basic model for the i-th sample.

[0012] As a further improvement of this technical solution, in S6, based on the model prediction result and the optimal value range of the feature variables, use the Monte Carlo simulation method to simulate different industrial production scenarios to help determine the optimal raw material ratio and process parameters.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. In this method for predicting the titanium dioxide grade based on the stacked model of integrated learning SVM-RM-GBM, through the stacked structure of the integrated learning model, the random forest, gradient boosting machine, and support vector machine are organically combined, enabling the model to handle high-dimensional and complex nonlinear relationships, effectively improving the accuracy of predicting the titanium dioxide grade. Through PCA dimensionality reduction and feature importance analysis, the model can screen out key feature variables, further improving the prediction accuracy and robustness, and reducing the interference of redundant features. Compared with traditional experimental methods, the present invention has significant improvements in prediction accuracy, cost savings, and time reduction.

[0014] 2. In the titanium dioxide grade prediction method based on the integrated learning SVM-RM-GBM stacking model, through the weight allocation strategy, adaptive learning and backpropagation algorithm to dynamically adjust the feature weights, and combined with the Monte Carlo simulation method, it can flexibly optimize the raw material ratio and process parameters according to different industrial scenarios, providing an environmentally friendly and efficient metallurgical process optimization strategy. This method shows obvious advantages in reducing resource waste and environmental impact during the experiment process, providing a more intelligent and sustainable development path for the actual industrial production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is the overall method flowchart of the present invention; Figure 2 It is the backbone network architecture diagram of the integrated learning model used in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment: Please refer to Figure 1 As shown, this embodiment provides a titanium dioxide grade prediction method based on the integrated learning SVM-RM-GBM stacking model, including the following steps: S1. Preprocess the factory raw material data set; S2. Based on the preprocessed data, perform feature derivation to generate new feature variables; S3. Use PCA dimensionality reduction technology and feature importance analysis to screen out the most critical feature set for titanium dioxide grade prediction from the derived features; S4. Build the backbone network of the integrated learning method, integrate the random forest, gradient boosting machine and support vector machine models, and use linear regression as the meta-model to form a stacking model; S5. Input the screened key features into the backbone network, use the weight allocation strategy of the attention mechanism, and weight the prediction results of the basic models through the meta-model to integrate the predictions of each model, and evaluate and determine the grade of titanium dioxide grade; S6. Use the Monte Carlo simulation method, combined with the best value range of the model prediction and feature variables, to simulate different industrial scenarios and optimize the raw material ratio and process parameters.

[0018] In S1, the preprocessing of the raw material dataset includes but is not limited to missing data processing, outlier identification and correction, data standardization and normalization to ensure the quality of the dataset and the accuracy of model training. The steps involved in the preprocessing are as follows: For missing data processing, the mean filling algorithm is used to fill in the missing data points. The specific mean filling algorithm is as follows: ; In the formula, represents the mean of the dataset; represents the i-th data point in the dataset; represents the number of non-missing values in the dataset; Based on the data filled by the mean filling algorithm, for outlier identification and correction, the standard deviation algorithm is used for data identification and correction. The specific standard deviation algorithm is as follows: ; In the formula, represents the population standard deviation; represents the mean of the dataset; represents the i-th data point in the dataset; represents the number of non-missing values in the dataset; Based on the mean filling algorithm and the standard deviation algorithm, data standardization and normalization are performed. The data standardization and normalization scale the data to a fixed range based on the Z-Score algorithm to make the data conform to the standard normal distribution. The specific Z-Score algorithm is as follows: ; In the formula, represents the standardized value; represents the original data point; represents the population standard deviation; represents the mean of the dataset; The raw material dataset includes chemical composition data and physical property data. The chemical composition data includes titanium dioxide content, main impurity components, and other trace elements; the physical property data includes particle size distribution, density, and specific surface area; Titanium dioxide content: The concentration or purity of titanium dioxide, which is the main target variable for prediction.

[0019] Main impurity components: The content of chemical elements such as FeO, Fe2O3, SiO2, MgO, CaO, etc. The proportion and combination of these components will affect the purification effect and grade of titanium dioxide (TiO2).

[0020] Other trace elements: May include minor impurity components such as Al2O3, Na2O, K2O, etc. Although the content is small, it may also have an indirect impact on the grade of titanium dioxide.

[0021] Particle size distribution: The size range of raw material particles, which affects the reaction surface area and reaction rate of the material.

[0022] Density: The density data of the raw material, which is used to judge the compactness and sedimentation characteristics of the material particles.

[0023] Specific surface area: The total surface area per unit mass, which is related to the particle size and affects the chemical reaction activity of the material and the separation effect of titanium dioxide.

[0024] In S2, based on the preprocessed data, the steps involved in generating new feature variables are as follows: Analyze the chemical composition of the raw material based on the molar ratio feature, and the specific molar ratio feature is as follows: ; In the formula, represents the molar ratio; represents the number of moles of titanium dioxide; represents the number of moles of element α; Considering the influence of particle size, density and specific surface area on the chemical composition of the raw material, therefore, by introducing particle size, density and specific surface area to optimize the molar ratio feature, the optimized molar ratio feature is specifically: ; In the formula, represents the weighted function of the physical properties of the chemical composition TiO2; represents the weighted function of the physical properties of the chemical composition α; represents the molar ratio extended feature; In the prediction of titanium dioxide grade, the correlation or influence between physical properties and the chemical composition of the raw material is reflected in the following aspects: Correlation and influence between particle size and chemical composition: Particle size determines the size of raw material particles, directly affecting the surface contact area of the raw material during smelting or reaction; the smaller the particle size, the larger the total surface area, making the reaction rate possibly higher and the chemical composition easier to release or react; during the purification of high-titanium slag, smaller particle size helps to separate titanium dioxide and other impurities more thoroughly, improving the titanium dioxide grade; the diffusion rate of smaller-sized particles in the reaction is higher, possibly improving the conversion efficiency of titanium dioxide; Correlation and Influence between Density and Chemical Composition: Density reflects the tightness of particles within a substance. Substances with higher density often contain more heavy metal elements or components. During the metallurgical process, raw materials with different densities often exhibit different physical sedimentation characteristics, which can affect the separation effect of titanium dioxide. Particles with high density settle faster, and density differences can be utilized during the smelting process to separate impurities and titanium dioxide components. This density difference helps to more effectively remove non-titanium dioxide components during the purification process, improving the quality of titanium dioxide. Correlation and Influence between Specific Surface Area and Chemical Composition: Specific surface area is closely related to particle size and represents the surface area per unit mass of particles. Materials with a large specific surface area are more likely to come into contact with other substances during chemical reactions, resulting in a faster reaction rate and higher conversion efficiency. During the metallurgical process of high-titanium slag, particles with a large specific surface area can react with reactants more quickly, thus affecting the conversion efficiency and purity of titanium dioxide. This plays an important role in optimizing reaction conditions and improving the extraction efficiency of titanium dioxide. By performing polynomial transformation on the original features, high-dimensional polynomial ratio features are generated to help capture the complex non-linear relationships between chemical components and between chemical components and physical properties. The specific polynomial ratio features are as follows: ; In the formula, represents the polynomial ratio feature; represents the binding feature of titanium dioxide; represents the binding feature of α.

[0025] In S3, PCA dimensionality reduction technology is used to reduce the dimensionality of features while retaining the most important information in the dataset, and feature importance analysis is used to determine the features that have a significant impact on predicting the quality of titanium dioxide. The PCA dimensionality reduction technology is based on the covariance matrix algorithm and is used to screen out the feature set for predicting the quality of titanium dioxide from the derived features, calculate the covariance matrix between features, and describe the correlation between features. The specific covariance matrix algorithm is as follows: ; In the formula, represents the covariance matrix; represents the data matrix; represents the sample mean vector; represents the number of samples; represents the transpose of the matrix.

[0026] Feature importance analysis is based on the Lasso objective function to select the features that have the most impact on predicting the quality of titanium dioxide. Specifically: ; In the formula, Indicates that the goal is to minimize the value of the entire expression; Indicates the number of samples; Indicates the number of features; Indicates the true target value of the i-th sample; Indicates the value of the j-th feature in the i-th sample; Indicates the regression coefficient of the j-th feature; Indicates the regularization parameter; i represents the index coefficient.

[0027] Such as Figure 2 As shown, in S4, the steps involved in constructing the new stacked model are: First, through the multi-decision tree integration structure of the random forest, high-dimensional data and non-linear relationships are processed; the random forest model training formula involved is: ; In the formula, Indicates the predicted value of the random forest for the sample xi; Indicates the total number of decision trees in the random forest; Indicates the predicted value of the j-th decision tree for the sample xi; Indicates the index coefficient; Secondly, the gradient boosting machine starts with a weak prediction model, and the formula derivation of the gradient boosting machine is: ; In the formula, Indicates the predicted value of the model for the i-th sample at the k-th iteration; Indicates the learning rate; Indicates the prediction residual of the weak learner for the i-th sample in the k-th iteration; Indicates the predicted value of the model for the i-th sample at the k-th iteration; Indicates the feature vector of the i-th sample; By gradually minimizing the residual of the loss function, the model parameters are dynamically adjusted. Through iterative optimization, the bias and variance are effectively reduced, and the prediction accuracy of the model is improved. The optimization formula for minimizing the loss function is: ; In the formula, Indicates the true target value of the i-th sample; Indicates the predicted value of the i-th sample after the k-th round of iteration; Indicates the total number of samples in the dataset; Indicates the index coefficient; Indicates the feature vector of the i-th sample; i represents the index coefficient; The support vector machine model then maximizes the margin of the data points to the decision boundary and finds an optimal hyperplane in the high-dimensional space to minimize the prediction error. Specifically: ; In the formula, represents the true target value of the i-th sample; represents the predicted value of the i-th sample; represents a predefined tolerance; represents the index coefficient; represents the loss function value of the support vector machine model; represents the total number of samples in the dataset; Finally, using linear regression as the meta-model, by weighting the predicted outputs of each base model and optimizing the weighting coefficients through the least squares method, an accurate prediction of the titanium dioxide content in high-titanium slag was achieved, specifically: ; In the formula, represents the predicted value of the final stacked model; represents the predicted value of the i-th sample of the random forest model; represents the predicted value of the i-th sample of the gradient boosting machine model; The i-th sample represents the predicted value of the support vector machine model; are the weight coefficients optimized through the least squares method and cross-validation.

[0028] In S5, by calculating the gradient of each feature with respect to the model output, the features that have the greatest impact on the prediction result are identified using the backpropagation algorithm, and their weights are adjusted accordingly; subsequently, the meta-model is used to comprehensively analyze the weighted features and integrate the prediction results of each base model.

[0029] The steps involved in integrating the prediction results of each base model are: Calculate the gradient of each feature with respect to the model output, specifically: ; In the formula, represents the loss function; represents the partial derivative; represents the predicted output of the model for the input x at the k-th iteration; represents the partial derivative of the loss function L with respect to the feature xi; represents the i-th feature; represents the partial derivative of the loss function L with respect to the current layer output Fk(x); represents the partial derivative of the current layer output Fk(x) with respect to the feature xi; Use the backpropagation algorithm to identify the features that have the greatest impact on the prediction result and adjust their weights accordingly. The specific backpropagation algorithm is: ; In the formula, represents the weight of the feature xi; represents the contribution degree of the i-th feature to the prediction result; represents the sum of all feature gradients; The meta-model is used to comprehensively analyze the weighted features and integrate the prediction results of each basic model into the weighted prediction output of the meta-model. Specifically: ; In the formula, represents the optimal weight of each basic model in the meta-model; represents the prediction output of each basic model for the i-th sample.

[0030] In S6, based on the model prediction results and the optimal value range of the feature variables, the Monte Carlo simulation method is used to simulate different industrial production scenarios to help determine the optimal raw material ratio and process parameters.

[0031] The role of the Monte Carlo simulation method is to simulate different production scenarios by randomly combining a large number of process parameters and raw material ratios, so as to identify the optimal parameter combination that can maximize the titanium dioxide grade. This method not only greatly improves the accuracy and stability of the prediction, but also effectively reduces the number of tests and costs in actual production, providing an efficient and feasible optimization strategy for the titanium dioxide purification process.

[0032] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model, characterized in that: It includes the following steps: S1. Preprocess the factory raw material dataset; S2. Based on the preprocessed data, perform feature derivation to generate new feature variables; S3. Use PCA dimensionality reduction technology and feature importance analysis to screen out the most critical feature set for titanium dioxide grade prediction from the derived features; S4. Build the backbone network of the ensemble learning method, integrate the random forest, gradient boosting machine, and support vector machine models, and use linear regression as the meta-model to form a stacked model; S5. Input the screened key features into the backbone network, use the weight assignment strategy of the attention mechanism to weight the prediction results of the base models through the meta-model, so as to integrate the predictions of each model, evaluate and determine the grade of titanium dioxide; S6. Use the Monte Carlo simulation method, combine model prediction and the optimal value range of feature variables, simulate different industrial scenarios, and optimize the raw material ratio and process parameters.

2. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In the above S1, the preprocessing of the raw material dataset sequentially includes missing data processing, outlier identification and correction, data standardization, and normalization; The raw material dataset includes chemical composition data and physical property data. The chemical composition data includes titanium dioxide content, main impurity components, and other trace elements; The physical property data includes particle size distribution, density, and specific surface area.

3. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In the above S2, the steps involved in generating new feature variables based on the preprocessed data are: Analyze the chemical composition of the raw materials based on the molar ratio feature. The specific molar ratio feature is as follows: ; In the formula, represents the molar ratio; represents the number of moles of titanium dioxide; represents the number of moles of element α; Optimize the molar ratio feature by introducing particle size, density, and specific surface area. The optimized molar ratio feature is specifically: ; In the formula, represents a weighted function of the physical properties of the chemical component titanium dioxide; represents a weighted function of the physical properties of the chemical component α; represents a molar ratio expansion feature; Generate high-dimensional polynomial ratio features by performing polynomial transformation on the original features. The specific polynomial ratio feature is: ; In the formula, represents the polynomial proportional feature; represents the binding feature of titanium dioxide; represents the binding feature of α.

4. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In the above S3, the PCA dimensionality reduction technology is used to reduce the dimension of the features while retaining the most important information in the dataset, and the features that affect the prediction of titanium dioxide grade are determined through feature importance analysis. The PCA dimensionality reduction technology is based on the covariance matrix algorithm to calculate the covariance matrix between features.

5. A prediction method for titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 4, characterized in that: The PCA dimensionality reduction technology based on the covariance matrix algorithm is used to screen out the feature set for titanium dioxide grade prediction from the derived features, specifically: ; In the formula, represents the covariance matrix; represents the data matrix; represents the sample mean vector; represents the number of samples; represents the transpose of the matrix.

6. A titanium dioxide grade prediction method based on an integrated learning SVM-RM-GBM stacking model according to claim 4, characterized in that: The feature importance analysis based on the Lasso objective function is used to select the features that have the most influence on the prediction of titanium dioxide grade, specifically: ; In the formula, indicates that the goal is to minimize the value of the entire expression; represents the number of samples; represents the number of features; represents the true target value of the i-th sample; represents the value of the j-th feature in the i-th sample; represents the regression coefficient of the j-th feature; represents the regularization parameter; i represents the index coefficient.

7. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In the above S4, the steps involved in building a new stacked model are: First, through the multi-decision tree integration structure of the random forest, process high-dimensional data and non-linear relationships. The random forest model training formula involved is: ; In the formula, represents the predicted value of the random forest for the sample xi; represents the total number of decision trees in the random forest; represents the predicted value of the j-th decision tree for the sample xi; represents the index coefficient; Secondly, the gradient boosting machine starts with a weak prediction model. The formula derivation of the gradient boosting machine is: ; Wherein, represents the predicted value of the model for the i-th sample at the k-th iteration; represents the learning rate; represents the prediction residual of the weak learner for the i-th sample in the k-th iteration; represents the predicted value of the model for the i-th sample at the k-th iteration; represents the feature vector of the i-th sample; By gradually minimizing the residuals of the loss function, dynamically adjust the model parameters, and through iterative optimization, effectively reduce the bias and variance, and improve the prediction accuracy of the model. The optimization formula for minimizing the loss function is: ; Wherein, represents the true target value of the i-th sample; represents the predicted value of the i-th sample after the k-th iteration; represents the total number of samples in the dataset; represents the index coefficient; represents the feature vector of the i-th sample; i represents the index coefficient; The support vector machine model maximizes the interval between data points and the decision boundary, and finds an optimal hyperplane in the high-dimensional space to minimize the prediction error, specifically: ; Wherein, represents the true target value of the i-th sample; represents the predicted value of the i-th sample; represents a predefined tolerance; represents an index coefficient; represents the loss function value of the support vector machine model; represents the total number of samples in the dataset; Finally, using linear regression as the meta-model, by weighting the prediction outputs of each basic model and optimizing the weighting coefficients by the least squares method, the accurate prediction of the titanium dioxide grade in high-titanium slag is achieved, specifically as follows: ; In the formula, it represents The predicted value of the final stacked model; It represents the predicted value of the random forest model for the i-th sample; It represents the predicted value of the gradient boosting machine model for the i-th sample; The predicted value of the support vector machine model for the i-th sample; Is the weight coefficient optimized by the least squares method and cross-validation.

8. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In S5, by calculating the gradient of each feature on the model output, the feature with the greatest impact on the prediction result is identified using the backpropagation algorithm, and its weight is adjusted accordingly; Subsequently, the meta-model is used to comprehensively analyze the weighted features and integrate the prediction results of each basic model.

9. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 8, characterized in that: The steps involved in integrating the prediction results of each basic model are as follows: Calculate the gradient of each feature on the model output, specifically as follows: ; In the formula, represents the loss function; represents the partial derivative; represents the predicted output of the model for the input x at the k-th iteration; represents the partial derivative of the loss function L with respect to the feature xi; represents the i-th feature; represents the partial derivative of the loss function L with respect to the output Fk(x) of the current layer; represents the partial derivative of the output Fk(x) of the current layer with respect to the feature xi; The feature with the greatest impact on the prediction result is identified using the backpropagation algorithm, and its weight is adjusted accordingly. The backpropagation algorithm is specifically: ; In the formula, represents the weight of the feature xi; represents the contribution degree of the i-th feature to the prediction result; represents the sum of all feature gradients; The meta-model is used to comprehensively analyze the weighted features and integrate the weighted prediction output of the prediction results of each basic model, specifically as follows: ; In the formula, represents the optimal weight of each basic model in the meta-model; represents the predicted output of each basic model for the i-th sample.

10. A method for predicting the titanium dioxide grade based on an integrated learning SVM-RM-GBM stacking model according to claim 1, characterized in that: In S6, based on the model prediction results and the optimal value range of the feature variables, the Monte Carlo simulation method is used to simulate different industrial production scenarios to help determine the optimal raw material ratio and process parameters.