Tire blank quality parameter prediction method based on dynamic feature selection and model stacking
Through the methods of dynamic feature selection and model stacking, the problems of low prediction accuracy of tire blank quality parameters and unstable performance in tire blank production are solved, and more accurate and stable prediction of tire blank quality parameters are achieved, which improves the quality control ability and market competitiveness of tire blanks.
Patent Information
- Application Number
- CN202411251498.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-07
- Publication Date
- 2025-06-06
AI Technical Summary
In the production process of tire blanks, it is difficult for the prior art to effectively predict the quality parameters of the tire blanks, mainly because the number of data samples is limited, the high characteristic dimensions are high, and the traditional feature selection methods are prone to large errors. In the face of large industrial fluctuations, a single model has low prediction accuracy and unstable performance when facing large industrial fluctuations.
A method of predicting the quality parameter of the tire blank based on dynamic feature selection and model stacking is adopted. Features are scored through random forests, feature combinations are selected dynamically, and multi-level prediction and data relationship mining are carried out through model stacking, combining base and meta-models.
It improves the accuracy and stability of the prediction of the quality parameters of the tire blank, can better process complex industrial data, enhances the control ability of the tire blank quality, and improves market competitiveness.
Smart Images

Figure CN120105046A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of green tire quality parameter prediction, and in particular relates to a green tire quality parameter prediction method based on dynamic feature selection and model stacking. Background Art
[0002] Despite the huge production capacity, there are still many challenges in the tire manufacturing process, especially in the quality control of green tires. As a key primary form of tire manufacturing, the quality of green tires is directly related to the performance and safety of finished tires. Therefore, improving the ability to predict green tire quality parameters is crucial to enhancing the competitiveness of China's tire industry.
[0003] In the field of tire manufacturing, machine learning is often used to predict green tire quality parameters. Machine learning predicts by establishing a mapping relationship between data collected during the production process and green tire quality. It can effectively handle nonlinear problems in data and has the advantages of simple implementation and low cost. Traditional machine learning includes decision trees, support vector machines, and neural networks. However, in actual production, it is difficult for prediction models to effectively learn prediction information from them. Researchers have found that this is mainly because the number of problem samples obtained during the green tire production process is limited and the feature dimension is high.
[0004] In terms of features, traditional feature selection methods usually rely on manually set thresholds to decide which features to select. However, in complex industrial production environments, due to the variable and uncertain relationships between features, this method is prone to large errors, affecting the accuracy of the model.
[0005] In terms of prediction models, many decision tree-based algorithms show good performance when dealing with complex industrial data problems, but different tree-based models have their own advantages and disadvantages when facing the same problem. Therefore, in order to obtain better prediction results, model stacking is used to combine multiple tree-based models to give full play to their respective advantages. Studies have shown that model stacking has higher prediction accuracy when dealing with complex data.
[0006] In summary, this paper proposes a green tire quality parameter prediction method based on dynamic feature selection and model stacking, aiming to accurately predict green tire quality parameters. First, all features are scored for importance using random forests, and various combinations from single features to all features are gradually explored through cross-validation, and the feature combination with the highest score is selected; next, the base model is used to make a preliminary prediction of the reduced-dimensional data, and then the meta-model is used to integrate the preliminary prediction results of the base model and further deeply mine the data relationship to obtain the final prediction result. Summary of the invention
[0007] The present invention aims to solve the defects of the prior art and provide a method for predicting green tire quality parameters based on dynamic feature selection and model stacking. It uses dynamic feature selection to pre-process data, and then predicts the processed data through model stacking to obtain the predicted value of green tire quality parameters. The prediction result is more accurate and has practical significance.
[0008] To achieve the above object, the present invention adopts the following technical solution, including the following steps:
[0009] Step 1: Obtain tire green quality parameter data, check whether there are extreme missing values or abnormal values in the data, and fill them with the average value of the previous and next two data to obtain the initial data sequence.
[0010] Step 2: Use the random forest algorithm to score the importance of all features.
[0011] Step 3: Use dynamic feature selection to sort the importance of high-dimensional data and establish different feature combinations with the number of features gradually increasing from 1 to all features. Use cross-validation to select the feature combination that has the most significant impact on the target value. Divide the selected data into a training set and a test set.
[0012] Step 4: Send the training set data to the base model layer in the model stack to obtain the meta-features of the preliminary prediction results of the base model.
[0013] Step 5: Use the meta-model to further train the meta-features to obtain the prediction model.
[0014] Furthermore, in step 2, using random forest to score feature importance includes the following steps:
[0015] Step 2-1: Use the average impurity reduction method through the random forest model to evaluate the importance of the feature and generate a feature importance score. The mathematical expression is as follows:
[0016]
[0017] Where parent represents the impurity of the parent node, left and right represent the impurity of the left and right child nodes after segmentation, respectively. left and N right is the number of samples of the left and right child nodes, N is the total number of samples of the parent node, and I represents the impurity function. (2) is the impurity function, where P i represents the probability of the i-th event in the system.
[0018] Furthermore, in step 3, selecting a feature combination includes the following steps:
[0019] Step 3-1: Based on the feature importance scores obtained in step 2-1, through dynamic feature selection, establish different feature combinations with the number of features gradually increasing from 1 to including all features.
[0020] Step 3-2: Use 5 rounds of cross-validation to validate feature subsets with different numbers of features and select the feature combination with the greatest impact.
[0021] Step 3-3: Divide 20% of the selected feature combinations into a test set and 80% into a training set.
[0022] Step 3-4: The training set is sent to the model stacking algorithm for training.
[0023] Furthermore, in step 4, obtaining meta-features includes the following steps:
[0024] Step 4-1: Use the base model layer to make a preliminary prediction on the data set after feature selection; in this layer, random forest, GBDT, and AdaBoost are three different base models that make independent predictions on the data and generate corresponding meta-features. The prediction mathematical expression is as follows:
[0025]
[0026] Among them, x is the tire quality parameter data, h i (x) represents the prediction result of the i-th decision tree for x, v is the step coefficient (learning rate), M is the total number of iterations, α m is the weight of the mth weak learner, h m (x) is the prediction result of the mth weak learner.
[0027] Furthermore, in step 5, obtaining the final prediction model includes the following steps:
[0028] Step 5-1: Integrate the meta-features obtained in step 4-1 through the meta-model and make further predictions to obtain a prediction model.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The method proposed in the present invention improves the problems of low prediction accuracy and poor anti-interference ability of traditional prediction models when faced with problems such as large data fluctuations, complex data relationships and high dimensions, and few problem samples in actual industrial production environments. Compared with traditional feature selection methods, dynamic feature selection selects more in-depth useful features through feature importance scoring and cross-validation of the performance of different feature combinations. Model stacking combines multiple prediction models to solve the problems of low prediction accuracy and unstable performance of a single model when faced with large industrial fluctuations. The tire quality parameter prediction method based on dynamic feature selection and model stacking has more accurate prediction capabilities, can better realize the early prediction of tire quality, further deepen the company's control over tire quality, and improve market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. The protection scope of the present invention is not limited to the following descriptions.
[0032] Figure 1 It is a work flow chart of the present invention.
[0033] Figure 2 It is a workflow diagram of dynamic feature selection of the present invention.
[0034] Figure 3 It is a schematic diagram of the stacking structure of the model of the present invention.
[0035] Figure 4-7 It is a prediction result diagram of the present invention. DETAILED DESCRIPTION
[0036] like Figure 1 As shown, the present invention combines the two methods of feature dimensionality reduction and prediction model to propose a design method:
[0037] First, the tire quality parameters are obtained, and the key data in the tire production process are screened through dynamic feature selection. These feature data are input into the base model layer in the model stack. The base model generates meta-features by deeply mining the complex relationship between features. The multi-layer perceptron (MLP) neural network is used as a meta-model to further train these meta-features to obtain the final prediction results.
[0038] Specifically, the prediction workflow is as follows Figure 1 As shown in Figure 2, the dynamic feature selection workflow is as follows: Figure 2 As shown, the model stacking structure is as follows Figure 3 As shown, the prediction results are Figure 4-7 shown.
[0039] The prediction steps are as follows:
[0040] The first step is to obtain tire quality parameter data and check whether there are extreme missing values or abnormal values in the data. If so, fill them with the average value of the previous and next two data to obtain the initial data sequence.
[0041] Take the tire quality data of a domestic company in 2023 as an example, which contains information such as tire production parameters and the working performance indicators of the forming machine. The data is a multi-objective regression problem, including 4 target values and 200 features. The collected data is screened for abnormal data and extreme outliers are eliminated.
[0042] In the second step, the preliminarily processed data is further processed through dynamic feature selection. The reduced-dimensionality dataset is obtained through feature importance sorting and cross-validation. The reduced-dimensionality dataset is divided into a test set and a training set.
[0043] Dynamic feature selection Figure 2 As shown, it includes the following three main steps:
[0044] (1) Feature importance ranking: Random forest is used to train the data after preliminary processing. In the random forest model, the feature importance score is calculated by evaluating the contribution of each feature to the reduction of impurity when splitting a node in the decision tree. During the training process, each tree will reduce the impurity of the node when splitting the feature. The total importance of the feature is the average of the reduction in impurity of the feature in all trees. The final score reflects the impact of each feature on the model prediction results. The scored features are ranked, and through dynamic feature selection, different feature combinations are established with the number of features gradually increasing from 1 to all features.
[0045] (2) Cross-validation: First, sort the features by importance and select the feature combination with the highest score for preliminary validation. Then, verify each feature combination step by step in descending order of score. During the validation process, continuously update and record the score of each feature combination until all combinations have completed validation.
[0046] (3) Feature selection: Based on the scores of each feature combination recorded, select the feature combination with the highest score.
[0047] Among these three steps, the feature importance score plays a major role. The expression is as follows:
[0048]
[0049] Where parent represents the impurity of the parent node, left and right represent the impurity of the left and right child nodes after segmentation, respectively. left and N rightis the number of samples of the left and right child nodes, N is the total number of samples of the parent node, and I represents the impurity function. (2) is the impurity function, where P i represents the probability of the i-th event in the system.
[0050] The third step is to send the training set into the model stacking algorithm for training to obtain the prediction model.
[0051] Model stacking is an ensemble learning method that improves overall prediction performance by combining the prediction results of multiple base learners. It uses different models to learn data in different ways, integrates the advantages of each model, and thus builds a stronger prediction model. The core principle of model stacking is to combine the meta-features generated by the base learners and feed them into the meta-model as input. The meta-model is trained with these meta-features to generate the final prediction model. The working principle of model stacking is as follows Figure 3 shown.
[0052] The model stacking algorithm is executed in the following steps:
[0053] (1) Read the input sample and send it to the base learner layer.
[0054] (2) The base learner generates meta-features. Different base learners generate different meta-features for the same sample.
[0055] (3) Combine meta-features, summarize all meta-features of the same sample, and input them into the meta-model layer.
[0056] (4) Train the meta-model. The meta-model uses these meta-features to train and construct the final prediction model.
[0057] The task of the base model is to make a preliminary prediction of the data after dimensionality reduction, so that the meta-model can make more accurate predictions on this basis. The base model mainly consists of the following parts:
[0058] (1) Random forest: A decision tree-based ensemble learning method that improves the accuracy and robustness of the model by building multiple decision trees and averaging their results. The expression of the random forest prediction result is as follows:
[0059]
[0060] where h i (x) represents the prediction result of the i-th decision tree for the input tire quality parameter x.
[0061] (2) GBDT: An ensemble learning method that gradually builds multiple weak learners (usually decision trees), with each new learner attempting to correct the errors of the previous model. The core idea is to connect decision trees in series, with each tree improving on the previous tree to improve the prediction accuracy of the overall model. The prediction expression is as follows:
[0062]
[0063]
[0064] Among them, .L is the loss function, using the residual ri,m as the target variable, and v is the step size coefficient (learning rate).
[0065] (3) AdaBoost: A boosting method that adjusts sample weights to make the model pay more attention to samples with larger prediction errors, thereby gradually improving the performance of the model. The training of each weak learner (decision tree) is based on an updated sample weight distribution, and the final model is obtained through weighted majority voting. In this study, the regression tree algorithm is used as the basic learner. The expression is as follows:
[0066]
[0067] w i,m+1 =w i,m exp(α m I(sh(x i )≠y i )) (11)
[0068]
[0069] in, is an indicator function, which takes the value 1 when the prediction is wrong and 0 otherwise. m Represents the weighted error rate of the mth weak learner. m represents the weight of the mth weak learner, H(x) is the prediction result of the final model, M is the total number of iterations, and α m is the weight of the mth weak learner, h m (x) is the prediction result of the mth weak learner.
[0070] The task of the meta-model is to synthesize the prediction results of the base model and further generate the final prediction results on this basis. Its main components are as follows:
[0071] Multilayer Perceptron (MLP): A basic feedforward neural network consisting of one or more intermediate layers (hidden layers) and an output layer. Each layer consists of several neurons, and information is transmitted between neurons through weighted connections. It can effectively learn and capture complex data patterns through its deep network structure and nonlinear activation function, thereby enhancing the ability to model nonlinear relationships.
[0072] The fifth step is to send the test set to the corresponding model for prediction to obtain the prediction results.
[0073] In order to verify the prediction performance of the method proposed in this paper, the tire quality data of a domestic company in 2023 will be analyzed. The data is a multi-objective regression task with 4 target values: DB1, DB2, DB4, and DB5. It is compared with RF, MLP, and the model stacking method without dynamic feature selection.
[0074] RF is a powerful machine learning model that improves the accuracy and stability of predictions by building and integrating multiple decision trees. When training each tree, it randomly selects data subsets and features. This randomness helps reduce overfitting of the model and enhances the generalization ability of new data. Random forests are applicable to various data types, can automatically handle feature selection and missing data problems, are often used for classification and regression tasks, and can provide useful insights into feature importance.
[0075] MLP is a basic neural network structure that consists of an input layer, multiple hidden layers, and an output layer. It handles complex nonlinear problems by applying nonlinear activation functions to neurons. MLP uses forward propagation to predict outputs and optimizes network weights through backpropagation and gradient descent methods to reduce prediction errors. This network structure is particularly suitable for processing tabular data and classification and regression tasks that require the model to learn complex patterns.
[0076] The mean absolute error (MAE) and root mean square error (RMSE) are used as evaluation indicators, and their expressions are:
[0077]
[0078] Where: e MAE is the absolute mean error; e RMSE is the root mean square error; n represents the number of test samples; y true is the actual power measurement value; y pred is the actual power prediction value.
[0079] Figure 4-7They are the prediction results of DB1, DB2, DB4 and DB5 data respectively. Table 1 shows the results of each evaluation index.
[0080] Table 1 Evaluation index table of each model
[0081]
[0082] like Figure 4 As shown in Table 2, the average absolute error and root mean square error of the prediction results of the DB data through the model stacking prediction model are significantly reduced compared with other prediction models. Among them, the average absolute error of the DB1 data is reduced by 31.381% and 20.622% compared with the random forest model and the MLP model, and the root mean square error is reduced by 26.989% and 19.311%. The average absolute error of the DB2 data is reduced by 19.950% and 10.933%, and the root mean square error is reduced by 19.962% and 5.258%. The average absolute error of the DB4 data is reduced by 16.467% and 19.882%, and the root mean square error is reduced by 17.077% and 22.124%. The average absolute error of the DB5 data is reduced by 3.504% and 15.294%, and the root mean square error is reduced by 5.256% and 22.953%.
[0083] The prediction model combining dynamic feature selection and model stacking has been further improved compared to model stacking alone. The average absolute error of DB1 data has been reduced by 24.295%, and the root mean square error has been reduced by 26.631%. The average absolute error of DB2 data has been reduced by 29.403%, and the root mean square error has been reduced by 25.327%. The average absolute error of DB4 data has been reduced by 30.842%, and the root mean square error has been reduced by 28.556%. The average absolute error of DB5 data has been reduced by 37.280%, and the root mean square error has been reduced by 36.543%.
[0084] In summary, this study proposes a prediction method that combines dynamic feature selection and model stacking to address the problems encountered in actual industrial production environments, such as large data fluctuations, complex data relationships and high dimensions, and few problem samples. This method can effectively reduce the data dimension, while selecting the feature combination with the highest contribution and exploring the potential relationship between the data. More accurate predictions can be achieved through the interaction between the base model and the metamodel. Compared with traditional feature selection methods, dynamic feature selection can deeply explore the relationship between features and retain useful features to the maximum extent. Model stacking solves the problem of low prediction accuracy and unstable performance of a single model when facing large industrial fluctuations through the linkage of the base model and the metamodel. Therefore, this combined prediction method is suitable for industrial production and has a higher reference value.
[0085] It can be understood that the above specific description of the present invention is only used to illustrate the present invention and is not limited to the technical solutions described in the embodiments of the present invention. Those skilled in the art should understand that the present invention can still be modified or replaced by equivalents to achieve the same technical effects; as long as the use requirements are met, they are within the protection scope of the present invention.
Claims
1. A tire green quality parameter prediction method based on dynamic feature selection and model stacking, characterized in that The following steps are involved: Step 1: Load the tire quality parameters from the data set and clean the data; Step 2: Use random forest to score the importance of features. Step 3: Dynamic feature selection sorts the scored data, establishes different feature combinations, and then selects the optimal solution through cross-validation verification, and divides the selected feature combination into a test set and a training set. Step 4: Send the training set to the base model layer in the model stacking algorithm to obtain preliminary meta-features. Step 5: Use the meta-model layer to train the meta-features to obtain a prediction model for tire quality parameters.
2. The method for predicting green tire quality parameters based on dynamic feature selection and model stacking according to claim 1, characterized in that: In step 2, scoring the feature importance includes the following steps: Step 2-1: Use the average impurity reduction method through the random forest model to evaluate the importance of the feature and generate a feature importance score. The mathematical expression is as follows: Where parent represents the impurity of the parent node, left and right represent the impurity of the left and right child nodes after segmentation, respectively. left and N right is the number of samples of the left and right child nodes, N is the total number of samples of the parent node, and I represents the impurity function. (2) is the impurity function, where P i represents the probability of the i-th event in the system.
3. The method for predicting green tire quality parameters based on dynamic feature selection and model stacking according to claim 1, characterized in that: Furthermore, in step 3, the feature importance ranking includes the following steps: Step 3-1: Score the feature importance obtained in step 2-1, and through dynamic feature selection, establish different feature combinations with the number of features gradually increasing from 1 to including all features. Step 3-2: Use 5 rounds of cross-validation to validate feature combinations with different numbers of features and select the feature combination with the greatest impact. Step 3-3: Divide 20% of the selected feature combinations into a test set and 80% into a training set. Step 3-4: The training set is sent to the model stacking algorithm for training.
4. The method for predicting green tire quality parameters based on dynamic feature selection and model stacking according to claim 1, characterized in that: Furthermore, in step 4, the preliminary training of the data includes the following steps: Step 4-1: Use the base model layer to make a preliminary prediction on the data set after feature selection; in this layer, random forest, GBDT, and AdaBoost are three different base models that make independent predictions on the data and generate corresponding meta-features. The prediction mathematical expression is as follows: Among them, x is the tire quality parameter data, h i (x) represents the prediction result of the i-th decision tree for input x, v is the step coefficient (learning rate), M is the total number of iterations, and α m is the weight of the mth weak learner, h m (x) is the prediction result of the mth weak learner, and H(x) is the prediction result.
5. The method for predicting green tire quality parameters based on dynamic feature selection and model stacking according to claim 1, characterized in that: Furthermore, in step 5, training the meta-features includes the following steps: Step 5-1: MLP is used as a meta-model to integrate the meta-features obtained in step 4-1 and further predict to obtain a prediction model.