Tire production key parameter prediction method based on autonomous optimization learning

By combining reinforcement learning and deep learning methods, key parameters in tire production are adjusted in real time, and the problem that traditional methods cannot adapt to dynamic changes is solved, efficient and accurate parameter control is achieved, and the consistency of production efficiency and product quality is improved.

CN120297467AActive Publication Date: 2025-07-11OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202510348640.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

In the existing tire production process, traditional supervised learning methods cannot effectively adapt to the dynamically changing production environment, resulting in large errors in key parameters prediction, serious waste of resources, and inefficient reliance on manual adjustment.

Method used

The PPO algorithm of reinforcement learning is adopted combined with the deep learning model LSTNet, and through independent optimization learning, key parameters are adjusted in real time, to reduce prediction errors, and improve production stability.

Benefits of technology

It realizes high-precision prediction of key parameters, reduces waste rate, improves production efficiency, reduces resource waste, and promotes the intelligence and automation of industrial manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297467A_ABST
    Figure CN120297467A_ABST
Patent Text Reader

Abstract

A tire production key parameter prediction method based on autonomous optimization learning comprises the steps of collecting key parameter data and influence factor data on a tire production line, training and verifying a deep learning model LSTNet, constructing an environment used for training a reinforcement learning agent, and compiling a reward function. An action space, an observation space and a state in the environment are defined, a reinforcement learning PPO intelligent agent is trained based on test data, so that the reinforcement learning PPO intelligent agent can continuously learn and adjust production parameters in the virtual dynamic production environment to optimize the prediction precision of key parameters, and the trained reinforcement learning intelligent agent and an LSTNet model are used for predicting and autonomously optimizing the key parameters. According to the method, two advanced technologies of deep learning and reinforcement learning are creatively combined, so that the intelligent agent can self-learn and optimize the behavior strategy during interaction with the environment, and the method is particularly suitable for industrial manufacturing environments which have complex system structures, dynamically change and are difficult to completely model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting key parameters in tire production based on autonomous optimization learning. This method combines a reinforcement learning algorithm to reduce the prediction error of key parameters by autonomously optimizing the remaining parameters in the data, belonging to the technical fields of deep learning and tire production. Background Art

[0002] Industrial manufacturing processes usually have complex structures and large scales, involving multiple subsystems working together to complete the overall manufacturing task. With the development of manufacturing technology, the operating parameters within the subsystems are increasing and redundant, and some parameters are uncontrollable. Therefore, how to achieve autonomous optimization of parameters while ensuring the efficient cooperation of each subsystem has become an important research topic. However, due to the complexity of the parameters and operating mechanisms within each subsystem, how to effectively perform autonomous optimization learning to improve the overall performance of the system remains a challenging problem.

[0003] With the rapid development of artificial intelligence technology, autonomous learning methods based on supervised learning and reinforcement learning have gradually become an important research direction in intelligent systems. In traditional industrial manufacturing processes, especially in complex production environments, prediction networks usually rely on traditional supervised learning methods. However, supervised learning methods usually require a large amount of labeled data, and there are certain limitations when dealing with dynamic and high-dimensional complex environments. Therefore, how to effectively cope with the ever-changing industrial manufacturing environment and improve the adaptive ability of the system has become a current problem.

[0004] Taking the tire production industry as an example, the production process is complex, involving multiple subsystems and key parameters (such as the width and thickness of the film, the temperature of the descending conveyor belt). The traditional production process cannot effectively adapt to the real-time changes of these key parameters, resulting in a large amount of resource waste. Existing parameter optimization methods usually rely on traditional mathematical models, and these models are often too complex to adapt to the changes in the production process. Summary of the Invention

[0005] The present invention provides a method for predicting key parameters in tire production based on autonomous optimization learning, aiming to optimize the error of key parameter prediction by interacting with the environment in a dynamically changing tire production environment. This method uses the PPO algorithm of reinforcement learning to adjust the parameters in real time, enabling the prediction network LSTNet to adapt to the changing factors in the production process, gradually reducing the prediction error of key parameters, and thus improving the prediction accuracy and the stability of the production process.

[0006] The method for predicting key parameters in tire production based on autonomous optimization learning includes:

[0007] Step 1. Collect the key parameter data and their influencing factor data on the tire production line, clean and preprocess the collected data, and divide it into training data and test data.

[0008] It is advisable that the data be no less than 10,000 pieces, and the division ratio of training data to test data is 3:1.

[0009] The key parameters include the width of the tire film, the thickness of the tire film, and the temperature of the tire descending conveyor belt.

[0010] The influencing factors of the width of the tire film are: the main motor current of the calender (A), the calender speed (mpm), the calender rubber output temperature (℃), the calender pulling roll speed (mpm), the take-up conveyor belt speed (mpm), the calender trimming knife spacing (mm), the upper roll temperature control temperature of the calender (℃), the lower roll temperature control temperature of the calender (℃), the main motor current of the extruder (A), the extruder speed (rpm), the head pressure of the extruder (Mpa), the head temperature of the extruder (℃), the end pressure of the extruder screw (Mpa), the head temperature control temperature of the extruder (℃), the plasticizing 1 temperature control temperature of the extruder (℃), the plasticizing 2 temperature control temperature of the extruder (℃), the barrel temperature control temperature of the extruder (℃), and the screw temperature control temperature of the extruder (℃). Among them, the width is the key parameter to be predicted; the calender trimming knife spacing is the manipulated variable and also the parameter we need to optimize.

[0011] The influencing factors of the thickness of the tire film are: the head temperature of the extruder (℃), the head pressure of the extruder (Mpa), the end pressure of the extruder screw (Mpa), the extruder current (A), the extruder speed (mpm), the calender roll gap (mm), the calender rubber output temperature (℃), the calender current (A), the calender speed (mpm), the temperature of the extruder screw section (℃), the plasticizing temperature of the extruder (℃), the extrusion section temperature of the extruder (℃), the upper roll temperature of the calender (℃), and the lower roll temperature of the calender (℃). Among them, the thickness is the key parameter to be predicted; the calender roll gap is the manipulated variable and also the parameter we need to optimize.

[0012] The influencing factors of the temperature of the tire descending conveyor belt are: the width measurement at the bonding place (mm), the return air temperature (℃), the supply air temperature (℃), the mixed air temperature (℃), the natural air temperature (℃), the air cooling 1 temperature (℃), the air cooling 2 temperature (℃), the air cooling 3 temperature (℃), the air cooling 4 temperature (℃), the return air valve opening (℃), the natural air valve opening (℃), and the air cooling percentage (%). Among them, the temperature of the descending conveyor belt is the key parameter to be predicted; the return air valve opening, the natural air valve opening, and the air cooling percentage are the manipulated variables and also the parameters we need to optimize.

[0013] Step 2. Use the training data collected in Step 1 to deeply learn the LSTNet model so that the model can accurately predict the key parameters from the given input parameters.

[0014] The structure of the LSTNet is as follows: LSTNet is a hybrid neural network architecture that combines a Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), and Skip GRU (Skip-GRU), mainly used to process time series data. The core idea of this model is to extract the local features of the time series through the convolutional layer, capture the time dependence through the LSTM, and enhance the modeling ability of local time dependence through the Skip-GRU. Finally, the model outputs the prediction result through the fully connected layer. This model uses the sliding window technique (window size of 96) to process time series data, extracts local features through three one-dimensional convolutional layers. After each convolution, the ReLU activation function is used to introduce non-linearity, and dimensionality reduction is performed through the max pooling layer. The output dimension of the convolutional layer gradually decreases, and finally, the local features of the time series are extracted. The output of the convolutional layer is used for time dependence modeling through a unidirectional LSTM layer, and is processed by layer normalization and Dropout regularization. The Skip_GRU module further enhances the modeling ability of local dependence through the skip window. The outputs of the LSTM and Skip-GRU are concatenated in the feature dimension (i.e., the last dimension), that is, the outputs of the LSTM and Skip-GRU are connected in the feature dimension, and the fully connected layer is used to map the outputs of the LSTM and Skip-GRU to a single output value to obtain the final prediction result.

[0015] Step 3. Use the test data collected in Step 1 to verify the deeply learned model in Step 3, evaluate its prediction error, so as to obtain the performance feedback of the prediction model.

[0016] Step 4. Construct an environment for training a reinforcement learning agent and write a reward function;

[0017] The action space in the environment is the set of actions that the agent can choose at each time step;

[0018] The observation space in the environment defines the environmental information that the agent can obtain at each time step.

[0019] The state in the environment represents the configuration of the parameters to be optimized at a certain time step.

[0020] The reward function calculates the error based on the prediction result of the LSTNet, and its calculation formula is:

[0021] When the error decreases, the calculation formula is:

[0022]

[0023] When the error increases, the calculation formula is:

[0024]

[0025] where gap represents that of the current time step, and gap old represents the error of the previous time step. The calculation formula of gap is:

[0026] gap = -abs(predict - true) # (3)

[0027] where predict is the key parameter value predicted by the LSTNet model, and true is the true key parameter value.

[0028] When the key parameter is the film width, the dimension of the action space is 2, that is, the agent has two actions, increasing or decreasing the distance between the trimming knives. The observation space is represented as an array containing multiple values of the distance between the trimming knives. Each value corresponds to the distance between the trimming knives at a certain moment in the production process. The state in the environment represents the configuration of the distance between the trimming knives at a certain time step.

[0029] When the key parameter is the film thickness, the dimension of the action space is 2, that is, the agent has two actions, increasing or decreasing the roll gap of the calender. The observation space is represented as an array containing multiple values of the roll gap of the calender. Each value corresponds to the roll gap of the calender at a certain moment in the production process. The state in the environment represents the configuration of the roll gap of the calender at a certain time step.

[0030] When the key parameter is the temperature of the tire descending conveyor belt, the dimension of the action space is 8, that is, the agent has 8 actions, and each parameter has two actions, respectively increasing or decreasing the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage. The observation space is represented as an array containing multiple values of the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage. Each value corresponds to the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage at a certain moment in the production process. The state in the environment represents the configuration of the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage at a certain time step.

[0031] Step 5. Train the reinforcement learning PPO agent based on the test data so that it can continuously learn and adjust the production parameters in the virtual dynamic production environment to optimize the prediction accuracy of the key parameters. This training process specifically includes:

[0032] (1) Select a piece of data;

[0033] (2) The agent selects an action to adjust the parameters in the data that need to be optimized, obtains a new optimized parameter, updates this data with it, and inputs the updated data into the model for prediction, obtaining a new key parameter.

[0034] (3) Calculate the current reward based on the prediction error of the model and update the policy regularly.

[0035] (4) Continuously iterate steps (2) and (3) to enable the agent to gradually improve its decision-making ability through continuous trial and error and learning, and update the policy regularly.

[0036] (5) When the training time of the agent for this data reaches the pre-set number of time steps, or the prediction error of the model is less than the specified error, the training of this data is completed, and it jumps to step (1) until all data is trained.

[0037] Step 6. Use the trained reinforcement learning agent in step 5 and the verified LSTNet model in step 3 to predict and autonomously optimize the key parameters.

[0038] Collect new data on the tire production line. After cleaning and preprocessing, it can be input into the prediction network LSTNet and the reinforcement learning agent for parameter optimization. After inputting the data, the optimization process is similar to the training process, but the policy is not updated. The specific training process includes:

[0039] (1) Select a piece of data;

[0040] (2) The agent selects an action to adjust the parameters in the data that need to be optimized, obtains a new optimized parameter, updates this data with it, and inputs the updated data into the model for prediction, obtaining a new key parameter;

[0041] (3) Calculate the current reward based on the prediction error of the model.

[0042] (4) Continuously iterate steps (2) and (3).

[0043] (5) When the training time of the agent for this data reaches the pre-set number of time steps, or the prediction error of the model is less than the specified error, the optimization of this data is completed, and it jumps to step (1) until all data is optimized.

[0044] Through the autonomous decision-making of the agent and the adjustment of parameters, the prediction model can be continuously optimized. When all the data is iterated once, the optimization ends. Ensure that in the actual production process, the key parameters can always be maintained within the predetermined range, improve production efficiency, reduce the scrap rate, and optimize the use of resources.

[0045] The present invention innovatively combines two advanced technologies, deep learning and reinforcement learning. By introducing reinforcement learning technology, the agent can self-learn and optimize its behavior strategy during the interaction with the environment. Compared with traditional supervised learning, reinforcement learning does not rely on a large amount of labeled data, but optimizes the decision-making process through exploration and feedback mechanisms, enabling continuous self-adjustment, parameter optimization, and improvement of prediction accuracy in complex and dynamic environments. This self-optimizing learning method is particularly suitable for industrial manufacturing environments with complex system structures, dynamic changes, and difficulty in complete modeling.

[0046] The deep learning model can effectively extract the complex relationships between key parameters and various production parameters from a large amount of historical data, while reinforcement learning dynamically adjusts the parameters through interaction with the environment. The combination of the two enables the prediction network to not only make accurate predictions under static conditions but also continuously optimize through real-time feedback during the production process, thereby further improving the prediction accuracy. In the traditional tire production process, it is often necessary for workers to adjust production parameters based on experience, which is not only inefficient but also easily affected by human factors. Through the present invention, the agent can automatically judge and adjust production parameters, reducing human error, improving production efficiency, and ensuring product quality consistency. Through the learning and optimization of the agent, the production process can be maintained in an optimal state, thereby significantly reducing the scrap rate, lowering production costs, and increasing production efficiency. Especially in the precise control of key parameters, it avoids resource waste caused by unstable or excessive parameter adjustment, promoting the efficient use of resources. The present invention promotes the intelligentization and automation of the industrial manufacturing process through the application of intelligent optimization algorithms, conforming to the intelligent manufacturing concept of Industry 4.0. Through autonomous learning and real-time optimization, the production system can more flexibly and efficiently respond to changes in market demand, enhancing the adaptive ability of the production line and improving the intelligent level of the industry. Developing a parameter self-optimizing learning method based on the combination of reinforcement learning and traditional supervised learning can not only effectively solve the key problems in current industrial manufacturing but also provide a new technical path for the intelligentization and automation process in the future industrial production field. Brief Description of the Drawings

[0047] Figure 1 It is the process framework diagram of the present invention.

[0048] Figure 2 It is the deep learning model architecture diagram of the present invention.

[0049] Figure 3 It is the prediction effect of the model taking the width of tire film as an example in the present invention.

[0050] Figure 4 It is the parameter self-optimization flow chart taking the width of tire film as an example in the present invention.

[0051] Figure 5 This is the optimization effect of the reinforcement learning agent taking the tire film width as an example in the present invention. Detailed implementation manners

[0052] The process of the present invention is as Figure 1 shown. First, collect the key parameter data and its influencing factor data on the tire production line, clean and preprocess the collected data, and divide it into training data and test data.

[0053] The required data includes key parameters and their influencing factors. The specific data is shown in the following table. The corresponding relationship between the labels and the parameters is: A: tire film width (mm), B: main motor current of calender (A), C: calender speed (mpm), D: calender rubber output temperature (°C), E: pulling roll speed of calender (mpm), F: take-up conveyor speed (mpm), G: calender trimming knife spacing (mm), H: upper roll temperature control temperature of calender (°C), I: lower roll temperature control temperature of calender (°C), J: main motor current of extruder (A), K: extruder speed (rpm), L: head pressure of extruder (Mpa), M: head temperature of extruder (°C), N: pressure at the end of extruder screw (Mpa), O: head temperature control temperature of extruder (°C), P: plasticization 1 temperature control temperature of extruder (°C), Q: plasticization 2 temperature control temperature of extruder (°C), R: barrel temperature control temperature of extruder (°C), S: screw temperature control temperature of extruder (°C). Among them, A: tire film width (mm) is the key parameter to be predicted, and G: calender trimming knife spacing (mm) is the manipulated variable and also the parameter we need to optimize. The collected data is shown in Table 1 below.

[0054] Table 1 Collect key parameter data and its influencing factor data on the tire production line

[0055]

[0056] These data reflect the variation law of the key parameters under different production conditions. The collected data needs to be cleaned and preprocessed to ensure the integrity and accuracy of the data. Then, divide these data into training data and test data according to a ratio of 3:1 to ensure that the model in the training process can be effectively verified.

[0057] Use the training data to train the LSTNet network. LSTNet is a neural network architecture based on long short-term memory (LSTM) units with a sliding window, and the sliding window size is 96. Its network architecture diagram is as Figure 2As shown. This model combines a Convolutional Neural Network (CNN), Long Short-Term Memory Network (LSTM), and Skip_GRU (Jumping GRU) network with the aim of effectively capturing local features and long-term dependencies in time series data for processing time series data. The sliding window technique is used to generate fixed-size windows on time series data, ensuring that the network can process the local features of the data. This method allows the model to use historical data in the time series for prediction without relying on the entire sequence. The network first extracts local features of the time series data through three one-dimensional convolutional layers. After each layer of convolution, a max-pooling layer follows to reduce the data dimension through the pooling operation while retaining the most important features. The ReLU activation function is used to introduce non-linearity. The features extracted by the convolutional layers are fed into a unidirectional LSTM layer (GRU1) for time-dependency modeling. LSTM is good at capturing long-term dependencies in the data, ensuring that the data is processed in the correct format when input. The output is processed by layer normalization and Dropout regularization is used to prevent overfitting. The network further enhances its ability to model local dependencies in the time series using the Skip_GRU module. This module captures features at different time steps by dividing the time series data into multiple jumping windows and modeling them through GRU. Finally, the outputs of LSTM and Skip-GRU are concatenated together as the input to the next layer. After the above feature extraction and modeling, the output of the network passes through a fully connected layer for the final prediction. Through this layer, the network can map the high-dimensional time series features to a single output value. Finally, the outputs of LSTM and GRU are further mapped to the final prediction result.

[0058] The deep learning model is trained using the collected training data. During the training process, the model adjusts the weights and biases in the network by optimizing the loss function (mean squared error) to minimize the prediction error. Through the backpropagation algorithm and gradient descent method, the model gradually optimizes its parameters so that the prediction results can be as close as possible to the true values. During the training process, the model continuously updates its internal weights to better learn the relationship between the input parameters and the key parameters. After training is completed, the deep learning model LSTNet can predict the key parameters according to different production conditions and provide a preliminary prediction result.

[0059] The trained deep learning model LSTNet is validated using the collected test data. By inputting the test data into the model, its prediction error is evaluated. The results are as Figure 3 shown. The mean absolute error (MAE) and the qualified rate of the prediction results are used to measure the prediction performance of the model. The purpose of validation is to judge the generalization ability of the model on unseen data and ensure that its prediction results have sufficient accuracy. If the prediction error is too large, it may be necessary to adjust the deep learning model or increase the amount of training data to improve the performance of the model.

[0060] Build an environment for training a reinforcement learning agent; taking the width of the tire film as an example, the action space in the environment is the set of actions that the agent can choose at each time step. Specifically, the agent can perform the following two actions: 0: increase the distance between the trimming knives; 1: decrease the distance between the trimming knives. The action space is in a discrete form, and the agent can only choose one of the actions to increase or decrease the parameter of the distance between the trimming knives each time. The observation space in the environment defines the environmental information that the agent can obtain at each time step. The observation space is represented as an array containing multiple values of the distance between the trimming knives. Each value corresponds to the distance between the trimming knives at a certain moment in the production process. The dimension of this array is determined by the sliding window size of LSTNet. When performing each action, the adjustment amount of the distance between the trimming knives by the agent is set to 0.1 mm. That is, when increasing or decreasing the distance between the trimming knives each time, the actual change amount of the trimming knives is 0.1 mm. The state in the environment represents the configuration of the distance between the trimming knives at a certain time step.

[0061] Write a reward function, which calculates the error according to the prediction result of LSTNet, and its calculation formula is:

[0062] When the error decreases, the calculation formula is:

[0063]

[0064] When the error increases, the calculation formula is:

[0065]

[0066] Among them, gap represents the error, and the calculation formula of gap is:

[0067] gap = -abx(predict - true) # (3)

[0068] Among them, predict is the predicted value of the LSTNet model, and true is the true value.

[0069] Based on the obtained model test results, use the test data to further train the reinforcement learning agent. The task of the agent is to continuously learn and adjust the production parameters in the virtual dynamic production environment to optimize the prediction accuracy of the key parameters. Select one piece of data for training each time.

[0070] Reinforcement learning agents learn by interacting with the environment. The decision-making process of the agent includes the following steps. Action selection: The agent selects an action based on the current state. The action is to adjust the production parameters related to the key parameters. Taking the tire film width as an example, it is to adjust the trimming knife spacing to change the state in the production environment. Reward calculation: The agent evaluates the selected action based on the current prediction error and the preset reward function. For example, when the production parameters selected by the agent reduce the prediction error of the key parameters, a positive reward is given; otherwise, a negative reward is given. State update and strategy adjustment: After selecting an action, the agent updates its state and adjusts its strategy based on the new reward information. Through this trial and error process, the agent continuously optimizes its decision-making process, thereby gradually improving the accuracy of production parameter adjustment. Through continuous interaction and learning, the agent can gradually identify the optimal parameter adjustment strategy in the production process, and ultimately achieve precise control of key parameters. The agent will perform this operation once in each time step. When the time step number of this data is used up or the error of the model prediction result is less than 1 mm, the training of this data is terminated. This cycle continues until all test data has been trained.

[0071] After the reinforcement learning agent is fully trained and the prediction error of the LSTNet model is small enough, it can enter the autonomous optimization stage. Figure 4 As shown. The optimization process is similar to the training process of the agent, but the strategy will not be updated. When the time step number of this data is used up or the error of the model prediction result is less than the specified error, the optimization of this data is terminated. This cycle is repeated until all data are optimized. At this point, the agent can use the optimization strategy and deep learning model it has learned to make real-time adjustments to ensure that key parameters are always within the predetermined standard range. In the actual production process, the deep learning model will continuously monitor the changes in key parameters and make predictions based on the current production data. At the same time, the reinforcement learning agent continuously adjusts the production parameters based on the real-time prediction results. Taking the tire film width as an example, the trimming knife spacing is adjusted to minimize the prediction error of key parameters. When the model prediction error is lower than the predetermined minimum error, the optimization process ends. The optimization results are shown as follows Figure 5 As shown. Through this self-optimization process, the production process can automatically adjust the production parameters without human intervention, maintaining the high accuracy of key parameters and the stability of the production process. This not only improves production efficiency, but also effectively reduces scrap rate and resource waste, and ultimately achieves the optimization and intelligence of the production process. The true width value, model prediction value and optimized width prediction value are shown in Table 2 below.

[0072] Table 2 Comparison of true width value, model prediction value and optimized width prediction value

[0073]

[0074]

[0075] After the agent is added for optimization, the error between the predicted value and the true value of the width is significantly reduced. Specifically, whether for a relatively small width value (such as 488.80 mm) or a relatively large width value (such as 523.20 mm), in the prediction model combined with agent optimization, the prediction error is effectively controlled within 1 mm. For example, when the true width value is 488.80 mm, the predicted value of the original LSTNet model may fluctuate between 487.02 mm and 487.91 mm, with a relatively large error. However, after the agent optimization is added, the predicted value is precisely adjusted to a range extremely close to the true value, and the error is reduced to less than 1 mm. This improvement is reflected in every data point. Whether for a relatively small width value or a relatively large width value, agent optimization ensures the accuracy and consistency of the prediction.

Claims

1. A method for predicting key parameters in tire production based on autonomous optimization learning, characterized in that It includes the following steps: Step 1. Collect the key parameter data and its influencing factor data on the tire production line, clean and preprocess the collected data, and divide it into training data and test data; Step 2. Use the data collected in Step 1 to train the deep learning model LSTNet so that the model can accurately predict the key parameters from the given input parameters; Step 3. Use the test data collected in Step 1 to verify the trained deep learning model in Step 3, evaluate its prediction error, and thus obtain the performance feedback of the prediction model; Step 4. Construct an environment for training the reinforcement learning agent and write a reward function; The action space Action Space in the environment is the set of actions that the agent can choose at each time step; The observation space Observation Space in the environment defines the environmental information that the agent can obtain at each time step; The state State in the environment represents the configuration of the parameters to be optimized at a certain time step; Step 5. Train the reinforcement learning PPO agent based on the test data so that it can continuously learn and adjust the production parameters in the virtual dynamic production environment to optimize the prediction accuracy of the key parameters; Step 6. Use the trained reinforcement learning agent in Step 5 and the verified LSTNet model in Step 3 to predict and autonomously optimize the key parameters, that is, collect new data on the tire production line, and after cleaning and preprocessing, it can be input into the prediction network LSTNet and the reinforcement learning agent for parameter optimization.

2. The key parameter prediction method for tire production based on autonomous optimization learning according to claim 1, characterized in that It is advisable that the data collected in Step 1 be no less than 10,000, and the division ratio of training data to test data is 3:

1.

3. The tire production key parameter prediction method based on autonomous optimization learning as described in claim 1 is characterized by the steps 1 The key parameters include the width of the tire film, the thickness of the tire film, and the temperature of the tire descending conveyor belt; among them, The influencing factors of the width of the tire film are: the temperature of each temperature control channel of the extruder in °C, the head pressure of the extruder in Mpa, the head temperature of the extruder in °C, the current of the extruder screw motor in A, the screw speed of the extruder in rpm, the temperature of the calender rolls in °C, the speed of the calender pressure roll in mpm, the current of the calender pressure roll in A, the temperature of the calender rubber output in °C, the distance between the trimming knives of the calender in mm, the roll gap of the calender in mm, the speed of the pulling roll of the calender in mpm, and the speed of the receiving conveyor belt in mpm; among them, the width is the key parameter to be predicted; the distance between the trimming knives of the calender is the manipulated variable and also the parameter to be optimized; The influencing factors of the thickness of the tire film are: the head temperature of the extruder in °C, the head pressure of the extruder in Mpa, the pressure at the end of the extruder screw in Mpa, the current of the extruder in A, the speed of the extruder in mpm, the roll gap of the calender in mm, the temperature of the calender rubber output in °C, the current of the calender in A, the speed of the calender in mpm, the temperature of the extruder screw section in °C, the plasticizing temperature of the extruder in °C, the temperature of the extruder extrusion section in °C, the temperature of the upper roll of the calender in °C, and the temperature of the lower roll of the calender in °C; among them, the thickness is the key parameter to be predicted; the roll gap of the calender is the manipulated variable and also the parameter to be optimized; The influencing factors of the temperature of the tire lowering conveyor belt are: width measurement at the bonding position (mm), return air temperature (°C), supply air temperature (°C), mixed air temperature (°C), natural air temperature (°C), air cooling 1 temperature (°C), air cooling 2 temperature (°C), air cooling 3 temperature (°C), air cooling 4 temperature (°C), return air valve opening (°C), natural air valve opening (°C), air cooling percentage (%); among them, the temperature of the lowering conveyor belt is the key parameter to be predicted; the return air valve opening (°C), natural air valve opening (°C), and air cooling percentage (%) are the manipulated variables and also the parameters to be optimized.

4. The method for predicting key parameters in tire production based on autonomous optimization learning as described in claim 1, wherein The LSTNet structure described in step 2 is as follows: LSTNet is a hybrid neural network architecture that combines a convolutional neural network - CNN, a long short-term memory network - LSTM, and a skip GRU - Skip-GRU for processing time series data; the core idea of this model is to extract local features of the time series through the convolutional layer, capture time dependencies through the LSTM, and enhance the modeling ability of local time dependencies through the Skip-GRU; the final model outputs the prediction result through the fully connected layer; this model uses the sliding window technique to process time series data, with a window size of 96, extracts local features through three one-dimensional convolutional layers, introduces non-linearity using the ReLU activation function after each convolution, and reduces the dimension through the max pooling layer. The output dimension of the convolutional layer gradually decreases, and then the local features of the time series are extracted; the output of the convolutional layer is used for time dependency modeling through a unidirectional LSTM layer, and is processed by layer normalization and Dropout regularization; the Skip-GRU module further enhances the modeling ability of local dependencies through the skip window; the outputs of the LSTM and Skip-GRU are concatenated in the feature dimension, that is, the outputs of the LSTM and Skip-GRU are connected in the feature dimension, and the fully connected layer is used to map the outputs of the LSTM and Skip-GRU to a single output value to obtain the final prediction result.

5. The key parameter prediction method for tire production based on autonomous optimization learning as claimed in claim 4, characterized in that In step 4 described above, When the key parameter is the film width, the dimension of the action space is 2, that is, the agent has two actions, increasing or decreasing the distance between the trimming knives; the observation space is represented as an array containing multiple values of the distance between the trimming knives; each value corresponds to the distance between the trimming knives at a certain moment during the production process. The state State in the environment represents the configuration of the distance between the trimming knives at a certain time step; When the key parameter is the film thickness, the dimension of the action space is 2, that is, the agent has two actions, increasing or decreasing the roll gap of the calender. The observation space is represented as an array containing multiple values of the roll gap of the calender; each value corresponds to the roll gap of the calender at a certain moment during the production process; the state State in the environment represents the roll gap configuration of the calender at a certain time step; When the key parameter is the temperature of the tire descending conveyor belt, the dimension of the action space is 8, that is, the agent has 8 actions, and each parameter has two actions, which are to increase or decrease the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage respectively. The observation space is represented as an array containing the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage. Each value corresponds to the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage at a certain moment during the production process; the state State in the environment represents the configuration of the opening degree of the return air valve, the opening degree of the natural air valve, and the air-cooling percentage at a certain time step.

6. The tire production key parameter prediction method based on autonomous optimization learning according to claim 1, characterized in that In the said step 4, the reward function calculates the error according to the prediction result of LSTNet, and its calculation formula is: When the error decreases, the calculation formula is: , When the error increases, the calculation formula is: , Among them, gap indicating the current time step, indicating the error of the previous time step, gap The calculation formula of is as follows: , Among them, is the key parameter value predicted by the LSTNet model, and is the real key parameter value.

7. The key parameter prediction method for tire production based on autonomous optimization learning according to claim 1, characterized in that In the said step 5, the training process specifically includes: (1) Select a piece of data; (2) The agent selects an action to adjust the parameter to be optimized in the data, obtains a new optimized parameter, updates this piece of data with it, and inputs the updated data into the model for prediction, and a new key parameter is predicted; (3) Calculate the current reward according to the prediction error of the model, and update the policy regularly; (4) Continuously iterate steps (2) and (3) to enable the agent to gradually improve its decision-making ability in continuous trial and error and learning, and update the policy regularly; (5) When the time for the agent to train this piece of data reaches the pre-set number of time steps, or the prediction error of the model is less than the specified error, the training of this piece of data is completed, and it jumps to step (1), until all data is trained.

8. The method for predicting key parameters in tire production based on autonomous optimization learning as described in claim 1, characterized in that In the said step 6 for prediction and autonomous optimization, after the data is input, the optimization process is similar to the training process, but the policy will not be updated. Specifically, it includes: (1) Select a piece of data; (2) The agent selects an action to adjust the parameter to be optimized in the data, obtains a new optimized parameter, updates this piece of data with it, and inputs the updated data into the model for prediction, and a new key parameter is predicted; (3) Calculate the current reward according to the prediction error of the model; (4) Continuously iterate steps (2) and (3); (5) When the time for the agent to train this piece of data reaches the pre-set number of time steps, or the prediction error of the model is less than the specified error, the optimization of this piece of data is completed, and it jumps to step (1), until all data is optimized.

9. The key parameter prediction method for tire production based on autonomous optimization learning according to claim 8, characterized in that In the said step 6, through the autonomous decision-making of the agent and the adjustment of the parameters, the prediction model is continuously optimized. When all the data is iterated once, the optimization ends.

Citation Information

Patent Citations

  • Training method of reinforced learning model, node, system and storage medium

    CN109952582A

  • Intelligent reflection surface assisted MIMO covert communication system and parameter optimization method

    CN113726471A

  • Milling parameter optimization method based on deep reinforcement learning

    CN114200889A

  • Reinforced learning cement production process control method based on near-end strategy optimization

    CN117850355A

  • Intelligent control method based on knowledge distillation and multi-agent reinforcement learning

    CN119126577A

Cited By

  • Rubber tire production group intelligent optimization method based on reinforcement learning

    CN121072836A