Training self-adjusting multi-module tobacco primary processing production outlet moisture content prediction method
By applying a deep learning model based on Transformer on the tobacco wire making production line, combining multi-sensor data and optimized model parameters, the problems of response speed limit, monitoring delay and noise interference in the tobacco wire drying process are solved, and high accuracy prediction of the moisture content of tobacco wire outlets and the stability of the production process are achieved.
Patent Information
- Application Number
- CN202510194377.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The traditional PID control system has problems such as response speed limitation, moisture monitoring delay, noise interference and integral saturation in the tobacco wire drying process, resulting in fluctuations in the moisture content of the tobacco wire outlet.
Using a deep learning model based on Transformer, we set up multiple sensors on the tobacco leaf recovery production line to collect data, perform data preprocessing and time segment division, build multiple original models, and optimize model parameters through Soft RIME and Hard RIME update mechanisms, and finally realize real-time prediction of the moisture content of tobacco silk production exports.
It significantly improves the accuracy of prediction of the moisture content of tobacco outlets, reduces fluctuations, improves the stability and response speed of the production process, and is better than the effects of traditional machine learning and deep learning models.
Smart Images

Figure CN120123644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the outlet moisture content of self - regulating multi - module tobacco leaf processing production, which is used to combine the tobacco leaf processing industry and deep learning technology. Background Art
[0002] Generally speaking, cut tobacco will experience heating, humidifying, drying and winnowing during the whole drying stage, as Figure 1 shown. In order to improve the consistency of tobacco before drying, it is first pretreated with hot steam in a humidity increasing regulator. The pneumatic dryer is a widely used dehydration equipment in tobacco processing, which can accelerate the loss of moisture by adjusting process parameters. The moisture content level of the inlet tobacco is about 18.6%, and the drying temperature in the drying cylinder is set in the range of 129°C to 134°C. The temperature in the workshop is kept constant at about 29°C throughout the year. The tobacco falls downward in the cylinder under the action of gravity. In this drying and processing mode, the original tobacco can be heated evenly. Compared with the thin - plate drying process, the pneumatic drying technology has a faster processing speed. During the whole process, the main factors affecting the moisture content include the valve opening, the hot air temperature and the steam pressure. After the dehydration operation, the temperature of the semi - finished product can be as high as 69°C, and the tobacco must be cooled in the machine and winnowed. According to the need of actual measurement accuracy, the humidity sensor is actually installed behind the winnower. The humidity measuring instrument used is the NDC TM710 measuring instrument, which is a near - infrared measuring device. In the method, the measured value of the moisture content is set to two decimal places, which is one of the most important evaluation indexes in the drying process. The measurement of the moisture content and the abnormal detection are carried out after the drying operation, and there is a large delay. As Figure 2 shown, the measuring point is located behind the winnower, and it takes several minutes (i.e., time T) to obtain the measured value from the dryer outlet. During the delay time, a large amount of under - dried tobacco flows through the production line before preventive and corrective measures are taken. Further, the parameters of the winnower are preset without feedback control, and the moisture of the tobacco after leaving the dryer is not controlled. Therefore, the time delay in the humidity measurement in the drying process control easily leads to the fluctuation of the outlet moisture content of the cut tobacco.
[0003] At present, the tobacco leaf processing production workshop adopts PID feedback control regulation to control the outlet moisture content of the cut tobacco during drying. The outlet moisture content of the cut tobacco is monitored in real time by a moisture meter. When the outlet moisture content fluctuates, the material flow rate input into the dryer is feedback - regulated through a PID controller, so as to realize the control of the stability of the outlet moisture of the cut tobacco.
[0004] The PID controller is a commonly used control algorithm in the tobacco leaf conditioning production workshop. Through the combined action of the proportional (P), integral (I), and derivative (D) parts, this controller precisely adjusts the material flow rate of the tobacco dryer. The proportional control part responds quickly in proportion to the error between the outlet moisture content and the set value; the integral control part is responsible for eliminating the steady-state error to ensure long-term stability; the derivative control part makes adjustments in advance according to the changing trend of the error to improve the stability and response speed of the system. Through this control strategy, the PID controller can effectively maintain the stability of the tobacco leaf outlet moisture content.
[0005] However, the method based on PID feedback control regulation still has the following defects:
[0006] 1). Response speed limitation: PID needs to adjust the control strategy based on historical data, which often results in the inability to handle quality anomalies in a timely manner.
[0007] 2). Monitoring delay exists: In actual production equipment, the moisture meter is installed at the outlet of the tobacco dryer, while the dehydration of tobacco leaves is carried out inside the tobacco dryer. There is a monitoring delay in monitoring the outlet moisture content of tobacco leaves, which causes a lag in PID adjustment.
[0008] 3). Sensitive to noise and external disturbances: Especially the derivative part is easily affected by high-frequency noise, which may lead to system instability. Summary of the Invention
[0009] The present invention aims to solve the problems of response speed limitation, moisture monitoring delay, noise interference, and integral saturation existing in the traditional PID control system in the tobacco drying process. By using a deep learning model based on Transformer to model the tobacco drying process with actual production line data, a training self-regulating multi-module prediction method for the outlet moisture content of tobacco leaf conditioning production is provided.
[0010] To achieve the invention purpose of this application, the following technical solutions are adopted in this application:
[0011] A training self-regulating multi-module prediction method for the outlet moisture content of tobacco leaf conditioning production of the present invention sets multiple sensors on the production line of tobacco leaf rewetting, which are used to monitor the outlet moisture content, tobacco leaf process flow rate, outlet temperature, winnover steam valve opening, anti-caking valve opening, purification gas valve opening, cleaning water valve opening, negative pressure, oxygen content, cut tobacco accumulation, winnover working steam flow rate, and anti-caking working steam flow rate, where: the method includes the following steps:
[0012] (I) Data collection and storage
[0013] Simultaneously collect the moisture content of the discharged material, the process flow rate of cut tobacco, the temperature of the discharged material, the opening degree of the winnover steam valve, the opening degree of the anti-caking valve, the opening degree of the purification gas valve, the opening degree of the cleaning water valve, the negative pressure, the oxygen content, the cumulative amount of cut tobacco, the working steam flow rate of winnover, and the working steam flow rate of anti-caking using the above-mentioned multiple sensors. Compose the above twelve data at each moment into a piece of original data and store it. Use at least 400,000 pieces of original data collected continuously for one year as the data for training the model;
[0014] (2) Data preprocessing
[0015] Perform the following preprocessing on the above original data:
[0016] (1). Remove the outliers from the original data
[0017] Calculate the respective average values of each parameter in each piece of original data among the above at least 400,000. When the above parameters are within the range of their average value ± 3 times the standard deviation, they are regarded as constant values and the above parameters are retained; otherwise, they are regarded as outliers and the piece of original data including the outlier data is deleted;
[0018] (2). Standardize the parameters processed in step (1)
[0019] Process the above parameters within their ranges according to the following formula to normalize the above parameters to between 0 and ±1,
[0020]
[0021] where: Z is the value after standardization (Z-Score); X is the original parameter; μ is the mean value (mean) of each parameter within its range; σ is the standard deviation (standard deviation) of each parameter within its range;
[0022] After the above original data is completed with data preprocessing, compose multiple pieces of preprocessed data in chronological order;
[0023] (3). Divide the above preprocessed data into time segments
[0024] Use a sliding window to divide the above multiple pieces of preprocessed data. According to the chronological order of the preprocessed data, set the size of the window to be at least 10. Each time it slides and moves one, which means each time extract at least 10 consecutive pieces of preprocessed data as input features. This input feature does not include the moisture content of the discharged material; the moisture content of the discharged material after at least 10 is used as the label value, that is, the target value to be predicted. Take the above input features and the above label value as a set of associated data. Obtain multiple sets of associated data according to the above method. Use 80% of the multiple sets of associated data as the training set and 20% of the multiple sets of associated data as the test set;
[0025] (4). Construct multiple original models
[0026] (3). Composition of the model
[0027] The model is composed of multiple Temporal Convolutional Network units, multiple Gated Recurrent Units, and a Transformer layer. Multiple Temporal Convolutional Network units are connected in series, and multiple Gated Recurrent Units are connected in series. The input features are input into the model simultaneously from the starting ends of multiple Temporal Convolutional Network units and multiple Gated Recurrent Units. The ending ends of multiple Temporal Convolutional Network units and multiple Gated Recurrent Units are input into the Transformer layer simultaneously. At the same time, the input features are connected to the Transformer layer through residual connections. After passing through the Transformer layer, the predicted moisture content of the discharged material is output;
[0028] (4). Construct multiple original model parameters under constraints
[0029] Set six initial parameters of the above model, which are respectively the number of units of multiple Temporal Convolutional Network units, and its value range is an integer from 32 to 256; the number of layers of multiple Temporal Convolutional Network units, and its value range is an integer from 3 to 8; the number of units of Gated Recurrent Units, and its value range is an integer from 32 to 256; the number of layers of Gated Recurrent Units, and its value range is an integer from 3 to 8; the number of convolutional filters, and its value range is an integer from 12 to 32; and the learning rate, and its value range is 0.0001 to 0.1; One group is composed of the above six set parameters to obtain an original model; at least ten groups of the above six parameters are set to obtain at least ten original models;
[0030] Set the initial number of iterations t, t = 0, and the test score of the best model is 0;
[0031] (5). Input the training set into at least ten original models composed above respectively and train them
[0032] In the way of step (three), input the training set into the above-mentioned one original model respectively. During the training process, input the input features of the training set into one original model in batches to perform forward propagation to calculate the model output, and then update the model parameters through backpropagation. At the same time, record the loss value of the training set. After traversing all the training set data, in order to optimize the training process, introduce a learning rate and dynamically adjust the learning rate according to the change of the loss value. If the loss value of this round is lower than the lowest value of the loss values of the previous rounds, save the current model as the candidate best model at this time. Repeat the above process at least 50 rounds to fully train the model. During the whole process, the model parameters are gradually optimized, and the learning rate ensures more refined adjustment of the weight parameters in the later stage of training, thereby improving the model performance and obtaining a trained model;
[0033] Repeat the above steps, input the training set into other original models respectively to obtain at least ten trained models;
[0034] Input the test set into the above at least 10 trained models respectively to obtain their prediction results. Calculate the reciprocal of the mean squared error (MSE) value of the test set of each of the above trained models as the score of each training set, and find the trained model with the highest score as the best model;
[0035] (6). Change the above six parameters through the following steps, which are the number of units in the temporal convolutional unit, the number of layers in the temporal convolutional unit, the number of units in the gated recurrent unit, the number of layers in the gated recurrent unit, the number of convolutional filters, and the learning rate
[0036] (5). Update the best model
[0037] Compare the test score of the best model of this iteration with the test score of the previous best model, and select the model with the highest test score as the best model, and record or update the best model;
[0038] (6). Calculate the adaptive coefficient E
[0039] In each t iteration, calculate an adaptive coefficient E according to the current iteration number t, and its value is
[0040]
[0041] This adaptive coefficient E gradually increases with the increase of the iteration number, where t is the current iteration number and T is the total number of iterations;
[0042] (7). Soft RIME search
[0043] In the i-th model among the above at least ten original models, randomly generate a random number r 2, , and its value range is (-1, 1). If r 2<E, update each parameter in the original model using the following Soft RIME formula;
[0044]
[0045] where are the six updated parameters, i and j respectively represent the j-th parameter in the i-th model, and R best,j is the j-th parameter of the above optimal model, Ub ij is the maximum value of the value range of the j-th parameter in the i-th model, Lb ij is the minimum value of the value range of the j-th parameter in the i-th model, r 1 is a random number with a value range of (-1, 1), and cosθ is a coefficient with a value of where t is the current iteration number, T is the total number of iterations, r 1 and cosθ together determine the update direction of the parameter, and h is a random number with a value range of (0, 1) to ensure that the generated parameters are randomly distributed within the constraint range;
[0046] β is a coefficient, rounds to an integer, and w is a manually set parameter with a value of 5;
[0047] For the original model that does not meet the r 2 <E condition, keep the parameters of the original model unchanged;
[0048] Repeat the above operations to perform Soft RIME updates on at least the above ten training models; obtain at least the above ten new original models;
[0049] (8), Hard RIME Penetration Mechanism
[0050] After Soft RIME update, in the i-th model among at least the above ten new original models, randomly generate another random number r 3 , whose value range is (-1, 1). If r 3 <F normr (S i ), use the following Hard RIME formula to update each parameter in this new original model;
[0051]
[0052] where are the six updated parameters, and R best,j is the j-th parameter of the above optimal model;
[0053]
[0054] Wherein: F normr (S i ) represents the normalized value of the test score of the current model, representing the probability that the i-th model is selected. F(S i ) is the test score of the original model in this round of iteration; F(S) min is the minimum test score of at least ten groups of original models in this round of iteration; F(S) max is the maximum test score of at least ten groups of original models in this round of iteration;
[0055] For the original models that do not meet the condition of r 3 <F normr (S i ), keep the parameters of the original model obtained in step (8) unchanged;
[0056] Repeat the above operations, perform Hard RIME update on the above at least ten new original models, and obtain the above at least ten updated original models;
[0057] Let t = t + 1,
[0058] When t <= T, repeat step (five), replace the at least ten original models or at least ten updated original models of the previous round of iteration with the at least ten updated original models obtained most recently, and retrain the above at least ten groups of updated original models with the training set and the test set;
[0059] When t > T, obtain the above best model, which is the trained model;
[0060] (VII) Use the trained model to predict the moisture content of the tobacco cut strand production at the outlet
[0061] Collect in real time the process flow rate of cut tobacco, the outlet temperature, the opening degree of the winnover steam valve, the opening degree of the anti-caking valve, the opening degree of the purification gas valve, the opening degree of the cleaning water valve, the negative pressure, the oxygen content, the cumulative amount of cut tobacco, the working steam flow rate of winnover, and the working steam flow rate of anti-caking at at least 10 consecutive time points on the production line of tobacco leaf conditioning, and preprocess the above real-time collected data by repeating (2) of step (two); input the above data into the trained model to predict the moisture content of the tobacco cut strand production at the outlet at the next moment of the cut tobacco drying process.
[0062] For the self-regulating multi-module method for predicting the moisture content of the tobacco cut strand production at the outlet of the present invention, wherein: the standard deviation Wherein: N is the total number of data points; x i is the data of the i-th data point; μ is the mean of the population, and the calculation formula is:
[0063] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: the size of the set window is 15, which means that every time 15 consecutive pre - processed data are extracted as input features, and the outlet moisture content of the 16th pre - processed data is used as the label value; each time, one data is moved along the time arrangement order of the pre - processed data. Then, 15 consecutive pre - processed data within the window are used as the next input feature, and the outlet moisture content of the 17th pre - processed data is used as the next label value, and so on.
[0064] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: the data type and quantity input into the trained model in step (seven) are the same as those input into the original model in step (five).
[0065] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: in step (six), the number of original models is 30, and the number of iterations T is 5.
[0066] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: in step (six), the number of units of the temporal convolutional unit, the number of layers of multiple temporal convolutional units, the number of units of the gated recurrent unit, the number of layers of the gated recurrent unit, and the number of convolutional filters are all rounded to integers.
[0067] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: the mean square error where: n is the total number of samples; y i is the true value of the i - th sample; is the predicted value of the i - th sample.
[0068] The method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention, wherein: the sampling frequency of the original data is 6 seconds. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is the flow chart of the method for predicting the outlet moisture content in the training of self - regulating multi - module tobacco leaf processing production of the present invention;
[0070] Figure 2 is the schematic diagram of the model composition. DETAILED DESCRIPTION OF THE INVENTION
[0071] Such as Figure 1As shown in the figure, for the method of predicting the moisture content at the outlet of the self-adjusting multi-module tobacco leaf processing production of the present invention, a plurality of sensors are provided on the production line of tobacco leaf rewetting, which are used to monitor the moisture content at the outlet, the process flow of cut tobacco, the outlet temperature, the opening of the winnover steam valve, the opening of the anti-caking valve, the opening of the purification gas valve, the opening of the cleaning water valve, the negative pressure, the oxygen content, the cumulative amount of cut tobacco, the working steam flow of Winnover, and the working steam flow of anti-caking. The method includes the following steps:
[0072] (I) Data acquisition and storage
[0073] Use the above-mentioned multiple sensors to simultaneously collect the moisture content at the outlet, the process flow of cut tobacco, the outlet temperature, the opening of the winnover steam valve, the opening of the anti-caking valve, the opening of the purification gas valve, the opening of the cleaning water valve, the negative pressure, the oxygen content, the cumulative amount of cut tobacco, the working steam flow of Winnover, and the working steam flow of anti-caking. Form a piece of original data with the above twelve data at each moment and store it. The sampling frequency of the original data is 6 seconds, and at least 400,000 pieces of original data collected continuously for one year are used as the data for training the model;
[0074] (II) Data preprocessing
[0075] Perform the following preprocessing on the above-mentioned original data:
[0076] (1) Remove the outliers in the original data
[0077] Calculate the respective average values of each parameter in each piece of original data among the above at least 400,000. When the above parameters are within the range of their average value ± 3 times the standard deviation, they are regarded as normal values, and the above parameters are retained; otherwise, they are regarded as outliers, and the piece of original data including the outlier data is deleted;
[0078] (2) Standardize the parameters processed in step (1)
[0079] Process the above parameters within their ranges according to the following formula to normalize the above parameters to between 0 and ±1,
[0080]
[0081] Among them: Z is the value after standardization (Z-Score); X is the original parameter; μ is the mean (mean) of each parameter within its range; σ is the standard deviation (standard deviation) of each parameter within its range;
[0082] Standard deviation Among them: N is the total number of data points; x i is the data of the i-th data point; μ is the mean of the population, and the calculation formula is:
[0083] After the above raw data is preprocessed, multiple preprocessed data are formed in chronological order;
[0084] (3). Divide the above preprocessed data into time segments
[0085] Use a sliding window to divide the above multiple preprocessed data. According to the chronological order of the preprocessed data, set the window size to 15, and move one by one each time. That is, each time 15 consecutive preprocessed data are extracted as input features, and the moisture content of the discharged material of the 16th preprocessed data is used as the label value; each time move one data along the chronological order of the preprocessed data. Then, 15 consecutive preprocessed data within the window are used as the next input feature, and the moisture content of the discharged material of the 17th preprocessed data is used as the next label value, and so on. The above input features and the above label values are used as a set of associated data, and multiple sets of associated data are obtained according to the above method. 80% of the above multiple sets of associated data are used as the training set, and 20% of the above multiple sets of associated data are used as the test set;
[0086] (4). Construct multiple original models
[0087] (3). Composition of the model
[0088] The model is composed of multiple Temporal Convolutional Network units, multiple Gated Recurrent Unit units, and a Transformer layer. Multiple Temporal Convolutional Network units are connected in series, and multiple Gated Recurrent Unit units are connected in series. The input features are simultaneously input into the model from the starting ends of multiple Temporal Convolutional Network units and multiple Gated Recurrent Unit units. Multiple Temporal Convolutional Network units and multiple Gated Recurrent Unit units input the processed features into the Transformer layer simultaneously. At the same time, the input features are connected to the Transformer layer through a residual connection. Finally, the predicted moisture content of the discharged material is output by the Transformer layer;
[0089] (4). Construct multiple original model parameters under constraints
[0090] Set six initial parameters of the above model, which are the number of units of multiple temporal convolutional units, and the value range is an integer from 32 to 256; the number of layers of multiple temporal convolutional units, and the value range is an integer from 3 to 8; the number of units of the gated recurrent unit, and the value range is an integer from 32 to 256; the number of layers of the gated recurrent unit, and the value range is an integer from 3 to 8; the number of convolutional filters, and the value range is an integer from 12 to 32; and the learning rate, and the value range is 0.0001 - 0.1; a set is composed of the above six set parameters to obtain an original model; at least set thirty groups of the above six parameters to obtain thirty original models;
[0091] Set the initial number of iterations t, t = 0, and the test score of the best model is 0;
[0092] (5). Input the training set into the above - composed thirty original models respectively and train them
[0093] In the way of step (3), input the training set into the above - mentioned one original model respectively. During the training process, input the input features of the training set into one original model in batches, perform forward - propagation to calculate the model output, then update the model parameters through back - propagation, and record the loss value of the training set at the same time. After traversing all the training set data, in order to optimize the training process, introduce the learning rate and dynamically adjust the learning rate according to the change of the loss value. If the loss value of this round is lower than the lowest value of the loss values of the previous rounds, save the current model as the candidate best model at this time. Repeat the above process 100 rounds to fully train the model. During the whole process, the model parameters are gradually optimized, and the learning rate ensures more fine - tuning of the weight parameters in the later stage of training, thereby improving the model performance to obtain a trained model;
[0094] Repeat the above steps, input the training set into other original models respectively to obtain thirty trained models;
[0095] Input the test set into at least 10 of the above - mentioned trained models respectively to obtain their prediction results. Calculate the reciprocal of the mean squared error (MSE) value of the test set of each of the above - mentioned trained models, which is the score of each training set. Find the trained model with the highest score as the best model;
[0096] Mean squared error where: n is the total number of samples; y i is the true value of the i - th sample; is the predicted value of the i - th sample;
[0097] (6). Change the above six parameters through the following steps, which are the number of units of the temporal convolutional unit, the number of layers of the temporal convolutional unit, the number of units of the gated recurrent unit, the number of layers of the gated recurrent unit, the number of convolutional filters, and the learning rate
[0098] (5), Update the best model
[0099] Compare the test score of the best model in this iteration with the test score of the previous best model, and select the model with the highest test score as the best model, record or update the best model;
[0100] (6), Calculate the adaptive coefficient E
[0101] In each iteration t, calculate an adaptive coefficient E according to the current iteration number t, and its value is
[0102]
[0103] This adaptive coefficient E gradually increases with the increase of the iteration number, where t is the current iteration number and T is the total number of iterations;
[0104] (7), Soft RIME search
[0105] In the i-th model among the above thirty original models, randomly generate a random number r 2, , and its value range is (-1, 1). If r 2 < E, update each parameter in this original model using the following formula Soft RIME;
[0106]
[0107] Among them are the six updated parameters, i and j respectively represent the j-th parameter in the i-th model, R best,j is the j-th parameter of the above best model, Ub ij is the maximum value of the value range of the j-th parameter in the i-th model, Lb ij is the minimum value of the value range of the j-th parameter in the i-th model, r 1 is a random number with a value range of (-1, 1), cosθ is a coefficient, and its value is where t is the current iteration number, T is the total number of iterations, T = 5, r 1 and cosθ together determine the update direction of the parameter, h is a random number with a value range of (0, 1), ensuring that the generated parameters are randomly distributed within the constraints;
[0108] β is a coefficient, Round to an integer, w is a manually set parameter, and its value is 5;
[0109] For the original models that do not meet the condition of r 2 < E, keep the parameters of the original model unchanged;
[0110] Repeat the above operations to perform soft RIME update on the above thirty training models; obtain the above thirty new original models;
[0111] (8), Hard RIME penetration mechanism
[0112] After Soft RIME update, in the i-th model among the above thirty new original models, randomly generate a random number r 3 , whose value range is (-1, 1). If r 3 <F normr (S i ), use the following formula to perform Hard RIME update on each parameter in this new original model;
[0113]
[0114] where are the six updated parameters, and R best,j is the j-th parameter of the above best model;
[0115]
[0116] where: F normr (S i ) represents the normalized value of the test score of the current model, represents the probability that the i-th model is selected, F(S i ) is the test score of this original model in this round of iteration; F(S) min is the minimum test score of at least ten groups of original models in this round of iteration; F(S) max is the maximum test score of at least ten groups of original models in this round of iteration;
[0117] For the original models that do not satisfy the condition of r 3 <F normr (S i ), keep the parameters of the original model obtained in step (8) unchanged;
[0118] Repeat the above operations to perform Hard RIME update on the above thirty new original models to obtain the above thirty updated original models;
[0119] Let t = t + 1,
[0120] When t <= T, repeat step (five), replace the thirty original models or thirty updated original models in the previous round of iteration with the latest obtained thirty updated original models, and retrain the above latest obtained at least ten groups of updated original models using the training set and the test set;
[0121] When t > T, the above optimal model is the trained model;
[0122] The number of units of the calculated temporal convolutional unit, the number of layers of multiple temporal convolutional units, the number of units of the gated recurrent unit, the number of layers of the gated recurrent unit, and the number of convolutional filters are each rounded to an integer;
[0123] (VII) Using the trained model to predict the outlet moisture content of tobacco primary processing
[0124] Real-time collect the process flow rate of cut tobacco, the outlet temperature, the opening degree of the winnover steam valve, the opening degree of the anti-caking valve, the opening degree of the purification gas valve, the opening degree of the cleaning water valve, the negative pressure, the oxygen content, the cumulative amount of cut tobacco, the working steam flow rate of winnover, and the working steam flow rate of anti-caking at at least 10 consecutive time points on the production line of tobacco leaf conditioning. Repeat step (2) of step (II) to preprocess the above real-time collected data; input the above data into the trained model to predict the outlet moisture content of tobacco primary processing at the next moment of the cut tobacco drying process.
[0125] In step (VII), the data type and quantity input into the trained model are the same as those input into the original model in step (V).
[0126] Transformer architecture: Construct a self-attention mechanism model based on Transformer to capture long-term dependencies and global information in the cut tobacco production process. For technical details, see the paper (Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Journal of Machine Learning Research, 18(1), 1 - 54.).
[0127] TCN (Temporal Convolutional Network): Introduce the TCN module to capture local temporal patterns and multi-scale information through convolutional operations. For technical details, please refer to the paper (Pantiskas L, Verstoep K, Bal H. Interpretable Multivariate Time Series Forecasting with Temporal Attention Convolutional Neural Networks [A]. 2020 IEEE Symposium Series on Computational Intelligence (SSCI) [C]. 2020.)
[0128] GRU (Gated Recurrent Unit): Combine the GRU recurrent neural network and utilize its gating mechanism to process long sequence data and prevent the problem of gradient disappearance. For technical details, please refer to the paper ([Yin Yanchao, Hong Zhimin, Gu Wenjuan, etc. A Multi-process Process Quality Prediction Method Combining Multi-channel CNN-BiGRU and Temporal Pattern Attention Mechanism [J]. Computer Integrated Manufacturing Systems, 2023: 1–19.).
[0129] When constructing the model, the processed time series data is defined as X and input into the TDC layer of the model, which converts the information contained in the data into a set of latent vectors. Denote this set of latent vectors as X1. Then, X1 is copied into two copies, one is input into the GRU layer and the other is input into the TCN layer. After the time feature extraction of the two copied X1 through the GRU and TCN layers, the outputs Y1 and Y2 are generated, indicating that they contain global and local time features.
[0130] Subsequently, the outputs Y1 and Y2 and the output X are combined into a set of fused features through vector concatenation. The specific vector fusion operation is implemented through the concatenate method in the open-source code package numpy. Finally, the fused feature vectors of Y1, Y2, and X will be handed over to the Transformer layer for calculation.
[0131] Finally, the high-dimensional features contained in the original input are transformed into a single output value.
[0132] Through model construction, the method for predicting the moisture content at the outlet of cut tobacco based on Transformer calculates the input time series data into a single predicted value. However, the initial setting values of the model cannot achieve accurate prediction of the target outlet moisture content. Therefore, it is necessary to train the model with the help of an optimizer and a loss function.
[0133] The MAE of the training self-regulating multi-module tobacco primary processing production outlet moisture content prediction method is 0.0163, the MSE is 0.0017, and the R 2 is 0.9812. It shows significantly better performance than other traditional machine learning and deep learning models in terms of the MAE and MSE metrics.
[0134] Moreover, this model has been successfully applied to the real-time data analysis of the primary processing production line, capable of processing and predicting the changes in key parameters during the production process in real time, providing accurate prediction results, and assisting in the optimization of the production process and decision-making support.
[0135] The above is only one implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any transformation that is easily conceivable by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for predicting the moisture content of tobacco shreds produced at the outlet of a training self-adjusting multi-module, wherein a plurality of sensors are arranged on the tobacco shreds drying production line, and the sensors are used to monitor the moisture content of the discharge, the process flow of shredded tobacco, the discharge temperature, the opening of the winnover steam valve, the opening of the anti-agglomeration valve, the opening of the purge gas valve, the opening of the cleaning water valve, the negative pressure, the oxygen content, the accumulated amount of shredded tobacco, the flow of the winnover working steam, and the flow of the anti-agglomeration working steam, and the method is characterized in that: The method comprises the following steps:
1. Data collection and storage The above multiple sensors are used to simultaneously collect the moisture content of the discharge material, the process flow rate of tobacco shreds, the discharge temperature, the opening of the Winnover steam valve, the opening of the anti-agglomeration valve, the opening of the purge gas valve, the opening of the cleaning water valve, the negative pressure, the oxygen content, the accumulated amount of tobacco shreds, the Winnover working steam flow rate and the anti-agglomeration working steam flow rate, and the above twelve data at each moment are combined into one piece of raw data and stored, and at least 400,000 pieces of raw data collected continuously for one year are used as data for training the model; (II) Data preprocessing The above raw data are preprocessed as follows: (1) Remove outliers from the original data Calculate the average value of each parameter in each piece of original data among the at least 400,000 pieces. When each parameter is within the range of ±3 times the standard deviation of its average value, it is considered a normal value and retained. Otherwise, it is considered an abnormal value and the piece of original data including the abnormal value data is deleted. (2) Standardize the parameters processed in step (1) When the above parameters are within their range, they are processed according to the following formula to normalize the above parameters to between 0 and ±1. Where: Z is the standardized value (Z-Score); X is the original parameter; μ is the mean of each parameter within its range; σ is the standard deviation of each parameter within its range; After the above raw data is preprocessed, multiple preprocessed data are formed in chronological order; (III) Divide the above preprocessed data into time segments The plurality of pre-processed data are divided by using a sliding window. According to the time sequence of the pre-processed data, the size of the window is set to at least 10. The window is moved one at a time, which means that at least 10 consecutive pre-processed data are extracted each time as input features, and the input features do not include the moisture content of the discharge material. The moisture content of the discharge material after at least 10 is used as a label value, that is, the predicted target value. The input features and the label values are used as a group of associated data. Multiple groups of associated data are obtained according to the method, and 80% of the plurality of associated data are used as a training set, and 20% of the plurality of associated data are used as a test set. (IV) Construct multiple original models (3) Model composition The model is composed of multiple temporal convolutional network units, multiple gated recurrent units and a transformer layer. Multiple temporal convolutional units are connected in series, and multiple gated recurrent units are connected in series. Input features are simultaneously input into the model from the starting ends of multiple temporal convolutional units and multiple gated recurrent units. Multiple temporal convolutional units and multiple gated recurrent units process the input features and simultaneously input them into the transformer layer. At the same time, the input features are connected to the transformer layer through residual connections. Finally, the transformer layer outputs the predicted moisture content of the output material. (4) Construct multiple original models under constraints Set six initial parameters of the above model, which are respectively the number of units of multiple temporal convolutional units, and its value range is an integer from 32 to 256; the number of layers of multiple temporal convolutional units, and its value range is an integer from 3 to 8; the number of units of the gated recurrent unit, and its value range is an integer from 32 to 256; the number of layers of the gated recurrent unit, and its value range is an integer from 3 to 8; the number of convolutional filters, and its value range is an integer from 12 to 32; the learning rate, and its value range is 0.0001 - 0.1; a group is composed of the above six set parameters to obtain an original model; at least set ten groups of the above six parameters to obtain at least ten original models; Set the initial number of iterations t, t = 0, and the test score of the best model is 0; (V). Input the training set into at least ten original models composed above respectively, and train them In the way of step (III), input the training set into the above one original model respectively. During the training process, input the input features of the training set into an original model in batches, perform forward propagation to calculate the model output, then update the model parameters through backpropagation, and at the same time record the loss value of the training set. After traversing all the training set data, in order to optimize the training process, dynamically adjust the learning rate according to the change of the loss value. If the loss value of this round is lower than the lowest value of the loss values of the previous rounds, save the current model as the candidate best model at this time. Repeat the above process at least 50 rounds to fully train the model. During the whole process, the model parameters are gradually optimized, and the dynamic adjustment mechanism of the learning rate ensures more fine-grained adjustment of the weight parameters in the later stage of training, thereby improving the model performance to obtain a trained model; Repeat the above steps, input the training set into other original models respectively to obtain at least ten trained models; Input the test set into at least 10 trained models above respectively to obtain their prediction results, calculate the reciprocal of the mean squared error (MSE) value of the test set of each above training model as the score of each training set respectively, and find the training model with the highest score as the best model; (VI). Change the above six parameters through the following steps, which are respectively the number of units of the temporal convolutional unit, the number of layers of the temporal convolutional unit, the number of units of the gated recurrent unit, the number of layers of the gated recurrent unit, the number of convolutional filters, and the learning rate (5). Update the best model Compare the test score of the best model of this iteration with the test score of the previous best model, select the model with the highest test score as the best model, and record or update the best model; (6). Calculate the adaptive coefficient E In each t iteration, calculate an adaptive coefficient E according to the current number of iterations t, and its value is This adaptive coefficient E gradually increases with the increase of the number of iterations, where t is the current number of iterations and T is the total number of iterations; (7). Soft RIME search In the i-th model among the above at least ten original models, a random number r is randomly generated 2, , and its value range is (-1, 1). If r2 < E, update each parameter in this original model using the following formula Soft RIME; in are the six updated parameters, i and j represent the jth parameter in the i-th model, R best,j is the jth parameter of the above optimal model, Ub ij is the maximum value of the jth parameter value range in the i-th model, Lb ij is the minimum value of the jth parameter in the i-th model, r1 is a random number in the range (-1, 1), and cosθ is a coefficient with a value of Where t is the current iteration number, T is the total iteration number, r1 and cosθ jointly determine the update direction of the parameters, and h is a random number in the range of (0, 1) to ensure that the generated parameters are randomly distributed within the constraint range; β is a coefficient, Round to the nearest integer. w is a manually set parameter with a value of 5. For the original models that do not meet the condition of r2 < E, keep the parameters of the original models unchanged; Repeat the above operations to perform Soft RIME update on at least ten trained models above; obtain at least ten new original models; (8). Hard RIME penetration mechanism After Soft RIME is updated, a random number r3 is randomly generated for the i-th model of the at least ten new original models, and its value range is (-1, 1). If r3 <F normr( S i ), use the following formula Hard RIME to update the parameters in the new original model; in are the updated six parameters, R best,j is the jth parameter of the above optimal model; Among them: F normr (S i ) represents the normalized value of the current model test score, indicating the probability of the i-th model being selected, F(S i ) is the test score of the original model in this round of iteration; F(S) min It is the minimum test score of at least ten groups of original models in this round of iterations; F(S) max It is the maximum test score of at least ten groups of original models in this round of iterations; For those who do not meet r3 <F normr (S i ) condition, keeping the parameters of the original model obtained in step (8) unchanged; Repeat the above operation to perform Hard RIME update on the at least ten new original models to obtain the at least ten updated original models; Let t = t + 1, When t<=T, repeat step (5), replace the at least ten original models or at least ten updated original models of the previous iteration with the at least ten updated original models obtained most recently, and retrain the at least ten groups of updated original models obtained most recently with the training set and the test set; When t>T, the best model is the trained model. (VII) Use the trained model to predict the moisture content of tobacco silk production export Collect in real time the tobacco process flow, discharge temperature, Winnover steam valve opening, anti-agglomeration valve opening, purge gas valve opening, cleaning water valve opening, negative pressure, oxygen content, leaf accumulation, Winnover working steam flow and anti-agglomeration working steam flow on the tobacco leaf rehumidification production line at least 10 consecutive time points, repeat step (ii) (2) to pre-process the above real-time collected data; input the above data into the trained model to predict the tobacco leaf production outlet moisture content at the next moment of the drying process.
2. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 1, characterized in that: The standard deviation Where: N is the total number of data points; x i is the data of the ith data point; μ is the mean of the population, and the calculation formula is:
3. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 2, characterized in that: The size of the window is set to 15, which means that 15 consecutive preprocessed data are extracted each time as input features, and the output moisture content of the 16th preprocessed data is used as the label value; each time a data is moved along the time arrangement order of the preprocessed data, then the 15 consecutive preprocessed data in the window are used as the next input feature, and the output moisture content of the 17th preprocessed data is used as the next label value, and so on.
4. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 3, characterized in that: The data type and quantity input into the trained model in step (seven) are the same as the data type and quantity input into the original model in step (five).
5. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 4, characterized in that: In step (six), the number of original models is 30 and the number of iterations T is 5.
6. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 5, characterized in that: In step (six), the calculated number of units of the temporal convolutional unit, the number of layers of the plurality of temporal convolutional units, the number of units of the gated recurrent unit, the number of layers of the gated recurrent unit, and the number of convolutional filters are rounded to integers.
7. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 6, characterized in that: The mean square error Where: n is the total number of samples; y i is the true value of the i-th sample; is the predicted value of the i-th sample.
8. The training self-adjusting multi-module tobacco shred production outlet moisture content prediction method according to claim 7, characterized in that: The sampling frequency of the original data is 6 seconds.
Citation Information
Cited By
Multi-quality index output prediction method and system for shred-making loosening and moisture-regaining process
CN115456460A