Method and system for predicting tobacco shred processing strength in drum tobacco shred drying process and storage medium

By introducing the heat conduction equation and multi-head self-attention mechanism into the drum wire drying process, and combining it with the mechanistic-constrained LSTM model, the problem of inaccurate prediction of processing intensity during the drum wire drying process is solved, and high-precision and stable prediction results are achieved.

CN121774246APending Publication Date: 2026-04-03CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512004414.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies for predicting processing intensity during the drum drying process suffer from poor interpretability, high sample dependence, and insufficient stability under dynamic nonlinear conditions, leading to inaccurate predictions.

Method used

A mechanism-based data fusion approach is adopted. By constructing a fusion mechanism knowledge model and a mechanism-constrained prediction model, and combining the heat conduction equation and multi-head self-attention mechanism, the MHA-TimeGAN-PiLSTM framework is built to predict the processing intensity of tobacco shreds.

Benefits of technology

It significantly improves the accuracy and stability of processing strength prediction, and can more accurately capture the dynamic impact of temperature changes on processing strength, thereby improving the model's prediction accuracy and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121774246A_ABST
    Figure CN121774246A_ABST
Patent Text Reader

Abstract

The invention relates to the field of tobacco processing, in particular to a mechanism data fused tobacco shred processing intensity prediction method and system in a few-sample roller tobacco shred drying process and a storage medium, and the method comprises the steps: obtaining data of a roller tobacco shred dryer in a normal operation state, and carrying out the preprocessing, so as to obtain a data set; constructing a fusion mechanism knowledge model and training to obtain an enhanced data set; constructing a mechanism constraint prediction model, and training the mechanism constraint prediction model by adopting the enhanced data set; performing joint training on the fusion mechanism knowledge model and the mechanism constraint prediction model to obtain a final cut tobacco processing strength prediction model; and performing processing strength prediction on the real-time process variable data by adopting the tobacco shred processing strength prediction model. According to the method, dynamic heat conduction mechanism knowledge and TimeGAN are deeply fused to generate high-quality training data with physical consistency, and the problems of insufficient data and inaccurate prediction in the drum cut tobacco drying process are effectively solved in combination with a mechanism constraint LSTM prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tobacco processing, and more specifically to a method, system, and storage medium for predicting the processing strength of tobacco shreds in a few-sample drum drying process using mechanistic data fusion. Background Technology

[0002] Rotary drum dryers are key pieces of equipment commonly used in tobacco processing. Their main function is to dry tobacco using heating and drum rotation to achieve uniform control of moisture content and improve flexibility. In cigarette production, drying is a crucial step in the initial processing of tobacco, and its heat treatment process directly affects the physical structure and quality stability of the tobacco. In existing technologies, processing intensity is typically used to characterize the combined effect of heat treatment intensity and duration experienced by the tobacco during drying. This parameter directly affects the moisture content, bulkiness, and structure of the tobacco, and also significantly influences changes in chemical composition, smoke release characteristics, and final sensory quality. Therefore, processing intensity is a key indicator in evaluating drying effectiveness. However, in actual production, the collection of processing intensity data has certain limitations, especially during operational transitions. Affected by changes in key process variables such as tobacco surface temperature and humidity, hot air temperature, and drum wall temperature, processing intensity and its related parameters often exhibit strong nonlinear characteristics and significant dynamic fluctuations. This makes existing methods insufficient for predicting processing intensity. In recent years, data-driven methods such as Multi-Layer Perceptron (MLP), Recurrent Neural Network (RNN), and Transformer have been used for processing intensity prediction tasks. However, due to the complex heat and mass transfer processes and dynamic mechanisms such as gas-solid coupling involved in the drum drying process, traditional data-driven models often suffer from insufficient interpretability and are heavily dependent on large-scale sample data. When the sample size is insufficient or when there are non-stationary, large dynamic fluctuations, and strong nonlinear dynamic characteristics, the prediction is prone to instability or bias.

[0003] Therefore, there is an urgent need for a hybrid prediction method that integrates dynamic mechanism knowledge and data augmentation strategies to solve the problems of weak interpretability, high sample dependence, and insufficient stability under dynamic nonlinear conditions in existing technologies. This method can achieve accurate, reliable, and interpretable prediction of key process parameters in the tobacco drying process, providing effective support for tobacco quality control and process optimization. Summary of the Invention

[0004] The purpose of this invention is to provide a method, system, and storage medium for predicting the processing strength of tobacco shreds in a few-sample drum drying process by fusing mechanistic data, in order to solve the technical problems of weak interpretability, high sample dependence, and insufficient stability under dynamic nonlinear conditions in the prior art.

[0005] To achieve the above objectives, embodiments of the present invention provide a method for predicting the processing strength of tobacco shreds during the drum drying process, comprising: Data from the normal operation of the rotary drum dryer is acquired and preprocessed to obtain a dataset. Construct and train a fusion mechanism knowledge model to obtain an augmented dataset; A mechanism constraint prediction model is constructed, and the mechanism constraint prediction model is trained using the augmented dataset; The fusion mechanism knowledge model and the mechanism constraint prediction model are jointly trained to obtain the final tobacco processing strength prediction model; The processing intensity prediction model for tobacco shreds is used to predict the processing intensity based on real-time process variable data.

[0006] Optionally, a fusion mechanism knowledge model is constructed and trained to obtain an augmented dataset, including: A knowledge model based on the fusion mechanism is constructed based on TimeGAN, and a multi-head self-attention mechanism is introduced for each module of the TimeGAN model. A heat conduction equation is introduced into the embedding-recovery network in the fusion mechanism knowledge model; The dataset is input into the fusion mechanism knowledge model to predict the outlet tobacco temperature; A reconstruction loss function is constructed based on the predicted export tobacco temperature; Based on the reconstruction loss function, unsupervised loss and supervised loss are introduced to obtain the optimization objective function; The fusion mechanism knowledge model is trained based on the optimization objective function; The dataset is input into the trained fusion mechanism knowledge model to obtain an augmented dataset.

[0007] Optionally, introducing a heat conduction equation into the embedding-recovery network in the fusion mechanism knowledge model includes: The conduction equation is defined according to formulas (1) and (2). (1) (2) in, For position and time temperature, The rate of change of the temperature of the exported tobacco shreds with respect to time. For the outlet tobacco temperature at Second derivative in the direction, For thermal diffusivity, For the heat transfer rate of tobacco, The density of the tobacco shreds, This refers to the specific heat capacity of tobacco shreds.

[0008] Optionally, the reconstruction loss function constructed based on the predicted exit tobacco temperature includes: The reconstruction loss function is obtained according to formulas (3) to (5). (3) (4) (5) in, To reconstruct the loss function, The reconstruction loss function without embedded mechanistic information. For the mechanism loss function, The temperature data loss function is... To predict the temperature field of the exported tobacco, This represents the actual temperature field of the exported tobacco shreds. To predict the time derivative of the temperature field of the exported tobacco shreds, To predict the spatial second derivative of the temperature field of the exported tobacco shreds, For the diffusion rate, This represents the number of training samples.

[0009] Optionally, based on the reconstruction loss function, unsupervised loss and supervised loss are introduced to obtain the optimization objective function, including: The objective function is obtained from formulas (6) to (7). (6) (7) in, To reconstruct the loss function, For the supervision loss function, For unsupervised loss functions, For parameters embedded in the network, To restore the network parameters, For the parameters of the generator network, For the parameters of the discriminator network, and This is a hyperparameter.

[0010] Optionally, constructing a mechanism-constrained prediction model and training the model using the augmented dataset includes: Based on expert prior knowledge, obtain the mechanism constraints; Based on the aforementioned mechanistic constraints, a monotonic constraint relationship loss function is constructed; A gradual constraint method is introduced, and a total loss function for the mechanistic constraint prediction model is constructed.

[0011] Optionally, based on the aforementioned mechanistic constraints, constructing a monotonic constraint relationship loss function includes: Based on formulas (8) and (9), a monotonic constraint relationship between processing intensity and hot air temperature and cylinder wall temperature is constructed. (8) (9) in, The processing strength versus hot air temperature loss function, The function is: processing strength - cylinder wall temperature loss. for Real-time hot air temperature for Real-time hot air temperature for The actual value of the cylinder wall temperature at any given time. for The actual value of the cylinder wall temperature at any given time. for Predicted processing strength at any time for Real value of processing strength at all times This represents the total number of training samples.

[0012] Optionally, a gradual constraint method is introduced, and the total loss function of the mechanistic constraint prediction model is constructed as follows: (10) (11) (12) in, The total loss function of the mechanism-constrained prediction model. The processing strength versus hot air temperature loss function, The function is: processing strength - cylinder wall temperature loss. For data-driven loss terms, For the gradient constraint function, for Real value of processing strength at all times for Predicted processing strength at any time for Real value of processing strength at all times This is the upper bound of the confidence interval. This is the lower bound of the confidence interval. The total number of training samples, for The weighting coefficients, for The weighting coefficients, for The weighting coefficients.

[0013] On the other hand, the present invention also provides a tobacco processing strength prediction system for a drum drying process, the system including a processor configured to perform any of the methods described above.

[0014] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.

[0015] The beneficial effects of this invention are as follows: This invention proposes a hybrid prediction model for processing intensity that combines data augmentation and mechanistic constraints. The model innovatively embeds the heat conduction equation, which describes the dynamic evolution of temperature during the drum drying process, into a TimeGAN generative adversarial network. This ensures that the generated data not only possesses statistical characteristics but also conforms to the physical changes in the temperature field during the drying process, significantly improving the realism of the generated samples and the accuracy of the prediction model. Addressing the strong nonlinearity and non-stationarity of processing intensity data during the drying process, the introduction of the heat conduction equation helps the model more accurately capture the dynamic impact of temperature changes on processing intensity. Furthermore, the model's ability to model temporal features is enhanced by incorporating a multi-head self-attention mechanism. Simultaneously, based on data augmentation, key mechanistic knowledge of the drum drying process is integrated to construct an LSTM prediction model with mechanistic constraints, achieving high-precision prediction of processing intensity. Experimental verification and model parameter optimization ensure the stability and accuracy of the prediction results. This invention significantly improves the prediction accuracy and interpretability of the model by combining data augmentation with mechanistic constraints. Compared with traditional methods that integrate static mechanistic knowledge, the dynamic mechanistic knowledge such as the heat conduction equation introduced in this invention can more realistically depict the spatiotemporal evolution characteristics of the processing process, and significantly improve the model's ability to model complex dynamic behaviors and its generalization performance.

[0016] Other features and advantages of the embodiments of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart of a method for predicting the processing strength of tobacco shreds during a drum drying process according to an embodiment of the present invention; Figure 2 This is a framework diagram of a tobacco processing strength prediction model according to one embodiment of the present invention; Figure 3 A flowchart of a method for constructing and training a fusion mechanism knowledge model according to an embodiment of the present invention; Figure 4 A flowchart of a method for constructing and training a mechanism-constrained prediction model according to an embodiment of the present invention; Figure 5 This is a schematic diagram illustrating the enhancement of processing intensity data using the proposed fusion mechanism information generation model according to an embodiment of the present invention; Figure 6 This is a schematic diagram comparing the prediction results of the tobacco processing strength of one embodiment of the present invention with those of other models. Detailed Implementation

[0018] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.

[0019] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.

[0020] like Figure 1 The diagram shows a flowchart of a method for predicting the processing strength of tobacco shreds during a drum drying process according to an embodiment of the present invention. Figure 1 The prediction method includes the following steps: In step S10, data under normal operating conditions of the roller dryer is acquired and preprocessed to obtain a dataset; In step S11, a fusion mechanism knowledge model is constructed and trained to obtain an augmented dataset; In step S12, a mechanism constraint prediction model is constructed, and the mechanism constraint prediction model is trained using an augmented dataset; In step S13, the fusion mechanism knowledge model and the mechanism constraint prediction model are jointly trained to obtain the final tobacco processing strength prediction model. In step S14, the processing intensity prediction model for tobacco processing is used to predict the processing intensity of real-time process variable data.

[0021] In such Figure 1 In the method for predicting the processing strength of tobacco shreds during the drum drying process, step S10 is used to acquire and preprocess data under normal operating conditions of the drum dryer. In this embodiment, acquiring data under normal operating conditions of the drum dryer can be achieved by collecting normal operating data of the drum dryer as the original dataset, which includes an original training set and an original test set. The original training set is denoted as... ,in, This indicates the number of samples in the original training set. The data dimension of each sample in the dataset is V, meaning that each sample corresponds to V process variables. There is Each sample corresponds to a label under normal operating conditions of a drum dryer. ,in, The data dimension of each label in the data is That is, each tag corresponds to There are 1 target variables. This portion of the samples and labels constitutes the original training sample set, denoted as: .

[0022] The original test set is denoted as ,in, This indicates the number of samples in the original test set. The data dimension of each sample in the dataset is V, meaning that each sample corresponds to V process variables. There is Each sample corresponds to a label under normal operating conditions of a drum dryer. ,in, The data dimension of each label is That is, each tag corresponds to There are 1 target variables. This portion of the samples and the labels constitute the original test sample set, denoted as . In this example, the initial training set size can be set. The original test set sample size is 720. The value is 180; the data dimension V for each sample is set to 20, meaning 20 process variables and the data dimension for each label is... Set to 1, meaning the target variable is processing intensity.

[0023] In this implementation, the method for preprocessing the original dataset can be to first normalize the original dataset, and then use a sliding window to divide the normalized original dataset to obtain the preprocessed dataset. Specifically, in this example, it can be to divide the dataset into parts separately. and Normalization is performed to normalize the data to the [0,1] interval. For example, formula (13) can be used to normalize the process variables: , =1,2,…, ; =1,2,…,20,(13) in, This represents the value of the v-th process variable in the i-th training sample of the original training set. , Let $v$ and $v$ represent the maximum and minimum values ​​of the $v$-th process variable in all the original training samples, respectively. This represents the normalized value of the corresponding sample and variable in the original training set. After normalizing the process variables in the original training set, the same normalization operation is then performed on the target variable, and its calculation formula is shown in formula (14): , =1,2,…, ; =1, (14) in, Represents the i-th training sample in the original training set. The values ​​of the target variables, , They represent the nth and nth samples in all the original training samples, respectively. The maximum and minimum values ​​of the target variables. This represents the normalized values ​​of the corresponding labels and variables in the original training set.

[0024] Next to After normalization, the calculation formula is formula (15): , =1,2,…, ; =1,2,…,20,(15) in, This represents the value of the v-th process variable in the i-th test sample of the original test set. This represents the normalized value of the corresponding sample and variable in the original test set. Similarly, after normalizing the process variables in the original test set, the same normalization operation is then performed on the target variable, and its calculation formula is shown in formula (16): , =1,2,…, ; =1, (16) in, Represents the i-th test sample in the original test set. The values ​​of the target variables, This represents the normalized values ​​of the corresponding labels and variables in the original test set.

[0025] After normalizing the dataset, we obtain the normalized original training sample set: , in, Indicates that there is The original training sample set consists of the labels corresponding to each sample. Similarly, the normalized original test sample set can be obtained: , in, Indicates that there is The original test sample set consists of labels corresponding to each sample. After normalization, the total number of samples in the original dataset is... In this example, This represents the total number of samples in the original dataset, set to 900.

[0026] Next, a sliding window is used to divide the normalized original dataset. Specifically, in this example, time-series data needs to be extracted by truncating the sliding window to obtain the following variable data pairs: (17) in, Let the size be represented by the window size, and let the training set be denoted as . : , in, Indicates that there is There are *n* samples, each containing historical information of size *l* windows, and each historical information has *l* dimensions. Let the labels corresponding to the training set be... , Indicates that there is There are 10 labels, and the data dimension of each label is 1. Let the test set be... : , in, Indicates that there is There are *n* samples, each containing historical information of size *l* windows, and each historical information has *l* dimensions. Let the labels corresponding to the test set be... , Indicates that there is There are 10 labels, and the data dimension of each label is 1. In this example, This indicates the window size; set it to 10.

[0027] By partitioning the dataset using variable data pairs, we can obtain the training sample set and the test sample set, respectively. , ;in, Zhongyou The training sample set consists of data pairs of several variables. Zhongyou The test sample set consists of pairs of variable data. In this example, The training sample set consists of 712 variable data pairs. The test sample set consists of 178 variable data pairs.

[0028] In this embodiment, the tobacco processing intensity prediction model includes a fusion of a mechanistic knowledge model and a mechanistic constraint prediction model. Specifically, in this example, it could be a processing intensity prediction model constructed using the MHA-TimeGAN-PiLSTM framework, such as... Figure 2 As shown, the input to the model is ,in, Depend on The sample size consists of, and each sample contains There are process variables, and the model output is: The target variable for prediction, namely processing intensity, is modeled using Time-series Generative Adversarial Network (TimeGAN) and Long Short-Term Memory (LSTM) as the data-driven model of this invention. As a parameter of the entire data-driven model. In this example, The sample size is 890. and The values ​​are set to 20 and 1 respectively. This means that the input of the LSTM model consists of 20 process variables and 1 corresponding target variable from the training sample set, and the output of the model is 1 predicted target variable, which is the predicted value of processing intensity.

[0029] Further, a fusion mechanism knowledge model is constructed and trained in step S11 to obtain an augmented dataset. In this embodiment, the specific method for constructing and training the fusion mechanism knowledge model in step S11 can be of various forms known to those skilled in the art. In one example of the present invention, step S11 may include, for example... Figure 3 The steps shown are described in this. Figure 3 In this context, step S11 may include: In step S20, a fusion mechanism knowledge model is constructed based on TimeGAN, and a multi-head self-attention mechanism is introduced for each module of the TimeGAN model. In step S21, a heat conduction equation is introduced into the embedding-recovery network in the fusion mechanism knowledge model; In step S22, the dataset is input into the fusion mechanism knowledge model and the outlet tobacco temperature is predicted; In step S23, a reconstruction loss function is constructed based on the predicted outlet tobacco temperature; In step S24, based on the reconstruction loss function, unsupervised loss and supervised loss are introduced to obtain the optimization objective function; In step S25, the fusion mechanism knowledge model is trained based on the optimization objective function; In step S26, the dataset is input into the trained fusion mechanism knowledge model to obtain an enhanced dataset.

[0030] In such Figure 3 In the method shown, step S20 introduces a multi-head self-attention mechanism into the TimeGAN model framework. Specifically, the multi-head self-attention mechanism is introduced into the Embedding Network, Generator, Recovery Network, and Discriminator modules respectively, thereby effectively modeling the global contextual information of the input features. Specifically, in this example, in each module, the input features are processed through a self-attention mechanism, where the self-attention is calculated using three learnable linear embedding matrices. , and What we received included: , and Based on this triple input, the self-attention mechanism obtains the following through dot product operation: (18) Among them, matrix , and These represent the query, key, and value, respectively. Represents the input feature matrix. This represents the transpose of the key matrix. , and This represents a trainable weight matrix. This represents the dimension of the key vector. After softmax processing, a linear combination of weighted values ​​can be calculated as the output of the self-attention mechanism.

[0031] To effectively characterize multi-level temporal dependencies and generate diverse feature representations through parallel computation, a multi-head self-attention mechanism computes multiple independent attention heads by performing different linear projections on the query, key, and value, and then concatenates the outputs of these attention heads: (19) in, This indicates a splicing operation. Indicates the output projection matrix. , h The number of self-attention heads.

[0032] By introducing a multi-head self-attention mechanism, the TimeGAN model can fully mine and utilize rich contextual information to improve the model's generative capabilities and the accuracy of time-series data modeling. In this example, the self-attention heads... h The number can be set to 3.

[0033] Step S21 is used to introduce a heat conduction equation related to the processing intensity of the roller drying process into the embedding-recovery network of TimeGAN. Specifically, the heat conduction equation can be defined according to formula (1): (1) in, For position and time temperature, The rate of change of the temperature of the exit tobacco shreds with respect to time represents the change of temperature over time. For the outlet tobacco temperature at The second derivative in the direction represents the diffusion behavior of temperature in space; Thermal diffusivity represents the rate at which heat diffuses from tobacco. Typically, thermal diffusivity is related to the thermal conductivity of the tobacco. ,density and specific heat capacity Related, that is .

[0034] When combining the multi-head self-attention mechanism with the TimeGAN network, the heat conduction equation is embedded in the network model as part of the mechanistic information, resulting in a TimeGAN generative model that integrates mechanistic information. In this example, It can be learned by the model itself, without the need to set the value manually.

[0035] Step S22 is used to process the dataset The input is fed into a TimeGAN network model that incorporates Physics-Informed Neural Networks (PINN), and the embedding-recovery network module predicts and outputs the exit tobacco temperature. The predicted exit tobacco temperature field is then represented as follows: The actual temperature field of the exported tobacco shreds is expressed as: .

[0036] Step S23 is used to construct the reconstruction loss function. In this example, to ensure that the temperature field output by the neural network conforms to the heat conduction equation, automatic differentiation is used to calculate its time derivative. Second derivative of space This is then used to construct the mechanistic loss function. By comparing the temperature field predicted by the network with the physical laws in the heat conduction equation, the following PINN mechanistic loss function is constructed: (4) in, This represents the number of training samples. The loss function is calculated by averaging the residuals of each sample point over all sample points. In this way, the mechanistic loss and the error of the actual temperature field work together to help the neural network better follow physical laws during the learning process, thereby obtaining more accurate temperature field predictions.

[0037] Next, the data loss caused by the difference between the actual temperature data and the temperature value predicted by the restorer module is calculated. This loss is defined as follows: (5) in, This indicates the predicted temperature of the tobacco shreds at the outlet of the restorer module. This represents the actual observed temperature of the exported tobacco shreds. This loss is used to measure the difference between the neural network's predictions and the actual data.

[0038] Then, the heat conduction equation described above is embedded into the predefined TimeGAN network in the form of PINN, thereby constructing a new reconstruction loss function in the embedding-recovery network. This reconstruction loss function is expressed as follows: (3) in, This represents the reconstruction loss function of the embedding-recovery network that embeds the heat conduction equation. This represents the reconstruction loss function without embedded mechanistic information. and These are hyperparameters used to control the mechanistic loss. The weights in the equation.

[0039] To guide the generation of samples with similar statistical properties and ensure the temporal relevance of the generated data, unsupervised loss is typically introduced. and monitoring losses Therefore, unsupervised loss and supervised loss are introduced in step S24 to obtain the optimized objective function.

[0040] The optimization objective of the TimeGAN network, which integrates the entire mechanism information, needs to consider simultaneously. , and The loss functions for these three factors, specifically the objective function for optimization, are expressed as follows: (6) (7) in, For parameters embedded in the network, To restore the network parameters, For the parameters of the generator network, For the parameters of the discriminator network, and This represents the hyperparameters that balance the three losses.

[0041] Finally, by jointly training and minimizing the objective function in step S25, and updating the model parameters using backpropagation, until the preset number of training rounds is reached, the optimal parameters of the TimeGAN model incorporating the underlying mechanism knowledge are obtained. In this example, the number of training epochs is set to 5000. and Set them to 0.1 and 10 respectively. and The hyperparameters are set to 1 and 10, respectively. Step S26 is used to input the dataset into the trained fusion mechanism knowledge model to obtain the augmented dataset.

[0042] Step S12 is used to construct a mechanism-constrained LSTM prediction model. Based on expert prior knowledge, the mechanism constraints are embedded into the data-driven model for model training. In this embodiment, the specific method for constructing and training the mechanism-constrained prediction model in step S12 can be of various forms known to those skilled in the art. In one example of the present invention, step S12 may include, for example... Figure 4 The steps shown are described in this. Figure 4 In this context, step S12 may include: In step S30, the mechanism constraints are obtained based on expert prior knowledge; In step S31, a monotonic constraint relationship loss function is constructed based on the mechanism constraints. In step S32, a gradual constraint method is introduced, and the total loss function of the mechanistic constraint prediction model is constructed.

[0043] In such Figure 4 In the method shown, step S30 is used to summarize and analyze the mechanistic knowledge required during model training based on expert prior knowledge. In this example, the relevant mechanistic knowledge is as follows: Based on the monotonic constraint relationship between the processing strength of tobacco shreds and the hot air temperature and cylinder wall temperature, they are expressed as follows: , ,in, Indicates the processing strength of tobacco shreds , Indicates hot air temperature , Indicates cylinder wall temperature As can be seen from the above constraints, the processing strength of tobacco shreds gradually increases with the increase of hot air temperature and cylinder wall temperature, and all three show a clear monotonically increasing relationship.

[0044] After clarifying the mechanistic constraints required for model training, step S31 embeds these constraints as penalty terms into the predefined LSTM model loss function for training, ensuring that the model training process also conforms to certain physical laws. Based on the monotonicity constraint relationship between tobacco processing intensity and hot air temperature and cylinder wall temperature, it can be expressed as: , ; in, express Predicted processing strength at any time express Real value of processing intensity at any given time. and They represent Time and Real-time hot air temperature and They represent Time and The actual value of the cylinder wall temperature at any given time. Based on the monotonic constraint relationship between processing intensity and hot air temperature and cylinder wall temperature, the loss function is defined as follows: (8) (9) in, and This represents a loss term that penalizes data where the relationship between processing intensity and hot air temperature and cylinder wall temperature does not exhibit a monotonicity. Represents the linearly modified activation function. for Real-time hot air temperature for Real-time hot air temperature for The actual value of the cylinder wall temperature at any given time. for The actual value of the cylinder wall temperature at any given time. for Predicted processing strength at any time for Real value of processing strength at all times The total number of training samples, In this example, the total number of training samples It can be set to 1602.

[0045] In step S32, a gradual constraint method is introduced, and the total loss function of the mechanistic constraint prediction model is constructed. After embedding mechanistic constraints into the LSTM model, a gradual constraint method is further introduced to effectively solve the dynamic fluctuation problem in the strength data of the roller drying process. Specifically, in this example, the first-order difference sequence is first calculated based on the original processing strength sequence. This is used to characterize short-term variation features; subsequently, the predicted processing intensity at time i is differencing the actual processing intensity at time i-1; then, at a given confidence level... Next, calculate the difference value. The confidence interval for the overall mean is defined, and predicted values ​​with differences exceeding the interval are considered outliers. To ensure that the predicted variability is within a reasonable range for changes in processing intensity, a parameter representing the confidence level of the true data distribution mean is manually adjusted. The upper and lower bounds of the corresponding confidence intervals are respectively and Finally, the loss function takes the form shown below: (12) loss function This effectively reduces outliers in predicted values, improves the model's robustness to non-stationary features of the process, and enhances overall prediction performance. During model training, the data-driven loss term is denoted as... Its specific form is as follows: (11) comprehensive , and The loss function, the total loss function of the constructed mechanism-constrained prediction model, is expressed as: (10) in, This represents the total loss function of the mechanistic-constrained LSTM prediction model. , and This represents the weight coefficient of each penalty term, used to balance the contributions of different mechanistic constraints during training. In this example, the confidence parameter... Set to 99%; the corresponding upper and lower bounds of the confidence interval are respectively and They were set to 0.1 and -0.1 respectively; , , Set them to 0.1, 0.1, and 0.01 respectively.

[0046] Step S13 is used to jointly train the fusion mechanism knowledge model and the mechanism constraint prediction model to obtain the final tobacco processing strength prediction model. Before training the LSTM model, the TimeGAN model, which integrates mechanism information, is used to augment the data, generating sample pairs of the same size as the original data. Subsequently, the generated samples are normalized and partitioned using a sliding window method according to step S10 to obtain the training set. The training set obtained in step S10 With enhanced samples The data are concatenated to form a new training set of process variables, denoted as: , ,in, This indicates a splicing operation. This represents the concatenated training set of process variables. This represents the test set.

[0047] Prepare training data pairs by pairing the data-augmented process variables with their corresponding target variables (labels) and inputting them into the mechanistic-constrained LSTM prediction model. After initializing the model parameters, use the combined mechanistic-constraint loss function. Calculate the current loss and update the model parameters through backpropagation. Repeat the gradient calculation and parameter update process until the preset number of training epochs is reached, finally obtaining the optimal model parameters after training. In this example, the number of training rounds is set to 100, the LSTM hidden layer size is set to 2, and the number of neurons per layer is set to 64.

[0048] Step S14 is used to perform online prediction using the trained model, predicting processing intensity in real time. In this example, a processing intensity prediction model composed of TimeGAN and LSTM is jointly trained, and the network parameters are optimized to obtain the optimal model parameters. This leads to a processing intensity prediction model that integrates mechanistic information. Real-time process variable data are collected from actual industrial sites. This information is then input into a processing intensity prediction model that integrates mechanistic information, and the model uses optimal model parameters. Finally, the model outputs a predicted value for processing intensity. .

[0049] The MHA-TimeGAN-PiLSTM processing intensity prediction model trained by this invention was tested on a test set containing 178 sample data points. This invention introduces a multi-head attention mechanism into the TimeGAN network and embeds the heat conduction equation, while embedding monotonicity constraints and gradual variation constraints into the LSTM model, forming a processing intensity prediction model that integrates mechanistic information. Figure 5 This diagram illustrates the results of data enhancement for processing intensity using the proposed Fusion Mechanism Information Generation Model (MHA-TimeGAN) in this invention. The diagram shows the visualization analysis of the processing intensity data using MHA-TimeGAN, employing Principal Component Analysis (PCA) (left) and Distributed Stochastic Neighbor Embedding (t-SNE) (right). Red represents the original processing intensity data, and blue represents the generated processing intensity data. The results show that after enhancing the processing intensity data using the MHA-TimeGAN model, the distribution of the generated samples closely matches the original data, exhibiting strong detail restoration capabilities, especially in edge and low-density regions. Through the visualization analysis of PCA and t-SNE, it can be clearly observed that the generated samples not only have a high degree of similarity to the original data in overall distribution but also maintain consistency in the structure of each feature space, further validating the effectiveness and data enhancement capabilities of the model.

[0050] To evaluate the performance of this model, various typical machine learning methods (including MLP, CNN, LSTM, PiLSTM, TimeGAN-PiLSTM, and the method proposed in this invention) were used for processing intensity prediction. The test results and performance are as follows: Figure 6 As shown in Table 1: Table 1. Results of experimental predictions of processing strength using different models

[0051] Figure 6This diagram illustrates the comparison of the processing intensity prediction results of the proposed model (MHA-TimeGAN-PiLSTM) with other models in this invention. The figure compares the performance of the MHA-TimeGAN-PiLSTM processing intensity prediction model with other models (including MLP, CNN, LSTM, PiLSTM, and TimeGAN-PiLSTM) through experiments. The black curve represents the measured processing intensity, and the orange dashed line represents the predicted processing intensity value of the proposed model. The results show that the proposed model performs best in prediction.

[0052] As shown in Table 1, the proposed method outperforms the suboptimal TimeGAN-PiLSTM model in all metrics. Specifically, the coefficient of determination (R²) is improved by 2.13%, the mean absolute percentage error (MAPE) is reduced by 7.96%, the mean absolute error (MAE) by 9.36%, and the root mean square error (RMSE) by 40.61%. It is particularly noteworthy that although the TimeGAN-PiLSTM model enhances sample diversity to some extent through generative modeling, its performance is still lower than the proposed MHA-TimeGAN-PiLSTM model. This performance improvement is mainly attributed to two aspects: First, the multi-head self-attention mechanism can more effectively capture non-stationary characteristics, dynamic fluctuations, and nonlinear dynamic segments. In these regions, the proposed model can more accurately capture feature information, and the generated samples more accurately reproduce the changing trends and fluctuation characteristics of the original data. Second, the method designed in this invention integrates mechanistic information into the model. By embedding heat conduction equations, monotonicity constraints, and gradual change constraints, it uses physical laws to guide the data-driven model training, avoiding complete reliance on data-driven models, thereby effectively improving the accuracy of processing intensity prediction. This demonstrates the feasibility and effectiveness of the method proposed in this invention.

[0053] On the other hand, the present invention also provides a tobacco processing strength prediction system for a drum drying process, the system including a processor configured to perform any of the methods described above.

[0054] In another aspect, the present invention also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement any of the methods described above.

[0055] The beneficial effects of this invention are as follows: This invention proposes a hybrid prediction model for processing intensity that combines data augmentation and mechanistic constraints. The model innovatively embeds the heat conduction equation, which describes the dynamic evolution of temperature during the drum drying process, into a TimeGAN generative adversarial network. This ensures that the generated data not only possesses statistical characteristics but also conforms to the physical changes in the temperature field during the drying process, significantly improving the realism of the generated samples and the accuracy of the prediction model. Addressing the strong nonlinearity and non-stationarity of processing intensity data during the drying process, the introduction of the heat conduction equation helps the model more accurately capture the dynamic impact of temperature changes on processing intensity. Furthermore, the model's ability to model temporal features is enhanced by incorporating a multi-head self-attention mechanism. Simultaneously, based on data augmentation, key mechanistic knowledge of the drum drying process is integrated to construct an LSTM prediction model with mechanistic constraints, achieving high-precision prediction of processing intensity. Experimental verification and model parameter optimization ensure the stability and accuracy of the prediction results. This invention significantly improves the prediction accuracy and interpretability of the model by combining data augmentation with mechanistic constraints. Compared with traditional methods that integrate static mechanistic knowledge, the dynamic mechanistic knowledge such as the heat conduction equation introduced in this invention can more realistically depict the spatiotemporal evolution characteristics of the processing process, and significantly improve the model's ability to model complex dynamic behaviors and its generalization performance.

[0056] Unlike existing technologies, this invention differs fundamentally in its modeling approach and technical path. Traditional methods often rely on static mechanistic knowledge to impose simple constraints on data-driven models, which fails to effectively reflect the dynamic evolution during processing, resulting in limited model interpretability and generalization ability. Therefore, this invention, for the first time, deeply integrates the dynamic heat conduction mechanism knowledge during drum drying with TimeGAN to generate physically consistent, high-quality training data. Combined with a mechanistic-constrained LSTM prediction model, this effectively solves the problems of insufficient data and inaccurate predictions during drum drying, demonstrating higher practical application value and technological innovation.

[0057] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0058] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0059] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0060] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0061] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0062] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0063] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0064] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0065] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for predicting the processing strength of tobacco shreds during the drum drying process, characterized in that, The prediction method includes: Data from the normal operation of the rotary drum dryer is acquired and preprocessed to obtain a dataset. Construct and train a fusion mechanism knowledge model to obtain an augmented dataset; A mechanism constraint prediction model is constructed, and the mechanism constraint prediction model is trained using the augmented dataset; The fusion mechanism knowledge model and the mechanism constraint prediction model are jointly trained to obtain the final tobacco processing strength prediction model; The processing intensity prediction model for tobacco shreds is used to predict the processing intensity based on real-time process variable data.

2. The prediction method according to claim 1, characterized in that, Constructing and training a fusion mechanism knowledge model to obtain augmented datasets includes: A knowledge model based on the fusion mechanism is constructed based on TimeGAN, and a multi-head self-attention mechanism is introduced for each module of the TimeGAN model. A heat conduction equation is introduced into the embedding-recovery network in the fusion mechanism knowledge model; The dataset is input into the fusion mechanism knowledge model to predict the outlet tobacco temperature; A reconstruction loss function is constructed based on the predicted export tobacco temperature; Based on the reconstruction loss function, unsupervised loss and supervised loss are introduced to obtain the optimization objective function; The fusion mechanism knowledge model is trained based on the optimization objective function; The dataset is input into the trained fusion mechanism knowledge model to obtain an augmented dataset.

3. The prediction method according to claim 2, characterized in that, Introducing a heat conduction equation into the embedding-recovery network in the fusion mechanism knowledge model includes: The conduction equation is defined according to formulas (1) and (2). ,(1) ,(2) in, For position and time temperature, The rate of change of the temperature of the exported tobacco shreds with respect to time. For the outlet tobacco temperature at Second derivative in the direction, For thermal diffusivity, For the heat transfer rate of tobacco, The density of the tobacco shreds, This refers to the specific heat capacity of tobacco shreds.

4. The prediction method according to claim 2, characterized in that, The reconstruction loss function is constructed based on the predicted export tobacco temperature, including: The reconstruction loss function is obtained according to formulas (3) to (5). ,(3) ,(4) ,(5) in, To reconstruct the loss function, The reconstruction loss function without embedded mechanistic information. For the mechanism loss function, The temperature data loss function is... To predict the temperature field of the exported tobacco, This represents the actual temperature field of the exported tobacco shreds. To predict the time derivative of the temperature field of the exported tobacco shreds, To predict the spatial second derivative of the temperature field of the exported tobacco shreds, For the diffusion rate, This represents the number of training samples.

5. The prediction method according to claim 2, characterized in that, Based on the aforementioned reconstruction loss function, unsupervised loss and supervised loss are introduced to obtain the optimization objective function, including: The objective function is obtained from formulas (6) to (7). ,(6) ,(7) in, To reconstruct the loss function, For the supervision loss function, For unsupervised loss functions, For parameters embedded in the network, To restore the network parameters, For the parameters of the generator network, For the parameters of the discriminator network, and This is a hyperparameter.

6. The prediction method according to claim 1, characterized in that, Constructing a mechanism-constrained prediction model and training the model using the augmented dataset includes: Based on expert prior knowledge, obtain the mechanism constraints; Based on the aforementioned mechanistic constraints, a monotonic constraint relationship loss function is constructed; A gradual constraint method is introduced, and a total loss function for the mechanistic constraint prediction model is constructed.

7. The prediction method according to claim 6, characterized in that, Based on the aforementioned mechanistic constraints, the monotonic constraint relationship loss function is constructed as follows: Based on formulas (8) and (9), a monotonic constraint relationship between processing intensity and hot air temperature and cylinder wall temperature is constructed. ,(8) ,(9) in, The processing strength versus hot air temperature loss function, The function is: processing strength - cylinder wall temperature loss. for Real-time hot air temperature for Real-time hot air temperature for The actual value of the cylinder wall temperature at any given time. for The actual value of the cylinder wall temperature at any given time. for Predicted processing strength at any time for Real value of processing strength at any time This represents the total number of training samples.

8. The prediction method according to claim 6, characterized in that, The total loss function of the mechanistic constraint prediction model, which introduces the gradual constraint method, includes: ,(10) ,(11) ,(12) in, The total loss function of the mechanism-constrained prediction model. The processing strength versus hot air temperature loss function, The function is: processing strength - cylinder wall temperature loss. For data-driven loss terms, For the gradient constraint function, for Real value of processing strength at any time for Predicted processing strength at any time for Real value of processing strength at any time This is the upper bound of the confidence interval. This is the lower bound of the confidence interval. The total number of training samples, for The weighting coefficients, for The weighting coefficients, for The weighting coefficients.

9. A system for predicting the processing strength of tobacco shreds during the drum drying process, characterized in that, The system includes a processor configured to perform the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 8.