Tobacco core temperature prediction method based on adaptive space-time diagram neural network
Through the method based on the adaptive spatio-temporal graph neural network, the adaptive graph structure is constructed using multi-sensor data, and combined with spatial and temporal feature extractors, the problem of difficult to predict the temperature abnormality of tobacco core in the prior art is solved, and high-accuracy temperature prediction and guarantee of alcoholization quality are achieved.
Patent Information
- Application Number
- CN202311611448.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively predict abnormal temperature of tobacco piles, which makes it difficult to ensure the quality of tobacco alcoholization.
The tobacco core temperature prediction method based on adaptive spatiotemporal graph neural network is adopted to construct an adaptive graph structure through multi-sensor data, combined with a spatial feature extractor and a time feature extractor, and the graph neural network and a long and short-term memory network are used to predict temperature.
It realizes high accuracy prediction of the tobacco core temperature, can warn of temperature abnormalities in advance, and ensures the quality of tobacco leaf alcoholization.
Smart Images

Figure QLYQS_1 
Figure QLYQS_2 
Figure QLYQS_4
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of tobacco core temperature prediction and graph neural networks, and particularly relates to a tobacco core temperature prediction method based on an adaptive spatio-temporal graph neural network. Background Art
[0002] With the development of social economy and the improvement of people's living standards, the demand for tobacco products has been increasing continuously. This requires tobacco production enterprises to strengthen management and improve quality in the storage of cut tobacco to meet the market demand. Tobacco aging is to store the artificially fermented tobacco leaves under certain conditions for a period of time to improve their aroma, taste and physical properties, making them more suitable for processing requirements. There are artificial aging and natural aging. The artificial aging has a fast speed but poor aging effect, and it is rarely used at present. Natural aging is to stack the tobacco leaves in the reserve warehouse according to certain requirements and carry out aging under the condition of controlling certain humidity and temperature. Therefore, strict control of the temperature and humidity environment is an important factor in the tobacco leaf aging process.
[0003] To create a suitable aging environment, generally, it is considered that the environmental temperature during the natural aging process of tobacco leaves is more suitable at 20-30 °C (Liu Bin, Zhu Xiaoqun, Huang Fu, etc. Application of vacuum cooling technology for cut tobacco bales [J]. Tobacco Science & Technology, 2007(08): 5-7+16.). At present, except for individual enterprises that have realized wireless monitoring of the core temperature through information means (Le Chengxing, Lai Lanfeng, (Shen Luheng, etc. Construction of an intelligent warehousing integrated management platform for tobacco leaf raw materials [J]. Fujian Computer, 2021, 37(02): 116-117.), most enterprise warehouses still rely on manual reading of the data on the temperature measuring probe on the cigarette box to detect the core temperature. Considering the workload and equipment cost, it is difficult to improve the detection spatial density and time frequency by this method.
[0004] At present, there are many monitoring schemes for aging temperature, but no scholars have realized the prediction of abnormal environmental temperature. However, temperature abnormality has great destructive power to the production of tobacco leaves, and the prediction of its temperature is extremely urgent.
[0005] In recent years, deep learning has achieved success in applications in many fields due to its powerful capabilities (J. Wang, L. Luo, W. Ye, and S. Zhu, "A defect-detection method of split pins in the catenary fastening devices of high-speed railway based on deep learning," IEEE Transactions on Instrumentation and Measurement, vol. 69, no. 12, pp. 9517-9525, 2020). Research in recent years has shown that deep learning has the powerful ability to automatically extract data features and establish end-to-end non-linear mappings, such as autoencoders (Ren L, Dong J, Wang X, et al. A data-driven auto-CNN-LSTM prediction model for lithium-ion battery remaining useful life[J]. IEEE Transactions on Industrial Informatics, 2020, 17(5): 3478-3487), recurrent neural networks (RNN, Kwon S J, Han D, Choi J H, et al. Remaining-useful-life prediction via multiple linear regression and recurrent neural network reflecting degradation information of 20Ah LiNixMnyCo1-x-yO2 pouch cell[J]. Journal of Electroanalytical Chemistry, 2020, 858: 113729) and its variant long short-term memory network (Y. Zhang, R. Xiong, H. He and M.G. Pecht, "Long Short-Term Memory Recurrent Neural Network for Remaining Useful Life Prediction of Lithium-Ion Batteries," in IEEE Transactions on Vehicular Technology, vol. 67, no. 7, pp. 5695-5705, July 2018, doi: 10.1109 / TVT.2018.2805189), etc. have all been applied to the prediction field and achieved good results. The RNN has the problem of vanishing gradients and only has short-term memory. When the sequence is long, the performance and accuracy of the RNN will deteriorate. Therefore, the LSTM was proposed to add an addition operation to the network using sophisticated gate control, which alleviates the problem of vanishing gradients to a certain extent. In addition, the prediction accuracy can also be improved by appropriately stacking the number of LSTM layers.
[0006] Graph neural networks have been widely applied to data represented in the form of graphs (Scarselli F, Gori M, Tsoi AC, et al. The graph neural network model [J]. IEEE transactions on neural networks, 2008, 20(1): 61-80) and are widely used in various fields, such as the field of time series prediction. Kong et al (Kong Z, Jin X, Xu Z, et al. Spatio-temporal fusion attention: A novel approach for remaining useful life prediction based on graph neural network [J]. IEEE Transactions on Instrumentation and Measurement, 2022, 71: 1-12) proposed a spatio-temporal fusion attention STFA method based on the graph neural network GNN, which uses the prior knowledge of the device to construct a graph structure to predict the remaining life of a turbofan generator. Li et al (Li P, Liu X, Yang Y. Remaining useful life prognostics of bearings based on a novel spatial graph-temporal convolution network [J]. Sensors, 2021, 21(12): 4217) used the average sliding root mean square as a health factor to identify the health state and degradation state of bearings, and then constructed a spatial graph based on the correlation strength between the obtained features to predict the remaining life of the bearings. Although most of the current prediction work based on graph neural networks is effective, the graph construction methods of these methods are not optimal in essence. Either a graph structure for training is determined in advance according to the prior knowledge of the device, or other artificial extraction methods of degradation features are used to construct the graph. These methods not only involve too much human subjectivity, resulting in a decrease in prediction accuracy due to the non-optimal learned graph structure, but also greatly limit the generality of the model and cannot work without providing prior knowledge. Summary of the Invention
[0007] The object of the present invention is to provide a method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network in view of the above problems existing in the prior art. This method can predict the core temperature based on factors such as the temperature, oxygen, and carbon dioxide concentration in the stack, realize early warning of temperature anomalies, and ensure the aging quality of stored tobacco leaves.
[0008] Another object of the present invention is to provide a method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network. Since there is no prior for the graph structure between multiple sensors, this method introduces a graph neural network and a spatial feature extractor. Specifically, an effective self-attention mechanism is added to quantify the importance of one node to another node, and it is improved to a multi-head attention mechanism; a long short-term memory network is introduced and spatio-temporal features are fused; finally, a fully connected layer is used as a predictor to complete the prediction task;
[0009] A method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network, including an optimal graph extractor, a spatial feature extractor, a temporal feature extractor, and a predictor;
[0010] The optimal graph extractor models each sensor as each node in the graph, and all elements of the adjacency matrix A are used as learnable parameters, and these elements are parameterized as a Bernoulli distribution. This generates a corresponding parameterized matrix where A ij ~Ber(p ij ), p ij represents the probability that the value of A ij is 1, and let θ represent the parameters of the downstream model. Then, this module generates an optimal graph structure by optimizing Equation (5).
[0011]
[0012] where X train represents multi-dimensional sensor data for training, represents the mean square error loss function. During the sampling process, the mean is estimated, which may lead to the probability p ijis non-differentiable. Therefore, we adopt the Gumbel-Softmax reparameterization technique introduced by Jang et al. (Sui X, He S, Vilsen S B, et al. A review of non-probabilistic machine learning-based state of health estimation techniques for Lithium-ion battery[J]. Applied Energy, 2021, 300: 117346) to ensure that the gradient of p ij can be calculated.
[0013] The overall framework of the proposed optimal graph extractor is as Figure 1 shown. It consists of two main parts: a feature extractor and an edge predictor. In the feature extractor, a neural network is used to perform non-linear feature extraction on the input training data. Considering that the input data has temporal characteristics, we perform convolution along the time dimension. Then, a fully connected layer is applied to non-linearly map the extracted features into the corresponding one-dimensional feature vectors, denoted as c i , i = (1, 2, 3......, n);
[0014]
[0015] is the training data of the i-th sensor.
[0016] The goal of the edge predictor is to learn a mapping function that takes a pair of feature vectors (c i , c j ) as input and outputs the binary edge probability p ij between the corresponding nodes, which is used to measure the possibility of the link between two nodes. During the entire iterative training process, the parameters of the optimal graph extractor are dynamically updated, so as to obtain the graph structure most suitable for downstream tasks;
[0017] The described spatial feature extractor uses GAT to perform convolution on the graph to achieve spatial feature extraction. Suppose there are n sensors and the time window length is m, then the central node v i has the node feature vector x i ∈R m . In the given time window, the set of all node features is merged to form the feature matrix X t ∈R n×m . The output of each layer of GAT is a new feature matrix X t ; Let v j represent a neighbor node of v i , and let x jRepresents the corresponding feature vector. Where j ∈ N i , N i is the neighbor set of node i. GAT uses the self-attention mechanism to calculate the attention coefficient, denoted as e ij :
[0018]
[0019]
[0020] where α ∈ R 2m′ and W ∈ R m×m′ are both learnable weight matrices, and {||} is the concatenation operation of vectors. To be able to evaluate the relative importance of adjacent nodes relative to the central node, we use the softmax function to normalize the values to obtain α ij Derivation of:
[0021]
[0022] The combination process of (6) and (7) is as Figure 3 shown. According to the above steps, the set of normalized attention coefficients of all nodes is obtained, and the output feature X′ of each node is calculated based on this i :
[0023]
[0024] where σ is the non-linear activation function ELU. Further, to facilitate obtaining and integrating different spatial features from different perspectives, we adopt the multi-head attention mechanism, which will simultaneously construct the above-mentioned multiple independent processes and finally concatenate the obtained features.
[0025]
[0026] In the formula, K represents the number of independent attention mechanisms. The weight matrix of the k-th attention mechanism is denoted by W k and the normalized attention coefficient obtained is denoted by α k , and its schematic diagram is as Figure 4 .
[0027] The spatial feature extractor receives the updated node features of the input sequence arranged in chronological order and sends them to the time feature extractor for further extraction of time features.
[0028] The time feature extractor uses the LSTM architecture as the main framework for extracting time features. Each LSTM unit consists of three different gates: an input gate, an output gate, and a forget gate. These gates are responsible for regulating the information flow inside the cell. The detailed calculation of LSTM is as follows:
[0029] f t = σ(W fh h t-1 + W fx x t + b f )(10)
[0030] i t = σ(W ih h t-1 + W ix x t + b i )(11)
[0031] C t = σ(f t C t-1 + i t tanh(W Ch h t-1 + W Cx x t + b C ))(12)
[0032] o t = σ(W oh h t-1 + W ox x t + b o )(13)
[0033] h t = o t tanh(C t )(14)
[0034] Wherein, W and b respectively represent the learnable weight matrix and the bias, and σ represents the sigmoid activation function. The forget gate can calculate the necessity of information retention by using the state C of the previous cell t-1 and the current input X' t to calculate the necessity of information retention. The input gate uses the output h of the previous cell t-1 and the current input X' t to determine the new information that must be stored in the cell state. The output gate calculates the output value by considering the current cell state and then transmits the output value to the next cell in the sequence. The time feature extractor is used to successfully extract the variation law of the sensor in the time domain and retain the important information at the time level. The goal of the predictor is to predict the core temperature of the tobacco package by using the extracted spatio-temporal features. The fully connected layer is used to map the features into the corresponding one-dimensional feature vectors
[0035] is the output of the model, yt is the true value, t is the time step, and T train represents the length of the training set, T represents the total length of the dataset, and T - T train represents the length of the test set. Then the loss function for model training is defined as:
[0036]
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. Different from traditional mathematical modeling methods, the present invention uses a graph-structured neural network to achieve temperature prediction, with strong adaptability, and also has stronger robustness and universality compared to traditional methods;
[0039] 2. Taking advantage of the ability of deep learning to automatically extract features, the present invention uses an optimal subgraph extractor to learn the optimal graph structure adapted to downstream prediction tasks, and regards the graph structure as a learnable parameter in the neural network, greatly reducing the training cost;
[0040] 3. The present invention regards the temperature output as a classification problem, and the model is simple and effective;
[0041] 4. The present invention fuses time features and spatial features as prediction indicators, makes full use of measured data, and can achieve high-accuracy prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the system block diagram of an embodiment of the present invention.
[0043] Figure 2 is the system block diagram of the optimal graph extractor of the present invention.
[0044] Figure 3 is the schematic diagram of the principle of the attention mechanism of the present invention
[0045] Figure 4 is the schematic diagram of the principle of the multi-head attention mechanism of the present invention
[0046] Figure 5 is the structural schematic diagram of the long short-term memory network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Figure 1 is the system block diagram of an embodiment of the present invention, Figure 2 is the system block diagram of the optimal graph extractor of the present invention, Figure 3 is the schematic diagram of the principle of the attention mechanism of the present invention, Figure 4 is the schematic diagram of the principle of the multi-head attention mechanism of the present invention, Figure 5Schematic diagram of the structure of the long short-term memory network of the present invention. Embodiments of the present invention include the following steps:
[0048] Step 1: During the training phase, the obtained sensor data is input into the optimal graph extractor. The input data for each training round is 256 * 16 * 4 (the number of samples is 256, the length of each sample sequence is 16, and the number of channels is 4). The feature extractor includes two convolutional layers. The first convolutional layer includes convolution conv1, non-linear activation function ReLU, and normalization function BatchNormalization1 (abbreviated as BN1, the same below), obtaining a feature map of 4 * 8 * 4991. Among them, the convolution kernel size of conv1 is 10, and the number of convolution kernels is 8. The second convolutional layer includes convolution conv2, non-linear activation function ReLU, and normalization function BN2, obtaining a feature map of 4 * 16 * 4982. Among them, the convolution kernel size of conv2 is 10, and the number of convolution kernels is 16. The obtained feature map is flattened into 4 * 79712 and passed through a fully connected layer (Fully Connected Layer, abbreviated as FC, the same below), mapped to 4 * 100, and then connected to ReLU and BN3 for normalization. The obtained feature map is an important parameter for constructing the Bernoulli distribution matrix, and finally a probability matrix p(θ) of 4 * 4 is generated. The binary adjacency matrix A of 4 * 4 is output by sampling p. A ij represents the connection between the i-th node and the j-th node, 0 represents no connection, and 1 represents a connection. θ is a trainable parameter and is iteratively updated during the training process.
[0049] During the sampling process, the mean is estimated, which may cause the probability p ij to be non-differentiable. Therefore, we use the gumbel-softmax reparameterization technique introduced by Jang et al. [4] to ensure that the gradient of p ij can be calculated. Finally, the optimal graph structure is generated by Equation (5):
[0050]
[0051] Step 2: The spatial feature extractor uses GAT to perform convolution on the graph to achieve spatial feature extraction. Assume that the central node v i has a node feature vector x i ∈R m . In a given time window, the set of all node features is combined to form a feature matrix X t ∈R n×m . The output of each layer of GAT is a new feature matrix X' t ; Let v j represent a neighbor node of v i , and let x j represent the corresponding feature vector. Among them, j ∈ N i, N i is the neighbor set of node i. GAT uses the self-attention mechanism to calculate the attention coefficients, normalizes the set of attention coefficients, and calculates the output features of each node based on this.
[0052] X' i . In the algorithm implementation, the framework uses the built-in graph attention convolution function GATConv encapsulated by PyG. The first GATConv of the network has an input feature number of 8, an output feature number of 16, and the number of heads of multi-head attention is 6. The second GATConv is defined as conv2, with an input feature number of 16, an output feature number of 32, and the number of heads of multi-head attention is 6.
[0053] Step 3: The time feature extractor uses the LSTM architecture as the main framework for extracting time features. The built-in function LSTM of the torch.nn module is used, with the input size defined as 4, the hidden size defined as 128, and the number of layers defined as 2, that is, it passes through two LSTM cells. The time feature extractor successfully extracts the variation law of the sensor in the time domain and retains the important information at the time level.
[0054] Step 4: The goal of the predictor is to predict the core temperature of the tobacco package using the extracted spatio-temporal features. A fully connected layer is used to map the features into corresponding one-dimensional feature vectors. The fully connected layer FC is embodied as 4 (number of dimensions) cyclic Linear1-ReLU-Linear2 layers in the network. Linear1 reduces the 128-dimensional features to 64 dimensions, and Linear2 reduces the 64 dimensions to 1 dimension to form the model output. The model performance is evaluated according to equations (2) and (3).
[0055] In this embodiment, with the data of 4 sensors as the input, the input data of the first 32 time points are used to predict the next data. The learning rate is set to 0.005, the maximum number of epochs is set to 200, the regularization parameter γ is set to 0.01, and the batch size is set to 256. This network is trained on the pytorch framework, and the graphics card used is NVIDIA GeForce GTX 2080TI.
[0056] The above is only a specific description of the preferred embodiment, but the present invention is not limited to the described embodiment. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these changes are all included in the scope defined by the claims of this application.
Claims
1. A method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network, characterized in that it adaptively obtains the optimal graph structure through the training process, learns spatio-temporal features, and predicts the core temperature of tobacco, including an optimal graph extractor, a graph attention network, a long short-term memory network, and a prediction layer; The optimal graph extractor follows the model training process, learns the optimal graph structure representing the dependencies between multiple sensors, and extracts the most relevant features; The graph attention network (Graph Attention Network, GAT) calculates attention coefficients using the self-attention mechanism at a given time point and extracts spatial features by updating the node feature vectors; The long short-term memory network (Long Short-Term Memory, LSTM) is used to learn temporal features. The LSTM network consists of multiple individual units, and each unit has an input gate, an output gate, and a forget gate to extract important information; The prediction layer is a fully connected layer in form, which maps the fused spatio-temporal features to the prediction output.
2. The method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network according to claim 1, characterized in that The optimal graph extractor accepts the sampled data from multiple sensors as input, including temperature sensors, humidity sensors, oxygen content monitors, carbon dioxide concentration monitors, etc. The time data is sampled by the sliding window method with a time step of 1, and the sensor data is separated based on time.
3. A method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network, characterized in that includes the following steps: 1) Taking the underlying data collected by each sensor as input, for a certain time window, modeling each sensor as different nodes in the graph, and the optimal graph extractor learns the optimal graph structure most suitable for the downstream prediction task through a feature extractor and an edge predictor; 2) For each time window, the spatial feature extractor performs graph convolution on the optimal graph obtained in step 1) using GAT to achieve spatial feature extraction. For each node in the graph, the self-attention mechanism is used to calculate the attention coefficients to quantify the correlation between nodes. In order to facilitate obtaining and integrating different spatial features from different perspectives, the multi-head attention mechanism is adopted in the scheme. 3) The temporal feature extractor uses the LSTM architecture as the main framework for extracting temporal features, which consists of two separate LSTM units, and each unit consists of three different gates: an input gate, an output gate, and a forget gate, to obtain the variation law of features over time. 4) The predictor uses a fully connected layer to output the predicted value of the core temperature of tobacco.
4. The method for predicting the core temperature of tobacco based on an adaptive spatio-temporal graph neural network according to claim 3, characterized in that In step 1), the data of multiple sensors are divided into multiple time windows based on the sliding window method. There is a corresponding feature matrix in each time window, which serves as the node feature matrix of each graph. Based on this, the multi-dimensional sensor data with temporal characteristics are modeled as a temporal graph as the model input. All elements of the adjacency matrix A of the nodes are used as learnable parameters and the element parameters are parameterized as a Bernoulli distribution: A ij ~Ber(p ij ) (1) Here, p ij represents A i the probability of taking the value 1. The optimal graph extractor consists of two main parts: a feature extractor and an edge predictor. The feature extractor uses a neural network to perform non-linear feature extraction on the input training data. The goal of the edge predictor is to learn a mapping function that takes a pair of feature vectors (c i , c j ) as input and outputs the binary edge probability between the corresponding nodes, which is used to measure the likelihood of a link between the two nodes. During the entire iterative training process, the parameters of the optimal graph extractor are dynamically updated so as to obtain the graph structure that is most suitable for the downstream task.
5. A method for predicting the core temperature of tobacco packages based on an adaptive spatio-temporal graph neural network as claimed in claim 3, characterized in that After obtaining the optimal graph in step 1), for the feature matrix X output at a given time window in step 2) t , the GAT layer uses the self-attention mechanism to calculate the attention coefficients, quantifying the importance of each node to another node. In order to evaluate the relative significance of adjacent nodes relative to the central node, we apply the softmax function to normalize the attention coefficient values. According to the above steps, the set of normalized attention coefficients of all nodes is obtained, and the updated feature matrix X′ is given based on this t . The updated node features arranged in chronological order are received by the spatial feature extractor and sent to the temporal feature extractor to further extract temporal features.
6. A method for predicting the core temperature of tobacco packages based on an adaptive spatio-temporal graph neural network as claimed in claim 3, characterized in that The time feature extractor in step 3) consists of multiple individual LSTM units. Each LSTM unit contains an input gate, an output gate and a forget gate. The forget gate can calculate the necessity of information retention by using the state of the previous unit and the current input. The input gate uses the output of the previous unit and the current input to determine the new information that must be stored in the unit state. The output gate calculates the output value by considering the current unit state and then transmits this output value to the next unit in the sequence. The LSTM regulates the information flow and extracts important information in time.
7. A method for predicting the core temperature of tobacco packages based on an adaptive spatio-temporal graph neural network as claimed in claim 3, characterized in that Step 3) successfully extracts the variation laws of each sensor over time and their dependencies on each other. On this basis, the spatial dependencies extracted by the spatial feature extractor are comprehensively processed to obtain comprehensive spatio-temporal fusion features. The predictor uses a fully connected layer to implement the prediction function. The root mean square error (RMSE) and the mean absolute error (MAE) are used as metrics to evaluate the prediction performance of the model. The two evaluation metrics are defined as follows: is the output of the model, y t is the true value, t is the time step, T train represents the length of the training set, T represents the total length of the dataset, T - T train represents the length of the test set. The division ratio of the training set and test set of the model is 8:
2. The loss function for model training is defined as:
Citation Information
Cited By
Optical module junction temperature cloud collaborative regulation and control system and method
CN121680529A
Optical module junction temperature cloud coordination and regulation system and method
CN121680529B