Grain pile temperature prediction method and device based on LSTM encoder-decoder and attention mechanism

The grain pile temperature prediction method based on LSTM encoder-decoder and attention mechanism, utilizing CBAM network and improved multi-head attention network, solves the problems of insufficient prediction accuracy and high computational resource consumption in existing technologies, and achieves efficient and flexible grain pile temperature prediction.

CN118798034BActive Publication Date: 2025-11-11ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410809291.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-11-11
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

Existing methods for monitoring and predicting grain pile temperature cannot simultaneously achieve high accuracy, high timeliness, and high flexibility. Traditional methods suffer from problems such as high computational resource consumption, insufficient prediction accuracy, and high manpower and material costs.

Method used

A grain pile temperature prediction method based on LSTM encoder and decoder and attention mechanism is adopted. Features are extracted by CBAM network, and feature decoding and prediction are performed by combining LSTM encoder and improved multi-head attention network. The temperature prediction model is optimized by using meteorological data.

Benefits of technology

It achieves high-precision single-point temperature prediction with low training overhead, and balances high accuracy, high timeliness and high flexibility, thereby improving the accuracy of feature extraction and prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118798034B_ABST
    Figure CN118798034B_ABST
Patent Text Reader

Abstract

This invention relates to the field of grain storage safety, specifically disclosing a method and apparatus for predicting grain pile temperature based on an LSTM encoder / decoder and an attention mechanism. The method includes: S1, training a pre-constructed grain pile temperature prediction model using historical data of the target grain pile to obtain a trained grain pile temperature prediction model; S2, using the trained grain pile temperature prediction model to predict the temperature of the target grain pile. The grain pile temperature prediction model includes a feature extraction and encoding unit and a feature decoding and prediction unit. This invention innovatively employs a self-designed grain pile temperature prediction model, based on an LSTM encoder / decoder and an attention mechanism, which can achieve high-precision single-point prediction with low training overhead. Simulation comparisons show that the prediction accuracy of this invention is higher than that of traditional methods, balancing the requirements of high accuracy, high timeliness, and high flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of grain storage safety, specifically to: 1. a grain pile temperature prediction method based on LSTM codec and attention mechanism, and 2. a grain pile temperature prediction device using this method. Background Technology

[0002] Grain storage requires a high level of safety. Currently, common methods for monitoring and predicting grain pile temperature include:

[0003] (1) Manually monitor the real-time data of the sensors.

[0004] Technology: By combining Internet of Things (IoT) technology, various sensors are deployed in grain piles to acquire real-time data, and the safety of stored grain is assessed by manually monitoring the data fed back from the sensors.

[0005] Disadvantages: It lacks predictive capabilities, which not only consumes a lot of manpower and resources, but also makes it difficult to predict potential food storage safety hazards in advance.

[0006] (2) Prediction based on mathematical calculation and analysis methods.

[0007] Technically, some researchers analyzed the relationship between real-time temperatures at meteorological stations and the average and maximum temperatures of grain piles, deriving the lag characteristics, correlations, and regression equations of grain surface temperatures at different levels relative to meteorological temperatures. Other researchers proposed a Hidden Scale (HCM) based on Fourier analysis, combined with the least squares method, using a Fourier series model to predict grain pile temperatures based on daily air temperatures.

[0008] Disadvantages: It makes inferences and predictions based solely on the mathematical relationships between data, without taking into account the impact of the spatiotemporal relationships between different locations. Therefore, the prediction accuracy for each sensor temperature measurement point is still significantly insufficient, making it difficult to adapt to complex and ever-changing grain conditions.

[0009] (3) Computer simulation based on thermodynamic field theory.

[0010] Technology: Some researchers have used existing models (energy equation, water conservation equation, rice yellowing model, etc.) to conduct numerical simulations and analyses of natural storage and ventilation combined with natural storage conditions. Other researchers have established a two-dimensional mathematical model based on the finite difference method to calculate heat transfer, simulating temperature changes in barley during storage in cylindrical silos.

[0011] Disadvantages: Accurate thermodynamic simulation requires a large number of physical formulas, which leads to huge consumption of computing resources, especially for large grain warehouses or complex environmental conditions. This makes it difficult to meet the application requirements of actual measurement and control scenarios.

[0012] (4) Prediction model based on simple neural network.

[0013] Technology: This method employs and improves simple models such as LSTM and BP neural networks, training them with large datasets to predict future data trends by learning from historical data. Researchers have proposed an improved BP algorithm, enhancing its speed and convergence, and achieving better prediction results.

[0014] Disadvantages: The network structure is relatively simple and lacks the ability to extract deeper feature information between data. This makes the network poor at processing and understanding contextual information when dealing with data with obvious spatiotemporal relationships and features, resulting in mediocre prediction performance.

[0015] In summary, the above methods cannot simultaneously meet the requirements of high accuracy, high timeliness, and high flexibility. Summary of the Invention

[0016] Therefore, it is necessary to address the problem that existing methods cannot simultaneously achieve high accuracy, high timeliness, and high flexibility by providing a method and device for predicting grain pile temperature based on LSTM codec and attention mechanism.

[0017] This invention is achieved using the following technical solution:

[0018] In a first aspect, this invention discloses a method for predicting grain pile temperature based on an LSTM codec and an attention mechanism, comprising the following steps:

[0019] S1, Use historical data of the target grain pile to train the pre-built grain pile temperature prediction model to obtain the trained grain pile temperature prediction model.

[0020] S2, use the trained grain pile temperature prediction model to predict the temperature of the target grain pile.

[0021] The historical data of the target grain pile includes: historical temperature data of the target grain pile itself and historical meteorological data of the environment in which the target grain pile is located.

[0022] The grain pile temperature prediction model includes a feature extraction and encoding unit and a feature decoding and prediction unit.

[0023] The feature extraction and encoding unit includes: CBAM network, stitcher one, MLP network, LSTM encoder, and stitcher two.

[0024] The CBAM network is used to extract features based on the historical temperature data of the target grain pile itself, resulting in the grain pile feature group F. cbam The splicer is used to group the grain pile features F. cbam With meteorological characteristic group F mf Perform corresponding concatenation to obtain the concatenated feature group F. inMLP networks are used to extract spliced ​​feature groups F in The association between them yields the association feature group F. mlp The LSTM encoder is used to extract the spliced ​​feature group F. in The encoded hidden state at each time step; splicer two is used to connect the associated feature group F mlp With splicing feature group F in The encoded hidden states at each time step are concatenated to obtain the feature encoding tensor F.

[0025] in, Let t represent the grain pile characteristics at time step t; t∈[1,T], where T represents the total number of time steps of the historical data of the target grain pile.

[0026] Meteorological characteristic group F mf This refers to the historical meteorological data of the environment in which the target grain pile is located; This represents the meteorological characteristics at time step t.

[0027] This represents the splicing feature at time step t.

[0028] This represents the output vector at time step t.

[0029] F = {f1; f2; ...; f T};f t Let represent the feature encoding tensor at time step t.

[0030] The feature decoding and prediction unit includes: an LSTM decoder, an improved multi-head attention network, a splicer three, and a predictive linear network.

[0031] The LSTM decoder consists of W LSTM decoding units; W > T.

[0032] The first T LSTM decoding units are used to decode the hidden state of each time step based on the feature encoding tensor F and the historical temperature data of the target grain pile itself.

[0033] The remaining LSTM decoding units are used in conjunction with an improved multi-head attention network, a splicer three, and a predictive linear network to predict the temperature of the target point.

[0034] The implementation of this grain pile temperature prediction method based on LSTM codec and attention mechanism is according to the method or process of embodiments of this disclosure.

[0035] In a second aspect, the present invention discloses a grain pile temperature prediction device based on LSTM codec and attention mechanism, which uses the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in the first aspect.

[0036] The grain pile temperature prediction device based on LSTM codec and attention mechanism includes: a model training module and a temperature prediction module.

[0037] The model training module is used to train a pre-built grain pile temperature prediction model using historical data from the target grain pile, resulting in a trained grain pile temperature prediction model. The temperature prediction module is used to predict the temperature of the target grain pile using the trained grain pile temperature prediction model.

[0038] The implementation of this grain pile temperature prediction device based on LSTM codec and attention mechanism is according to the method or process of embodiments of this disclosure.

[0039] Thirdly, the present invention discloses a readable storage medium storing computer program instructions, which are read and executed by a processor to perform the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in the first aspect.

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. The method of this invention innovatively adopts a self-designed grain pile temperature prediction model, which is based on LSTM encoder and decoder and attention mechanism, and can achieve high-precision single-point prediction with low training overhead. Through simulation comparison, the prediction accuracy of this invention is higher than that of traditional methods, and takes into account the requirements of high accuracy, high timeliness and high flexibility.

[0042] 2. The grain pile temperature prediction model used in this invention innovatively employs a CBAM network to optimize feature extraction. This allows for weighting of features at different levels, better capturing the correlation between data, and thus improving the expressive power of the feature extraction results. Furthermore, based on the analysis of the actual temperature change patterns of the grain pile and the correlation between the temperatures of sensors at different locations, this invention ultimately selects the number of layers in the temperature sensor distribution network in the grain pile as the number of channels, thereby enabling the rational use of the CBAM network to extract and optimize features.

[0043] 3. This invention takes into account the different spatiotemporal relationships between different data, and processes grain pile data and meteorological data separately, so that even when adding reference data, it will not add too much extra calculation.

[0044] 4. This invention employs an LSTM encoder and an MLP network to encode features, enhancing the expressive power of feature encoding and enabling the encoding results to contain more spatiotemporal information about the target point, which is beneficial for further processing and prediction during subsequent decoding.

[0045] 5. This invention improves the multi-head self-attention network based on the multi-head self-attention network in the Transformer model, and uses it in conjunction with the LSTM decoder to adaptively construct the most suitable context vector to complete the prediction, effectively improving the performance of the model and the accuracy of the prediction. Attached Figure Description

[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A flowchart of the grain pile temperature prediction method based on LSTM codec and attention mechanism provided in Embodiment 1 of the present invention;

[0048] Figure 2 for Figure 1 Structural diagram of COFCO's compost temperature prediction model;

[0049] Figure 3 for Figure 2 Structure diagram of the CBAM network;

[0050] Figure 4 for Figure 2 Structure diagram of a medium-sized LSTM encoder;

[0051] Figure 5 for Figure 2 Structure diagram of the LSTM decoder;

[0052] Figure 6 for Figure 5 Structure diagram of the improved multi-head attention network;

[0053] Figure 7 This is a structural diagram of the target grain pile provided in Embodiment 2 of the present invention;

[0054] Figure 8 This is a comparison chart of the predictions of target point 1 by the four methods in Embodiment 2 of the present invention;

[0055] Figure 9 This is a comparison chart of the predictions of target point two by the four methods in Embodiment 2 of the present invention;

[0056] Figure 10 This is a comparison chart of the predictions of target point three using four methods in Embodiment 2 of the present invention. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.

[0060] Example 1

[0061] Please refer to Figure 1 , Figure 1 This is a simplified flowchart of a grain pile temperature prediction method based on an LSTM encoder / decoder and an attention mechanism, which aims to predict the temperature of a target grain pile.

[0062] It should be noted that several temperature sensors for monitoring temperature are evenly distributed in the target grain pile (equally spaced along the length, width, and height directions), forming a temperature sensor distribution network.

[0063] like Figure 1 As shown, the grain pile temperature prediction method based on LSTM codec and attention mechanism includes:

[0064] S1. Use historical data of the target grain pile to train the pre-built grain pile temperature prediction model to obtain the trained grain pile temperature prediction model.

[0065] The total time step for the historical data of the target grain pile is T.

[0066] It is important to note that the historical data for the target grain pile includes not only the historical temperature data of the grain pile itself, but also the historical meteorological data of the surrounding environment. This is because meteorological data influences the long-term or short-term temperature change trends at each temperature measurement point (temperature measurement points located on the outside of the grain pile tend to show short-term trends, while those located on the inside tend to show long-term trends), and this needs to be taken into account to improve the accuracy of predictions.

[0067] Meteorological data typically includes temperature and humidity. While other types of meteorological data can be added, they are generally far less than the temperature data of the grain pile itself. Therefore, adding meteorological data will not significantly increase the computational load.

[0068] See Figure 2 The grain pile temperature prediction model includes: a feature extraction and encoding unit, and a feature decoding and prediction unit.

[0069] First, such as Figure 2 As shown, the feature extraction and encoding unit includes: CBAM network, stitcher one, MLP network, LSTM encoder, and stitcher two.

[0070] ①The CBAM network is used to extract features based on the historical temperature data of the target grain pile itself, resulting in the grain pile feature group F. cbam .

[0071] Among them, F cbam It can be represented as:

[0072] Let represent the grain pile characteristics at time step t; t∈[1,T].

[0073] For details, please refer to Figure 3 The CBAM network includes: channel attention network, weight adjuster 1, spatial attention network, and weight adjuster 2.

[0074] The channel attention network is used to process the historical temperature data of the target grain pile itself to obtain the channel feature weights on different layers of the target grain pile. Specifically, the channel attention network uses the number of layers in the temperature sensor distribution network within the target grain pile as the number of channels.

[0075] More specifically, in the channel attention network, the historical temperature data of the target grain pile corresponding to each channel is used as input; firstly, average pooling and max pooling are performed on the input in each channel to obtain the average pooling result and max pooling result; then, these two results are processed and output through a shared MLP network layer and added together; finally, the summed result is output through the Sigmoid activation function to obtain the channel feature weights on different layers of the target grain pile.

[0076] The weight adjuster is used to adjust the historical temperature data of the target grain pile itself according to the channel feature weights, so as to obtain the channel-optimized feature data set. In other words, multiplying the channel feature weights by the corresponding historical temperature data yields the channel-optimized feature data set.

[0077] Spatial attention networks are used to process the feature data set after channel optimization to obtain the position feature weights at different locations of the target grain pile.

[0078] More specifically, in the spatial attention network, the channel dimension of the channel-optimized feature data group is first compressed, and the corresponding maximum and average values ​​are calculated to obtain the maximum value result and the average value result; then the maximum value result and the average value result are concatenated along the channel dimension and calculated through a convolutional layer; next, the result calculated by the convolutional layer is output through the Sigmoid activation function to obtain the position feature weights at different positions of the target grain pile.

[0079] Weight adjuster 2 is used to adjust the channel-optimized feature data set according to the position feature weights, to obtain the position-optimized feature data set, which is then used as the grain pile feature set F. cbam In other words, multiplying the position feature weights by the corresponding channel-optimized feature data set yields the position-optimized feature data set.

[0080] ②See Figure 2 The splicer is used to group the grain pile features F cbam With meteorological characteristic group F mf Perform corresponding concatenation to obtain the concatenated feature group F. in .

[0081] Among them, meteorological characteristic group F mf This refers to the historical meteorological data of the environment in which the target grain pile is located.

[0082] So, F mf F in It can be represented as:

[0083] This represents the meteorological characteristics at time step t.

[0084] This represents the splicing feature at time step t.

[0085] in, The calculation formula is designed as follows:

[0086]

[0087] In the formula, COMBINE represents the splicing operation.

[0088] ③See Figure 2 MLP networks are used to extract spliced ​​feature groups F in The association between them yields the association feature group F. mlp .

[0089] MLP networks employ a typical design with two fully connected layers combined with an activation function, the implementation of which is effective for F... in Perform correlation extraction.

[0090] Among them, F mlp It can be represented as:

[0091] This represents the output vector at time step t.

[0092] in, The calculation formula is designed as follows:

[0093]

[0094] In the formula, W1, b1, W2, and b2 are all parameter matrices; M represents the vector length, H represents the number of hidden layers in the LSTM encoder, and RELU(.) represents the ReLU activation function.

[0095] ④ See Figure 2 The LSTM encoder is used to extract the concatenated feature group F. in The encoded hidden state at each time step.

[0096] Compared to complex models such as convolutional neural networks and Transformer encoders and decoders, which have slower training speeds, LSTM encoders are more efficient and can extract the dependency features (i.e., hidden states and unit states) of each time step in both short and long time.

[0097] For details, please refer to Figure 4 The LSTM encoder consists of T LSTM encoding units.

[0098] Wherein, the t-th LSTM coding unit is based on the concatenation features of the t-th time step. The encoded hidden state at time step (t-1) The state of the coding unit at time step t-1 The encoded hidden state at time step t is calculated. State of the encoding unit at time step t

[0099] In other words, The calculation formula is designed as follows:

[0100]

[0101] In the formula, LSTM represents the LSTM network computation process.

[0102] in, All are zero vectors.

[0103] ⑤ See Figure 2 The second splicer is used to connect the associated feature groups F mlp With splicing feature group F in The encoded hidden states at each time step are concatenated to obtain the feature encoding tensor F.

[0104] Where F = {f1; f2; ...; f T};f t Let represent the feature encoding tensor at time step t.

[0105] Among them, f t The calculation formula is designed as follows:

[0106]

[0107] In the formula, COMBINE represents the splicing operation.

[0108] II. Figure 2 As shown, the feature decoding and prediction unit includes: an LSTM decoder, an improved multi-head attention network, a splicer three, and a prediction linear network.

[0109] See Figure 5 The LSTM decoder consists of W LSTM decoding units; W > T.

[0110] Similar to LSTM encoders: Compared to complex models such as convolutional neural networks and Transformer encoders and decoders, which have slower training speeds, LSTM decoders are more efficient and can extract the dependency features (i.e., hidden states and unit states) of each time step in both short and long time.

[0111] 2.1 The first T LSTM decoding units are used to decode the hidden state of each time step based on the feature encoding tensor F and the historical temperature data of the target grain pile itself.

[0112] Specifically, in the first T LSTM decoding units, the t-th LSTM decoding unit encodes the tensor f based on the features of the t-th time step. t The historical true temperature r of the target point at time step t-1 t-1 The decoding hidden state h at time step t-1 t-1 The state c of the decoding unit at time step t-1 t-1The decoded hidden state h at time step t is calculated. t The state c of the decoding unit at time step t t .

[0113] Since the LSTM decoding unit can only have 3 inputs, the feature encoding tensor f at the t-th time step needs to be... t The historical true temperature r of the target point at time step t-1 t-1 The vectors are concatenated and combined into a single vector.

[0114] In other words, h t c t The calculation formula is designed as follows:

[0115] (h t ,c t =LSTM[COMBINE(f) t ,r t-1 ),h t-1 ,c t-1 ];

[0116] In the formula, COMBINE represents the splicing operation, and LSTM represents the LSTM network computation process.

[0117] Where h0 and c0 are both zero vectors.

[0118] 2.2 The remaining LSTM decoding units are used in conjunction with the improved multi-head attention network, splicer three, and prediction linear network to predict the temperature of the target point.

[0119] ① The improved multi-head attention network is obtained by structurally adjusting the multi-head attention network in the Transformer model. Its specific construction method includes:

[0120] S1, Obtain the multi-head attention network in the Transformer model and use it as the original multi-head attention network;

[0121] S2, an improved multi-head attention network is obtained by adjusting the original multi-head attention network;

[0122] Specifically, the adjustment methods include:

[0123] 1. The linear layers from the input end to the formation of input data Q, K, and V in the original multi-head attention network have been removed. In the original multi-head attention network, the same data input from the input end is transformed into input data Q, K, and V through different linear layers; these linear layers have been removed here.

[0124] 2. Directly use different inputs as input data Q, K, and V, and add a linear layer to transform Q and a linear layer to transform K in each head. That is, instead of using the same data to generate Q, K, and V, different inputs are used directly; among them, Q and K need to be transformed for subsequent calculations.

[0125] Specifically, when making a prediction at the w-th time step (w>T):

[0126] Decode the hidden state group H as input data Q, H = {h1; h2; ...; h T};

[0127] The state c of the decoding unit at time step w-1 w-1 K is the input data;

[0128] The feature encoding tensor F is used as the input data V.

[0129] Thus, the improved multi-head attention network calculates the output vector based on the input data Q, K, and V.

[0130] Specifically, the improved multi-head attention network consists of N heads. All N heads share the same set of Q, K, and V. See [link to relevant documentation]. Figure 6 The calculation process for the nth (n∈[1,N]) head is as follows:

[0131] 1) Perform linear transformations on Q and K respectively; where the dimension of the transformed vector K is denoted as d. k ;

[0132] 2) Process the linear transformation result in 1) to perform matrix multiplication and obtain the similarity matrix;

[0133] 3) Divide each element of the similarity matrix in 2) by Performing a scaling operation can reduce the variance and thus improve the stability of training.

[0134] 4) Calculate the influence weight of each time step with respect to the current time step using the Softmax function on the Scale result in 3).

[0135] 5) Perform matrix multiplication between the result of the Softmax function in 4) and V, and use the result as the output m of this head. n .

[0136] Finally, the calculation results of the N heads are concatenated into an output vector m, and the desired output dimension is obtained through a linear layer transformation, thus yielding the output vector.

[0137] ②See Figure 4 The splicer will output vector and the temperature value z at the (w-1)th time step w-1 By concatenating the corresponding vectors, we obtain the context vector for the w-th time step.

[0138] It should be noted that when w = T+1, z w-1 =r w-1 When w > T+1, z w-1 =p w-1 .

[0139] Where, r w-1 This represents the historical true temperature of the target point at the (w-1)th time step;

[0140] p w-1 This represents the predicted temperature value of the target point at the (w-1)th time step.

[0141] In other words, The calculation formula can be designed as follows:

[0142]

[0143] In the formula, COMBINE represents the splicing operation.

[0144] ③See Figure 5 The w-th LSTM decoding unit is based on the context vector at the w-th time step. The decoded hidden state h at the (w-1)th time step w-1 The state c of the decoding unit at time step w-1 w-1 The decoded hidden state h at the w-th time step is calculated. w The state c of the decoding unit at the w-th time step w .

[0145] In other words, h w c w The calculation formula can be designed as follows:

[0146]

[0147] In the formula, LSTM represents the LSTM network computation process.

[0148] ④ See Figure 5 The predictive linear network is based on the decoded hidden state h at the w-th time step. w Calculate the predicted temperature value p at the target point in the w-th time step. w .

[0149] Where, p w The calculation formula can be designed as follows:

[0150] p w =W p h w +b p ;

[0151] In the formula, W p b p Both are parameter matrices;

[0152] The grain pile temperature prediction model based on the above structure is trained using historical data of the target grain pile, and the training method is the conventional model training method:

[0153] The historical data of the target grain pile is divided proportionally into: training set and test set;

[0154] The training set was used to iterate the grain pile temperature prediction model for multiple rounds, and a loss function (using the mean squared error function) was constructed for the target point to adjust the model parameters through backpropagation.

[0155] After each iteration, performance testing is performed on the test set until the model converges;

[0156] The model with the best performance test results is selected as the trained grain pile temperature prediction model.

[0157] S2, use the trained grain pile temperature prediction model to predict the temperature of the target grain pile.

[0158] Referring to the working principle of the model in S1, the trained grain pile temperature prediction model can continuously predict the temperature of the target point at future time steps with high accuracy. The target point can be any temperature measurement point within the grain pile, offering high flexibility.

[0159] Example 2

[0160] This embodiment 2 discloses a grain pile temperature prediction device based on LSTM codec and attention mechanism, which uses the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in embodiment 1.

[0161] The grain pile temperature prediction device based on LSTM codec and attention mechanism includes: a model training module and a temperature prediction module.

[0162] The model training module is configured to train a pre-built grain pile temperature prediction model using historical data from the target grain pile, resulting in a trained grain pile temperature prediction model. The temperature prediction module is configured to use the trained grain pile temperature prediction model to predict the temperature of the target grain pile.

[0163] Since this grain pile temperature prediction device based on LSTM codec and attention mechanism uses the grain pile temperature prediction method based on LSTM codec and attention mechanism in Example 1, it also has the same effect as Example 1, and will not be repeated here.

[0164] Example 3

[0165] This embodiment 3 performs simulation verification on the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in embodiment 1 (referred to as the Proposed approach), and compares its performance with three other existing methods.

[0166] The other three methods include:

[0167] 1. The LSTM baseline method (abbreviated as LSTM) is mainly used to process time series data and can effectively learn and extract long-term and short-term dependencies between data.

[0168] 2. The CNN-LSTM method (abbreviated as CNN-LSTM) combines the advantages of CNN networks and LSTM. The CNN network is responsible for extracting local features from the input sequence and converting them into a higher-level representation, which is then fed into the LSTM for time series processing.

[0169] 3. The ResNET-LSTM method (referred to as ResNET-LSTM) combines the characteristics of ResNET networks and LSTM. Residual connections allow information to be transferred more efficiently within the network, helping to alleviate the vanishing gradient problem in deep networks. The ResNET network is responsible for extracting sequence features, while LSTM is responsible for processing sequence relationships.

[0170] In this embodiment 3, a grain depot in the Jianghuai region was used as the target grain pile. The distribution of temperature sensors inside the grain depot is as follows: Figure 7 As shown: there are 7 points in the length direction, 5 points in the width direction, and 4 points in the height direction, for a total of 140 temperature measurement points. That is to say, the distribution network of temperature sensors has 4 layers.

[0171] A continuous 960-day period of data, from May 16, 2020 to December 31, 2022, was selected as the data source for the verification experiment. The data for each day included temperature values ​​read by all 140 temperature sensors in the grain depot, and the corresponding historical meteorological data (including temperature and humidity) was obtained from the China Meteorological Administration website.

[0172] In addition, before conducting the verification experiment, the historical data was preprocessed to remove missing data values ​​and outliers caused by unavoidable factors, and numerical interpolation was performed on these removed time points to obtain the experimental dataset.

[0173] The experimental dataset is divided into three sets according to the following proportions: training set, test set, and validation set.

[0174] The same training and testing sets were used to train and test the models of the four methods. Three test points were selected in the target grain pile: test point 1 (2,3,4), test point 2 (5,5,3), and test point 3 (7,1,4).

[0175] The validation set serves as the actual measurement (i.e., observation) corresponding to the prediction results, and the RMSE metric is used to characterize the prediction accuracy of the four methods.

[0176] The RMSE indicator is calculated using the following formula:

[0177]

[0178] In the formula, y w This represents the actual grain temperature data at time w (the w-th time step). This represents the predicted grain temperature data at time w (the w-th time step), where x represents the number of temperature samples.

[0179] See the comparison results of the four temperature prediction methods. Figures 8-10 The corresponding RMSE metrics are shown in Table 1 below.

[0180] Table 1 Comparison of RMSE metrics for the four methods

[0181]

[0182] Based on the comparison, we can conclude that:

[0183] 1) For test point 1 (2,3,4), the RMSE of the prediction results of the method disclosed in Example 1 is reduced by 38.2% compared with LSTM, 22.3% compared with CNN-LSTM, and 22.7% compared with ResNET-LSTM.

[0184] 2) For test point 2 (5,5,3), the RMSE of the prediction results of the method disclosed in Example 1 is reduced by 71.5% compared with LSTM, 71.3% compared with CNN-LSTM, and 44.3% compared with ResNET-LSTM.

[0185] 3) For test point 3 (7,1,4), the RMSE of the prediction results of the method disclosed in Example 1 is reduced by 72.2% compared with LSTM, reduced by 68.5% compared with CNN-LSTM, and increased by 5.4% compared with ResNET-LSTM.

[0186] 4) The RMSE of the three test points was averaged and compared. The method disclosed in Example 1 reduced the RMSE by 65.7% compared to LSTM, 62.3% compared to CNN-LSTM, and 25.8% compared to ResNET-LSTM.

[0187] In summary, the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in Example 1 is significantly better than the other three schemes in terms of overall prediction accuracy.

[0188] Furthermore, this embodiment 3 also investigates the advantages and disadvantages of four methods:

[0189] While LSTM and CNN-LSTM are relatively flexible and efficient, their accuracy is poor.

[0190] While ResNET-LSTM boasts excellent accuracy, its overhead is prohibitively high.

[0191] The grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in Example 1 combines high accuracy, high timeliness and high flexibility.

[0192] Example 4

[0193] This embodiment 4 also discloses a readable storage medium storing computer program instructions. When the computer program instructions are read and run by a processor, the grain pile temperature prediction method based on LSTM codec and attention mechanism disclosed in embodiment 1 is executed.

[0194] When applying the method of Example 1, it can be applied in the form of software, such as by designing a program that can run independently on a computer-readable storage medium, which can be a USB flash drive. The program of the entire method can be designed to be launched by an external trigger.

[0195] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0196] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for predicting grain pile temperature based on LSTM codec and attention mechanism, characterized in that, Includes the following steps: S1, Use historical data of the target grain pile to train the pre-built grain pile temperature prediction model to obtain the trained grain pile temperature prediction model. The historical data of the target grain pile includes: the historical temperature data of the target grain pile itself and the historical meteorological data of the environment in which the target grain pile is located; The grain pile temperature prediction model includes: a feature extraction and encoding unit, and a feature decoding and prediction unit; The feature extraction and encoding unit includes: a CBAM network, a stitcher one, an MLP network, an LSTM encoder, and a stitcher two; the CBAM network is used to extract features based on the historical temperature data of the target grain pile itself, obtaining the grain pile feature group F. cbam The splicer is used to group the grain pile features F. cbam With meteorological characteristic group F mf Perform corresponding concatenation to obtain the concatenated feature group F. in MLP networks are used to extract spliced ​​feature groups F in The association between them yields the association feature group F. mlp The LSTM encoder is used to extract the spliced ​​feature group F. in The encoded hidden state at each time step; splicer two is used to connect the associated feature group F mlp With splicing feature group F in The encoded hidden states at each time step are concatenated to obtain the feature encoding tensor F; in, The grain pile features are represented at time step t; t∈[1,T]; T represents the total number of time steps in the historical data of the target grain pile; meteorological feature group F mf This refers to the historical meteorological data of the environment in which the target grain pile is located; This represents the meteorological characteristics at time step t. This represents the splicing feature at time step t; F represents the output vector at time step t; F = {f1; f2; ...; f...} T };f t This represents the feature encoding tensor at time step t; The feature decoding and prediction unit includes: an LSTM decoder, an improved multi-head attention network, a stitcher three, and a prediction linear network; the LSTM decoder includes: W LSTM decoding units; W > T; the first T LSTM decoding units are used to decode the hidden state of each time step based on the feature encoding tensor F and the historical temperature data of the target grain pile itself; the remaining LSTM decoding units are used to combine with the improved multi-head attention network, the stitcher three, and the prediction linear network to predict the temperature of the target point; Improved methods for constructing multi-head attention networks include: S100, obtain the multi-head attention network in the Transformer model and use it as the original multi-head attention network; S200, an improved multi-head attention network is obtained by adjusting the original multi-head attention network; The adjustment methods include: The linear layers from the input end to the formation of input data Q, K, V in the original multi-head attention network were removed; different inputs were directly used as input data Q, K, V, and a linear layer that transforms Q and a linear layer that transforms K were added to each head; S2, the temperature of the target grain pile was predicted using the trained grain pile temperature prediction model.

2. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 1, characterized in that, The CBAM network includes: channel attention network, weight adjuster 1, spatial attention network, and weight adjuster 2; The channel attention network is used to process the historical temperature data of the target grain pile itself to obtain the channel feature weights on different layers of the target grain pile; the channel attention network uses the number of layers of the temperature sensor distribution network in the target grain pile as the number of channels; The weight adjuster is used to adjust the historical temperature data of the target grain pile itself according to the channel feature weights to obtain the channel-optimized feature data set. Spatial attention networks are used to process the feature data set after channel optimization to obtain the position feature weights at different locations of the target grain pile; Weight adjuster 2 is used to adjust the channel-optimized feature data set according to the position feature weights, to obtain the position-optimized feature data set, which is then used as the grain pile feature set F. cbam .

3. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 2, characterized in that, The LSTM encoder consists of T LSTM encoding units; Wherein, the t-th LSTM coding unit is based on the concatenation feature f at the t-th time step. t in The encoded hidden state at time step t-1 The state of the coding unit at time step t-1 The encoded hidden state at time step t is calculated. State of the encoding unit at time step t All are zero vectors.

4. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 2, characterized in that, The calculation formula is designed as follows: In the formula, COMBINE represents the splicing operation; The calculation formula is designed as follows: In the formula, W1, b1, W2, and b2 are all parameter matrices; M represents the vector length, H represents the number of hidden layers in the LSTM encoder; ReLU(.) represents the ReLU activation function; The calculation formula is designed as follows: In the formula, LSTM represents the LSTM network computation process; f t The calculation formula is designed as follows: In the formula, COMBINE represents the splicing operation.

5. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 1, characterized in that, In the first T LSTM decoding units, the t-th LSTM decoding unit encodes the tensor f based on the features of the t-th time step. t The historical true temperature r of the target point at time step t-1 t-1 The decoding hidden state h at time step t-1 t-1 The state c of the decoding unit at time step t-1 t-1 The decoded hidden state h at time step t is calculated. t The state c of the decoding unit at time step t t ; h0 and c0 are both zero vectors.

6. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 5, characterized in that, When making a prediction at the w-th time step, the decoded hidden state group H is used as the input data Q; the decoding unit state c at the (w-1)-th time step... w-1 The input data is K; the feature encoding tensor F is the input data V; where w > T; H = {h1; h2; ...; h T }; The improved multi-head attention network calculates the output vector based on input data Q, K, and V. The splicer will output vector and the temperature value z at the (w-1)th time step w-1 By concatenating the corresponding vectors, we obtain the context vector for the w-th time step. The w-th LSTM decoding unit is based on the context vector of the w-th time step. The decoded hidden state h at the (w-1)th time step w-1 The state c of the decoding unit at time step w-1 w-1 The decoded hidden state h at the w-th time step is calculated. w The state c of the decoding unit at the w-th time step w ; The predictive linear network is based on the decoded hidden state h at the w-th time step. w Calculate the predicted temperature value p at the target point in the w-th time step. w ; Where, when w = T+1, z w-1 =r w-1 When w > T+1, z w-1 =p w-1 .

7. The grain pile temperature prediction method based on LSTM encoder / decoder and attention mechanism according to claim 6, characterized in that, h t c t The calculation formula is designed as follows: (h t ,c t )=LSTM[COMBINE(f t ,r t-1 ),h t-1 ,c t-1 ]; In the formula, COMBINE represents the splicing operation, and LSTM represents the LSTM network computation process; The calculation formula is designed as follows: In the formula, COMBINE represents the splicing operation; h w c w The calculation formula is designed as follows: In the formula, LSTM represents the computation process of the LSTM network; p w The calculation formula is designed as follows: p w =W p h w +b p ; In the formula, W p b p Both are parameter matrices; 8. A grain pile temperature prediction device based on LSTM codec and attention mechanism, characterized in that, The grain pile temperature prediction method based on LSTM codec and attention mechanism as described in any one of claims 1-7 was used; The grain pile temperature prediction device based on LSTM codec and attention mechanism includes: The model training module is used to train a pre-built grain pile temperature prediction model using historical data of the target grain pile, so as to obtain a trained grain pile temperature prediction model. as well as The temperature prediction module is used to predict the temperature of a target grain pile using a trained grain pile temperature prediction model.

9. A readable storage medium, characterized in that, The readable storage medium stores computer program instructions, which are read and executed by a processor to perform the steps of the grain pile temperature prediction method based on LSTM codec and attention mechanism as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Ocean temperature field spatio-temporal data generation method based on separation attention mechanism

    CN115271199A

  • Facility environment multi-step prediction method based on pooling attention

    CN116578862A