A refrigeration system load prediction method based on a sequence-to-sequence model
The sequence-to-sequence model with sparse self-attention and distillation modules addresses the inefficiencies of existing load prediction algorithms, enhancing accuracy and reducing computational complexity for medium to long-term load forecasting in high-efficiency power systems.
Patent Information
- Application Number
- CN202211626813.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-16
AI Technical Summary
The existing load prediction algorithms have problems such as accumulation of errors and low accuracy in medium and long-term prediction, and traditional models have large amounts of calculation and too many parameters, which lead to difficulty in training.
Using a sequence-to-sequence model method, a multi-layer encoder composed of sparse self-attention module and distillation module is used to combine sparse self-attention and full-quantity self-attention mechanism to optimize the time complexity and feature extraction of the Transformer model, and feature fusion is performed through multi-dimensional time series embedding encoding and residual connection modules.
The accuracy of medium- and long-term load prediction is improved, the calculation complexity and model parameters are reduced, and the prediction accuracy and training efficiency are improved.
Smart Images

Figure CN115982567B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the energy efficiency of refrigeration machine rooms, and particularly relates to a method for predicting the load of a refrigeration system based on a sequence-to-sequence model. Background Art
[0002] In view of the fact that the energy consumption of the power stations of many high-energy-consuming enterprises accounts for nearly 50% of the total energy consumption of the enterprises, therefore, the energy efficiency optimization of high-efficiency power systems has become increasingly important. However, an efficient energy efficiency optimization strategy depends on the accurate prediction of the load of high-efficiency power systems.
[0003] Existing load prediction algorithms mostly adopt regression algorithms and short-time series prediction methods, such as recurrent neural networks like LSTM and GRU. When predicting medium- and long-time series prediction problems, these methods have poor effects due to the sharp accumulation of errors. Currently, with the success of sequence-to-sequence algorithms such as autoencoders and Transformers in the field of natural language processing, etc., it has brought new ideas for us to carry out medium- and long-time series prediction of the load of high-efficiency power systems. However, such models have the following problems:
[0004] 1. Due to the use of the attention mechanism, the computational cost and time complexity of the network become huge compared with traditional deep learning methods such as CNN and RNN;
[0005] 2. For the embedding encoding of time series, the long-order dependence, position dependence, special time points, and mutual weight influence are not fully considered;
[0006] 3. In order to obtain more features, many models adopt the method of stacking multiple encoders, which not only increases a large number of network parameters, but also causes excessive parameters and makes it difficult to train and converge. Summary of the Invention
[0007] In order to solve the above problems, the present invention proposes a method for predicting the load of a refrigeration system based on a sequence-to-sequence model, which solves the problems of error accumulation and low accuracy in medium- and long-time series prediction of the load of high-efficiency power systems, realizes more accurate prediction of the medium- and long-term load of the power station system, and meets the accuracy requirements of energy efficiency optimization applications for engineering-level power systems.
[0008] The present invention can be realized through the following technical solutions:
[0009] A method for predicting the load of a refrigeration system based on a sequence-to-sequence model, which organizes a large amount of historical data into a multi-dimensional time series MTS with a timestamp attribute, trains a sequence-to-sequence prediction model with this, and then uses the time series data of a given window as input, and uses the trained prediction model to predict the load trend in the next period of time.
[0010] Among them, the prediction model includes an encoder and a decoder. The encoder adopts a multi-layer structure composed of a sparse self-attention module and a distillation module. It takes the vector formed by embedding and encoding the time series data as the input of the initial layer structure, and then the output of the previous layer structure as the input of the next layer structure until the final feature map is output. The input of each layer structure is divided into two paths. The first path passes through the sparse self-attention module and the distillation module in sequence to output features Figure Ⅰ , and the second path passes through downsampling to output features Figure Ⅱ . Then the features Figure Ⅰ and the features Figure Ⅱ undergo feature fusion through a residual connection module to output the feature map of the current layer;
[0011] The decoder first fills the target element to be predicted with zeros, then embeds and encodes the generated vector into the masked sparse self-attention module. After that, the generated feature map is used as a query vector and sequentially input into the full self-attention module and the fully connected layer together with the final feature map output by the encoder, and finally the predicted target element is output in a generative manner in real time.
[0012] Furthermore, the sparse self-attention module uses a fully connected layer to project the features after fusing the query matrix Q and the key matrix K into a new probability space, and then calculates the importance score of the query using the formula I(Q) = FC(Q + K). By setting n = clnL Q , L Q represents the number of rows of matrix Q. Select the top n query vectors with the highest scores as the attention scores for subsequent calculations. Among them, FC(·) represents the fully connected operation. The number of input channels of the FC layer is the feature size, and the number of output channels is 1. The shape of I(Q) is L Q ×1.
[0013] Furthermore, the "refinement" process of the features Figure Ⅰ from the j-th sparse self-attention module to the j + 1-th sparse self-attention module is defined as:
[0014]
[0015] where t represents the current time period, [·] AB represents the attention block, γ represents a learnable parameter, DS(·) represents the downsampling operation, Conv1d(·) represents one-dimensional convolution filtering in the time dimension using the ELU(·) activation function.
[0016] Furthermore, representing the multi-dimensional time series MTS in matrix form, first perform z-score normalization processing, then divide it into batches by rows, and then use the following formula to perform embedding encoding to form a vector as the input of the prediction model,
[0017]
[0018] Among them, represents the result after multi-dimensional time series embedding encoding, where i ∈ {1, …, L x}, and t and L x respectively represent the current time period and the number of rows of data. Let the encoded feature dimension be D model ;
[0019] represents that the input feature dimension is d model of the multi-dimensional time series after being projected by one-dimensional convolution, and the vector with the feature dimension of D model ; represents the time encoding;
[0020] represents the position encoding:
[0021]
[0022]
[0023] Among them then use one-dimensional convolution to project it to D model dimensions, and pos represents the current position.
[0024] represents the parameter used to adjust the weights of the position encoding and the time encoding, and the calculation method is expressed as:
[0025]
[0026] Among them, Relu() is the activation function, and Conv1() is the one-dimensional convolution, whose input channel number is D model , and the output channel number is 1.
[0027] Furthermore, perform z-score normalization processing using the following equation,
[0028]
[0029] Among them, d (i,j) is the value of the i-th row and j-th column in the multi-dimensional time series MTS, and D (,) represents all the values in the j-th column. Mean(·) and Std(·) respectively represent the average value and the standard deviation of the j-th column in the dataset.
[0030] The beneficial technical effects of the present invention are as follows:
[0031] 1. A sparse self-attention mechanism based on a neural network is proposed. By using a learnable neural network to obtain prominent contribution attention dot products, the ability of the Transformer self-attention mechanism in terms of time complexity and memory usage is further optimized;
[0032] 2. The way of stacking traditional deep model encoders is improved, and a multi-class pooling and residual distillation mechanism is adopted to obtain as many features as possible without stacking multiple encoders;
[0033] 3. A new time series embedding encoding method is proposed to make the local position encoding and global time encoding have stronger robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic diagram of data demonstration of the present invention;
[0035] Figure 2 is a schematic diagram of the overall structure of the prediction model of the present invention;
[0036] Figure 3 is a partial detailed diagram of the encoder structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following describes in detail the specific embodiments of the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0038] As Figure 1 shown, the present invention provides a refrigeration system load prediction method based on a sequence-to-sequence model. A large amount of historical data is organized into a multi-dimensional time series MTS with timestamp attributes, and a sequence-to-sequence prediction model is trained with this. Then, the time series data of a given window is used as the input, and the trained prediction model is used to predict the load trend in the next period of time.
[0039] Among them, the prediction model includes an encoder and a decoder. The encoder adopts a multi-layer structure composed of a sparse self-attention module and a distillation module. It takes the vector formed by time series data embedding encoding as the input of the initial layer structure, and then the output of the previous layer structure is used as the input of the next layer structure until the final feature map is output. The input of each layer structure is divided into two paths. The first path passes through the sparse self-attention module and the distillation module in sequence to output features Figure Ⅰ , and the second path passes through downsampling to output features Figure Ⅱ , and then the features Figure Ⅰ and the features Figure Ⅱ undergo feature fusion through a residual connection module to output the current layer feature map;
[0040] The decoder first fills the target element to be predicted with zeros, then inputs the generated vector after embedding encoding into the masked sparse self-attention module. After that, the generated feature map is used as the query vector and is input into the full self-attention module and the fully connected layer in sequence with the final feature map output by the encoder. Finally, the predicted target element is output in real time in a generative manner.
[0041] The specific steps are as follows:
[0042] Step 1: Temporal data preparation and preprocessing
[0043] Historical data in an efficient power system is a key factor in model training and establishment. Therefore, the present invention needs to further process the data. The present invention organizes a large amount of historical data into multi-dimensional multi-temporal series MTS with timestamp attributes, and trains the prediction model proposed by the present invention, namely the deep network, based on this. Then, given a certain window of temporal data, the trained deep network can be used to predict the load trend in the next period of time, as Figure 1 shown.
[0044] When the maximum and minimum values of a certain attribute in MTS are unknown, or there are outliers, the min-max normalization of the data is not applicable. Therefore, we define the z-score normalization method for the data as follows:
[0045]
[0046] where d (i,) is the value of the i-th row and j-th column in the MTS dataset, as follows:
[0047]
[0048] D (,) represents all the values in the j-th column. Importantly, Mean(·) and Std(·) represent the mean and standard deviation of the j-th column in the dataset, respectively. It should be noted that D′ is only used for model training to prevent the data gap from being too large and affecting training. The validation and test datasets need to be split from the entire dataset and do not require z-score normalization.
[0049] Generally speaking, we convert the training data D′ into several small batches by rows. For example, if our training data has 100 rows and the minimum batch size can be set to 10 rows, then B = 100 / 10, which means there are a total of 10 batches. The row numbers in each batch are k. The input of the deep network can be defined as:
[0050]
[0051] Among them, b refers to the index of the mini-batch (b ∈ 1, …, B), k is the row index in the mini-batch (k ∈ 1, …, L x ), Lx is the total number of rows of the data time series, and the feature dimension of the mini-batch is D x .
[0052] The predicted value of the output is:
[0053]
[0054] Among them, H is the number of time steps after the current timestamp, k is the row index in the output (k ∈ 1, …, L y ), Ly is the total number of rows of the predicted time series, and the feature dimension of the output is D y .
[0055] Step 2: Construct a sequence-to-sequence model;
[0056] The overall structure of the present invention is as Figure 2 shown and follows the encoder-decoder architecture. During the encoding process, the input is embedded to form a vector, which then enters the sparse self-attention module proposed by the present invention. The output of the sparse self-attention module needs to pass through the distillation module and then output a feature map. To ensure a reduction in loss during the forward propagation of features, the present invention downsamples the embedded vector and fuses it with the output of the distillation module to obtain a feature map. This process is called the residual connection module.
[0057] The decoder receives a long sequence input, first fills the target element to be predicted with zeros, and performs the same embedding encoding as the encoder. The generated vector is input into the masked sparse self-attention module, and the generated feature map is used as a query vector and sent together with the key vector and value vector output by the encoder into the full self-attention module. The fully connected layer instantaneously predicts the output element in a generative manner.
[0058] 2.1 Embedding encoding
[0059] A multi-dimensional time series is a data sequence arranged in chronological order, and its values are in a continuous space. In most cases, the original data is used as the model input after embedding. Therefore, the embedding encoding determines the performance of the data. Previous work designed data embedding by manually adding different time windows, lag operators, and other manual feature derivations. However, this method is too cumbersome and requires domain-specific knowledge. In deep learning models, neural network-based embedding methods have been widely used. In particular, considering that positional semantics and timestamp information will affect the embedding of data, the present invention proposes the following embedding encoding method:
[0060]
[0061] Among them, is the result after encoding the multi-dimensional time series by embedding, where \(i\in\{1,\ldots,L\}\) x}, and let the encoded feature dimension be \(D\) model ;
[0062] represents the input multi-dimensional time series (with a feature dimension of \(d\) model ) is the vector after being projected by a one-dimensional convolution (with a feature dimension of \(D\) model );
[0063] is the positional encoding:
[0064]
[0065]
[0066] where Then, a one-dimensional convolution is used to project it to the \(D\) model dimension. In other words, once the length \(L\) of the input sequence x and the feature dimension \(d\) model are determined, the positional embedding is fixed, and \(pos\) represents the current position.
[0067] Each global timestamp is adopted by a learnable embedding Specifically, the one-hot encoding of year, month, day, hour, minute, second, and holiday is mapped to a vector with the same feature dimension (\(D\) ) through a fully connected layer. For example, for a time of 11:11:11 on November 11, 2022, and if all times are from 2000 to 2022, model ) of the vector. For example, for the year: \([0,0,\ldots,0]\) is a 23-dimensional vector, and for 2022, the first 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0068] For the month: \([0,0,\ldots,0]\) is a 12-dimensional vector, and for November, the 11th 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0069] For the day: \([0,0,\ldots,0]\) is a 31-dimensional vector, and for the 11th, the 11th 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0070] For the hour: \([0,0,\ldots,0]\) is a 24-dimensional vector, and for 11 o'clock, the 11th 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0071] For the minute: \([0,0,\ldots,0]\) is a 60-dimensional vector, and for 11 minutes, the 11th 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0072] For the second: \([0,0,\ldots,0]\) is a 60-dimensional vector, and for 11 seconds, the 11th 0 is set to 1, and then projected to 512 dimensions through a fully connected layer
[0073] Second: [0, 0, … 0] is a 60 - dimensional vector. At the 11th second, the 11th 0 is set to 1, and then it is projected to 512 dimensions through a fully - connected layer.
[0074] In addition, the present invention uses parameters to adjust the weights of position encoding and time encoding, and its calculation method can be expressed as:
[0075]
[0076] where Relu() is the activation function, Conv1d() is the one - dimensional convolution, the number of input channels is D model , and the number of output channels is 1. Such an embedding method can not only mine more MTS features but also facilitate training. The embeddings of both the encoder and the decoder use this method.
[0077] 2.2 Sparse self - attention module
[0078] The traditional full - scale self - attention module is based on tuple inputs, namely query, key, and value, and can be described as:
[0079]
[0080] where Q, K, and V are the matrices of query, key, and value respectively, and d k is the input dimension. In addition, if q i , k i , v i are used to represent the i - th row in the Q, K, and V matrices respectively, then the i - th row of the output can be expressed as:
[0081]
[0082] where k(q i, k j ) is actually an asymmetric exponential kernel function This also means that it can weight the summation of value vectors (V matrix). It requires quadratic dot - product calculations and O(L Q L K ) memory usage, which is the main limitation when expanding the prediction ability.
[0083] Numerous studies have shown that sparse self - attention scores form a long - tailed distribution, that is, a few dot - product pairs contribute the main attention, and other dot - product pairs can be ignored. In this case, if we can calculate the most important n query vectors through the relationship between Q and K, then O(L Q L K) can be optimized. Therefore, the present invention proposes a method for implementing the filtering process of queries through neural network learning, which is defined as follows:
[0084] I(Q) = FC(Q + K)
[0085] where I(Q) represents the importance score of the query, and its shape is L Q ×1; FC(·) represents the fully connected operation. The number of input channels of the FC layer is the feature size, and the number of output channels is 1.
[0086] In the above way, we abandon the method of calculating the attention score in the traditional Transformer. Instead, we project the features after fusing Q and K into a new probability space by using a fully connected layer. In addition, we obtain the score of the query, and by setting n = clnL Q , select the top n query vectors with the highest scores to implement this process. We name this method the sparse self-attention method. In this way, the time and space complexity of the self-attention module can be improved from O(L 2 ) to O(LlnL).
[0087] All in all, our method has the following advantages:
[0088] 1. Reduce the computational workload in the process of filtering query vectors;
[0089] 2. Obtain faster training speed and lower GPU usage;
[0090] 3. Achieve good continuity in the feature domain.
[0091] 2.3 Encoder
[0092] In order to extract the robust long-distance dependencies of long sequence inputs, we propose a feature extraction method of a single encoder and improve the distillation operation, which is shown in Figure 3 . After the vector encoded by the input embedding is calculated by our sparse self-attention module, we obtain the N-head weight matrix of the attention module in the figure, and the "refinement" process from the j-th attention block to the (j + 1)-th attention block can be defined as:
[0093]
[0094] where [·] AB represents the attention block, γ is a learnable parameter; DS(·) represents the downsampling operation, and we use global average pooling (stride = 2) for DS(·). In addition, is calculated as:
[0095]
[0096] Among them, Conv1d(·) performs one-dimensional convolutional filtering in the time dimension using the ELU(·) activation function (with a convolutional kernel size of 3). Although downsampling can reduce the dimension of features, some semantic information will be lost. To mitigate this effect, we obtain as much semantic information as possible (with a stride of 2 for both) by adding a max pooling layer (Maxpool) and an average pooling layer (AvePool) in parallel. We also add a learnable γ to adjust the importance of these two set operations. In addition, to prevent the disappearance of gradients and features, we add a residual connection. After encoding, the length of the feature map becomes one-fourth of the original. Compared with the method of stacking encoders, our method has fewer parameters, faster computational speed, and can also obtain as many features as possible.
[0097] 2.4 Decoder
[0098] The input of the decoder consists of two parts. One part is the output of the encoder (keys and values), and the other part is the query vector calculated by passing the embedded vector filled with target elements as 0 through the masked sparse self-attention module. Compared with the sparse self-attention module, the masked sparse self-attention module masks the future part before calculating Softmax(·) and fills it with the cumulative sum of the V vectors at all time points before each query. This filling method can prevent the model from paying attention to future information. Finally, the query, key, and value are passed to the traditional full self-attention module and passed through a fully connected layer to obtain the prediction result.
[0099] Step 3: Experimental settings
[0100] 3.1 Dataset
[0101] The dataset used in the present invention is the data of the cold machine room refrigeration system, with a time span from May 1, 2022 to August 31, 2022, a data interval of 1 minute, and a total of 177,120 data entries. The data dimensions are: time, cold machine load (kw), primary side load of the cold machine (kw), cooling tower load (kw), cooling pump load (kw), a total of 5 dimensions. We divide the training set, validation set, and test set according to a ratio of 6:2:2. We use a sliding window to process the dataset, with the input sequence length being N, and the next M steps being the ground truth (for MSE training with the predicted values of the model).
[0102] 3.2 Experimental settings
[0103] The deep network model proposed in the present invention is implemented under the Pytorch framework and trained using the Adam optimizer with an initial learning rate of 10 -4 , and a weight decay of 5e -4, with a momentum of 0.9, a batch size of 32, iterated 20 times, and the learning rate decayed by 0.5 every 5 epochs. The training was implemented on an NVIDIA Geforce GTX 3090Ti GPU and an Intel(R) Core(TM) i9-10900K CPU.
[0104] 3.3 Evaluation Metrics
[0105] For the evaluation metrics of the deep model of the present invention, we use CORR, MAE, and MSE, where CORR represents the empirical correlation coefficient, MAE is the mean absolute error, and MSE represents the mean squared error. Their definitions are as follows:
[0106]
[0107]
[0108]
[0109] where y and are the ground truth signal and the system prediction signal respectively. In addition, we set y = y1, y2,..., y n
[0110] and n represents the number of samples.
[0111] Step 4: Model Execution
[0112] Based on the above chiller plant sequence-to-sequence load prediction model, the historical load data of the chiller plant to be processed is used as the input to obtain the corresponding load prediction result for optimizing the energy efficiency of the chiller plant.
[0113] Due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0114] 1) Compared with the traditional prediction method, the CORR obtained from the medium- and long-term prediction (more than 30 steps) is increased by 3%, the MAE is increased by 25%, and the MSE is increased by 70%;
[0115] 2) Compared with the self-attention module of the traditional Transformer, the time and space complexity of the present invention can be reduced from O(L 2 ) to O(LlnL);
[0116] 3) Compared with the traditional Transformer model, the present invention reduces the number of parameters of the model by more than 50%
[0117] above;
[0118] 4) Compared with the self-attention module of traditional Transformers, when the prediction step exceeds 100 steps, the present invention reduces the training time by more than 5 times and the video memory occupancy by more than 2 times.
[0119] In addition, we also conducted application research experiments in more fields, as follows:
[0120] Example 1: Load prediction of the refrigeration system in a cold machine room
[0121] The present invention is applied to the load prediction of the refrigeration system. The input data has a step length of 60, a dimension of 5, a prediction step length of 30, and a prediction dimension of 5, which are time, cold machine load (kw), primary side load of the cold machine (kw), cooling tower load (kw), cooling pump load (kw), respectively. The total load is the sum of the loads of each device. The model construction method is as Figure 2 shown. The CORR, MAE, and MSE results between the predicted values and the true values we obtained are corr: 0.954, mae: 0.188, mse: 0.079, while the traditional Transformer method obtained corr: 0.917, mae: 0.237, mse: 0.136.
[0122] Example 2: Exchange rate prediction based on the deep algorithm of the present invention
[0123] We collected the daily exchange rates of eight countries including Australia, the UK, Canada, Switzerland, China, Japan, New Zealand, and Singapore from 1990 to 2016. There are a total of 7588 data points. We split them into a training set, a validation set, and a test set according to a 6:2:2 ratio. We applied this patent to exchange rate prediction, set the input data step length to 120, the prediction step length to 60, and Figure 2 constructed the model of the present invention. The CORR, MAE, and MSE results between the predicted values and the true values we obtained are corr: 0.911, mae: 0.241, mse: 0.142, while the traditional Transformer method obtained corr: 0.882, mae: 0.275, mse: 0.204.
[0124] Example 3:
[0125] We used a publicly available dataset of residential electricity consumption statistics for testing. This dataset provides two years of data, recorded every minute, and comes from a region in our country. The dataset contains 1,051,200 data points over 365 days in two years. Each data point contains 8-dimensional features, including the recording date of the data point, the predicted value "oil temperature", and 6 different types of external load values. We split the dataset into training set, validation set, and test set in a ratio of 6:2:2. We applied this patent to exchange rate prediction, set the input data step size to 720, the prediction step size to 360, and Figure 2 Construct the model of the present invention. The results of CORR, MAE, and MSE between the predicted values and the true values we obtained are corr: 0.772, mae: 0.308, mse: 0.379, while the traditional Transformer method obtained corr: 0.737, mae: 0.377, mse: 0.458.
[0126] Those skilled in the art should understand that these are only examples. Without departing from the principles and essence of the present invention, various changes or modifications can be made to these embodiments. Therefore, the protection scope of the present invention is defined by the appended claims.
Claims
1. A refrigeration system load prediction method based on a sequence-to-sequence model, characterized in that: A large amount of historical data is organized into a multi-dimensional time series MTS with timestamp attributes, and a sequence-to-sequence prediction model is trained with this. Then, using the time series data within a given window as input, the trained prediction model is used to predict the load trend in the next period of time. Among them, the prediction model includes an encoder and a decoder. The encoder adopts a multi-layer structure composed of a sparse self-attention module and a distillation module. It takes the vector formed by embedding and encoding the time series data as the input of the initial layer structure, and then the output of the previous layer structure is used as the input of the next layer structure until the final feature map is output. The input of each layer structure is divided into two paths. The first path passes through the sparse self-attention module and the distillation module in sequence to output Feature Map Ⅰ, and the second path passes through downsampling to output Feature Map Ⅱ. Then, Feature Map Ⅰ and Feature Map Ⅱ undergo feature fusion through a residual connection module to output the current layer feature map. The decoder first fills the target element to be predicted with zeros, then inputs the generated vector after embedding encoding into the masked sparse self-attention module. After that, the generated feature map is used as the query vector and is input into the full self-attention module and the fully connected layer in sequence with the final feature map output by the encoder. Finally, the predicted target element is output in a generative manner in real time. The "refinement" process of Feature Map Ⅰ from the j-th sparse self-attention module to the j + 1-th sparse self-attention module is defined as: where \(t\) represents the current time period, \([\cdot]\) AB denotes the attention block, \(\gamma\) represents a learnable parameter, and \(DS(\cdot)\) denotes the downsampling operation. \(Conv1d(\cdot)\) denotes one-dimensional convolutional filtering in the time dimension with the \(ELU(\cdot)\) activation function. The multi-dimensional time series MTS is represented in matrix form, first undergoes z-score normalization processing, then is batch-divided by row, and then the following formula is used for embedding encoding to form a vector as the input of the prediction model. Among them, represents the result after the multi-dimensional time series is embedded and encoded, where i ∈ {1, …, L x}, t and L x respectively represent the current time period and the number of rows of data; let the encoded feature dimension be D model ; Indicates that the input feature dimension is d model of the multi-dimensional time series After being projected by one-dimensional convolution, the feature dimension is D model vector; Indicates time encoding; Indicates position encoding: Among them Then, use one-dimensional convolution to project it onto dimension D model Dimension, and pos represents the current position; The parameter used to adjust the weights of the positional encoding and the temporal encoding is calculated as follows: Among them, Relu() is the activation function, and Conv1d() is the one-dimensional convolution, with the number of input channels being D model , and the number of output channels being 1.
2. The load prediction method for a refrigeration system based on a sequence-to-sequence model according to claim 1, wherein: The sparse self-attention module projects the features after fusing the query matrix Q and the key matrix K into a new probability space using a fully connected layer, and then calculates the importance score of the query using the formula I(Q) = FC(Q + K). By setting n = clnL Q , L Q represents the number of rows of the matrix Q, and the top n query vectors with the highest scores are selected as the attention scores for subsequent calculations. Among them, FC(·) represents the fully connected operation. The number of input channels of the FC layer is the feature size, and the number of output channels is 1. The shape of I(Q) is L Q ×1.
3. The load prediction method for a refrigeration system based on a sequence-to-sequence model according to claim 1, wherein: The z-score normalization processing is carried out using the following equation. where d (i,j) is the value at the i-th row and j-th column in the multi-dimensional time series MTS, D (,j) represents all the values in the j-th column, and Mean(·) and Std(·) represent the mean and standard deviation of the j-th column in the dataset, respectively.
Citation Information
Patent Citations
Method for optimizing parameters of refrigerating machine room cooling water system
CN112413762A
Non-intrusive load decomposition method based on Informer model coding structure
CN113393025A