Multi-region load forecasting method and system integrating global enhancement and local attention features
By combining the load feature extraction model of long and short-term memory network and spatial attention mechanism, as well as the meteorological feature extraction model of token to token and Vision Transformer models, the problem of global and local feature fusion in multi-region load prediction is solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202411887001.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In the prior art, traditional load prediction methods fail to fully consider the complex coupling characteristics and mutual influence between loads and meteorology in multiple regions, resulting in insufficient accuracy and robustness of load prediction, especially when power demand fluctuates between multiple regions.
A load feature extraction model based on long and short-term memory networks and spatial attention mechanisms is adopted, and a meteorological feature extraction model of the token to token model and Vision Transformer model is combined to establish a multi-scale LSTM feature extraction network and a multi-layer perceptron network, and a multi-region load prediction is carried out by integrating global and local features.
The accuracy and reliability of multi-region load prediction has been improved, the difficulties of lack of meteorological data and joint forecasting of multi-region loads have been overcome, and the prediction accuracy has been improved.
Smart Images

Figure CN119834215B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of load forecasting, and specifically, relates to a multi-region load forecasting method and system that integrates global enhancement and local attention features. Background Art
[0002] The reliability and stability of the power system are crucial for modern society, and load forecasting, as the basis of power system dispatching and planning, plays a vital role in achieving efficient energy utilization and providing a stable power supply. With the growth of energy demand and the continuous development of new energy technologies, load forecasting has received increasing attention.
[0003] However, in the prior art, traditional load forecasting methods usually focus on the load conditions of a single region, ignoring the complex coupling characteristics between multi-region loads and meteorology, as well as the correlation between multi-region loads. At the same time, the prior art does not perform multi-region interactive forecasting, resulting in the inability to fully consider the mutual influence between different regions, thereby restricting the load forecasting model from comprehensively capturing the dynamic changes of the overall power system, and further reducing the accuracy and robustness of load forecasting. This deficiency is manifested as the model being difficult to provide accurate forecasting results when dealing with power demand fluctuations between multi-regions, thus affecting the planning and dispatching decisions of the power system.
[0004] In the prior art, for the excavation of regionally lightning-sensitive loads in multi-region aggregation, the sample set data of all regions where lightning-sensitive loads are to be excavated are numbered; the long-term growing trend change component in the daily load curve of the currently to-be-excavated region is stripped; the fitting result function of the daily load curve after stripping the long-term growing trend change component in the load curve with respect to meteorological conditions and date type indication information is obtained through fitting; based on the average value of the difference between the predicted load and the actual load within the time interval affected by lightning, the lightning-sensitive load of the region's load on the current day is calculated; the lightning-sensitive loads of all regions are regionally aggregated to obtain all the lightning-sensitive load values of all regions where lightning-sensitive loads are to be excavated, and the fitting from relevant external meteorological and date information to the target day load curve provides a reference for analyzing the user's electricity consumption behavior characteristics, but the accuracy of predicting the load based on the fitting result is relatively low.
[0005] Therefore, in-depth research on the multi-region load forecasting problem and exploration of its mutual influence with multi-region meteorology are of great significance for improving the accuracy and reliability of load forecasting. Summary of the Invention
[0006] To solve the deficiencies in the prior art, the present invention provides a multi-region load forecasting method and system that integrates global enhancement and local attention features, and improves the accuracy of multi-region load forecasting results by fully excavating the load data characteristics under different regional conditions.
[0007] The present invention adopts the following technical solutions.
[0008] The present invention proposes a multi-region load forecasting method that fuses global enhancement and local attention features, including:
[0009] Obtain load data and multivariate meteorological data of multiple regions;
[0010] Based on the long short-term memory network and the spatial attention mechanism, establish a load feature extraction model to extract load features from the load data of different regions; based on the token to token model, the Vision Transformer model, and the Transformer model, establish a meteorological feature extraction model to extract meteorological features from the multivariate meteorological data of different regions;
[0011] Establish a multi-region load forecasting model, including: a multi-scale LSTM feature extraction network composed of multiple LSTM models with different scales, a multi-layer perceptron network, and a multi-region load forecasting network established based on the transformer model; among them, the multi-scale LSTM feature extraction network extracts load global features and load local features from the load features, and the multi-layer perceptron network fuses the load global features and load local features with the meteorological features to obtain the input data set of the multi-region load forecasting network; according to the input data set, the multi-region load forecasting network outputs the multi-region load forecasting result.
[0012] Preferably, the multivariate meteorological data includes: wind speed, wind direction, wind force, humidity, atmospheric pressure, temperature.
[0013] Preferably, the processing of the load data and multivariate meteorological data of multiple regions includes:
[0014] Delete outliers, fill in missing values, and perform time series smoothing on the data set after deleting outliers and filling in missing values;
[0015] Among them, the moving average method is used to perform time series smoothing on the data set after deleting outliers and filling in missing values.
[0016] Preferably, the load feature extraction model includes: an input layer, N spatial attention modules, N long short-term memory network modules, and an output layer;
[0017] The multi-region load data is processed by the first spatial attention module and then input into the first long short-term memory network module. The output data of the first long short-term memory network module is concatenated with the multi-region load data and then input into the second spatial attention module. The data processed by the second spatial attention module and the output data of the first long short-term memory network module are together input into the second long short-term memory network module, and so on. The output data of the (N - 1)-th long short-term memory network module is concatenated with the multi-region load data and then input into the N-th spatial attention module. The data processed by the N-th spatial attention module and the output data of the (N - 1)-th long short-term memory network module are together input into the N-th long short-term memory network module; the data output by the N-th long short-term memory network module is the multi-region load feature; the multi-region load feature output by the N-th long short-term memory network module is transmitted to the output layer.
[0018] Preferably, the structures of each spatial attention module are the same and each includes: a load data input unit, a linear mapping layer, and an activation function unit;
[0019] The multi-region load data is input into the linear mapping layer after passing through the load data input unit. The linear mapping layer includes three linear mapping units. The multi-region load data respectively obtains three weighted matrices through the three linear mapping units, satisfying the following relational expressions:
[0020] Q = W Q ·X
[0021] K = W K ·X
[0022] V = W V ·X
[0023] In the formula, Q is the weighted sum matrix of the first weight matrix W Q output by the first linear mapping unit and the multi-region load data X, defined as the query matrix, and K is the weighted sum matrix of the second weight matrix W K output by the second linear mapping unit and the multi-region load data X, defined as the key matrix, and V m is the weighted sum matrix of the third weight matrix W V output by the third linear mapping unit and the multi-region load data X, defined as the value matrix;
[0024] The activation function unit Softmax obtains the output data of the m-th spatial attention module according to the following relational expression:
[0025]
[0026] Where Attention(Q, K, V) is the output data of the spatial attention module, dk is the scaling factor, and the value of the scaling factor dk is the number of columns of the query matrix Q and the key matrix K.
[0027] Preferably, the meteorological feature extraction model includes: an input layer, N T2T modules, N Vision Transformer modules, N Transformer modules, and an output layer;
[0028] After passing through the first T2T module, the multi-source meteorological data is first patched and then input into the first Vision Transformer module. The data output by the first Vision Transformer module is transmitted to the first Transformer module; the data output by the first Transformer module and the multi-source meteorological data are passed through the second T2T module and then patched and input into the second Vision Transformer module. The data output by the second Vision Transformer module and the data output by the first Transformer module are transmitted to the second Transformer module together; and so on. The data output by the (N - 1)th Transformer module and the multi-source meteorological data are passed through the Nth T2T module and then patched and input into the Nth Vision Transformer module. The data output by the Nth Vision Transformer module and the data output by the (N - 1)th Transformer module are transmitted to the Nth Transformer module together; the meteorological features output by the Nth Transformer module (Transformer block N) are transmitted to the output layer.
[0029] Preferably, each T2T module includes: f processing units and f - 1 conversion units;
[0030] In the conversion unit, the i-th element in the sequence is updated using the weighted sum of the i-th element and other elements in the sequence. After all elements are updated, a new sequence is obtained;
[0031] Among them, the weights used for weighted summation satisfy the following relational expression:
[0032] α[i, j] = exp(score[i, j]) / Sum_k(exp(score[i, j]))
[0033] In the formula, α[i, j] is the weight between the i-th element and the j-th element, score[i, j] is the correlation degree between the i-th element and the j-th element, and Sum_k() is the summation.
[0034] Preferably, the multi-scale LSTM feature extraction network extracts the load global feature and the load local feature from the load feature, including:
[0035] Processing the load feature and the meteorological feature into the form of a confusion matrix;
[0036] The multi-scale LSTM feature extraction network includes N LSTM models with different scales; the first LSTM model takes the load feature in the form of a confusion matrix as the input and outputs the first local feature of the load; the second LSTM model takes the first local feature of the load as the input and outputs the second local feature of the load; and so on, the Nth LSTM model takes the (N-1)th local feature of the load as the input and outputs the Nth local feature and the load global feature.
[0037] Preferably, the multi-layer perceptron network fuses the load global feature and the load local feature with the meteorological feature to obtain the input data set of the multi-region load prediction network, including:
[0038] The multi-layer perceptron network includes N multi-layer perceptron layers;
[0039] In the first multi-layer perceptron layer, feature fusion is performed on the first local feature of the load to obtain the first local fusion feature, and the first local fusion feature matches the dimension of the meteorological feature in the form of a confusion matrix; in the second multi-layer perceptron layer, feature fusion is performed on the second local feature of the load to obtain the second local fusion feature, and the second local fusion feature matches the dimension of the meteorological feature in the form of a confusion matrix; and so on, in the Nth multi-layer perceptron layer, feature fusion is respectively performed on the Nth local feature of the load to obtain the Nth local fusion feature, and feature fusion is performed on the load global feature to obtain the global fusion feature. The Nth local fusion feature and the global fusion feature both match the dimension of the meteorological feature in the form of a confusion matrix; the first local fusion feature, the second local fusion feature,..., the Nth local fusion feature and the global fusion feature constitute the input data set of the multi-region load prediction network.
[0040] Preferably, according to the input data set, the multi-region load prediction network outputs the multi-region load prediction result, including:
[0041] The multi-region load prediction network includes N transformer models. The meteorological feature input data and the first local fusion feature are transmitted to the first transformer model together. The data output by the first transformer model and the second local fusion feature are transmitted to the second transformer model together, and so on. The data output by the (N-1)th transformer model and the Nth local fusion feature, the global fusion feature are transmitted to the Nth transformer model together, and the Nth transformer model outputs the multi-region load prediction result.
[0042] The present invention also provides a multi - region load forecasting system integrating global enhancement and local attention features, including:
[0043] A data processing module, configured to obtain load data and multivariate meteorological data of multiple regions; based on a long - short - term memory network and a spatial attention mechanism, establish a load feature extraction model to extract load features from the load data of different regions; based on a token - to - token model, a Vision Transformer model, and a Transformer model, establish a meteorological feature extraction model to extract meteorological features from the multivariate meteorological data of different regions;
[0044] A load forecasting module, configured to establish a multi - region load forecasting model, including: a multi - scale LSTM feature extraction network composed of multiple LSTM models with different scales, a multi - layer perceptron network, and a multi - region load forecasting network established based on a transformer model; wherein, the multi - scale LSTM feature extraction network extracts load global features and load local features from the load features, the multi - layer perceptron network fuses the load global features and load local features with the meteorological features to obtain an input data set for the multi - region load forecasting network; according to the input data set, the multi - region load forecasting network outputs a multi - region load forecasting result.
[0045] The present invention also relates to a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method.
[0046] The present invention also relates to a computer - readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method are implemented.
[0047] The beneficial effects of the present invention are at least as follows compared with the prior art. For numerical weather prediction (NWP) data such as temperature, humidity, wind speed, and atmospheric pressure, the method proposed by the present invention uses an NWP processing model based on T2T and Vision Transformer to extract features from meteorological data. The global and local features of multi - region loads extracted by a multi - layer perceptron MLP are fused with NWP features and input into the model for training, constructing a multi - region load forecasting method integrating global enhancement and local attention features, overcoming the problems of lack of meteorological data and difficulty in jointly forecasting multi - region loads, and improving the forecasting accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flowchart of a multi - region load forecasting method integrating global enhancement and local attention features proposed by the present invention;
[0049] Figure 2 The structural diagram of the load feature extraction model in the embodiment of the present invention;
[0050] Figure 3 The structural diagram of the NWP data feature extraction model in the embodiment of the present invention;
[0051] Figure 4 The flowchart of the MLP feature fusion model in the embodiment of the present invention;
[0052] Figure 5 The comparison diagram between the model load test curve and the actual value in the embodiment of the present invention. Detailed implementation manners
[0053] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0054] The present invention proposes a multi-region load forecasting method that fuses global enhancement and local attention features, including:
[0055] Step 1: Obtain the load data and multi-source meteorological data of multiple regions,
[0056] In the embodiment, as Figure 1 shown, the input layer Input layer includes the load data of Region 1, the load data of Region 2,..., the symbol data of Region M, and also includes the multi-source meteorological data of Region 1, the multi-source meteorological data of Region 2,..., the multi-source meteorological data of Region M; the multi-source meteorological data is from NWP; where M is the number of regions;
[0057] Specifically, preprocess the load data and NWP data of multiple regions to obtain initial input data, including:
[0058] 1), Delete outliers:
[0059] There may be outliers in the load data and NWP data. Therefore, after detecting the outliers in the collected data, delete the outliers; in the embodiment, the methods for detecting outliers include but are not limited to: the 3σ method; specifically as follows: use the load data and NWP data of each region to form a data set, calculate the mean μ and standard deviation σ of the data set; select three times the standard deviation 3σ as the threshold; when the data sample x satisfies |x - μ| > 3σ, then the data sample x is determined to be an outlier.
[0060] Among them, the calculation formulas for the mean μ and the standard deviation σ are as follows:
[0061]
[0062] In the formula, x i is the i-th data sample in the dataset, and n is the total number of data samples.
[0063] 2) Missing value filling:
[0064] Use the mean filling method for missing values, specifically as follows: Calculate the mean μ′ of the non-missing data in the dataset, and use the mean μ′ of the non-missing data to fill all missing values.
[0065] Among them, the calculation formula for the mean μ′ of the non-missing data is as follows:
[0066]
[0067] In the formula, n′ is the total number of non-missing data samples, and x j is the j-th non-missing data sample.
[0068] 3) Perform time series smoothing on the dataset after removing outliers and filling missing values to eliminate noise;
[0069] Use the moving average method to perform time series smoothing on the dataset after removing outliers and filling missing values to eliminate noise, specifically as follows:
[0070] Set the sliding window width k, where k is an odd number greater than 1; the size of the sliding window determines the degree of smoothing. The larger the sliding window, the more obvious the smoothing effect, but it may lead to over-smoothing of the data;
[0071] Process the data samples into a time series data sequence according to the time series relationship of the sampling moments; when the center data point of the sliding window is aligned with the sampling moment t, calculate the average value of the data within the sliding window as the moving average value at the sampling moment t; among them, when the sliding window is at the head or tail of the time series data sequence, there is a situation where the number of data samples within the sliding window is less than k, then calculate the moving average value based on the actual existing data samples and the number of data samples;
[0072] The calculation formula for the moving average value at the sampling moment t is as follows:
[0073]
[0074] In the formula, MA t is the moving average value at the sampling moment t, and x i is the i-th data sample within the sliding window;
[0075] When the number of data samples within the sliding window is less than k, the calculation formula for the moving average at sampling time t is as follows:
[0076]
[0077] Where k1 and k2 are the numbers of data samples before and after sampling time t within the sliding window respectively, and k1 + k2 + 1 ≤ k;
[0078] The moving averages at each sampling time in the time series data sequence are calculated in turn, and data smoothing and noise removal are achieved by a method that reflects the long-term trend.
[0079] Step 2: Based on the long short-term memory network (LSTM) and the spatial attention mechanism, establish a load feature extraction model to extract load features from load data in different regions; based on the token to token model, the Vision Transformer model, and the Transformer model, establish a meteorological feature extraction model to extract meteorological features from multivariate meteorological data in different regions.
[0080] Figure 1 Among them, in the processing layer, using the spatial attention module Space attention, the hidden layer, and the long short-term memory network modules LSTM Block1, LSTM Block1,..., LSTM BlockN, establish a load feature extraction model, and use the T2T module and the Vision Transformer modules Transformer Block1, Transformer Block2,..., Transformer Block N; among them, the long short-term memory network modules LSTMBlock1, LSTM Block1,..., LSTM Block n are connected in series in turn, and the output of LSTM Block n is the output of the load feature extraction model; the Vision Transformer modules Transformer Block1, TransformerBlock2,..., Transformer Block N are connected in series in turn, and the output of Transformer Block N is the output of the meteorological feature extraction model; among them, the numbers of the spatial attention module Space attention, the hidden layer, the long short-term memory network module, the T2T module, and the Vision Transformer module are all N;
[0081] Specifically, Step 2 includes:
[0082] Step 2.1: Based on the long short-term memory network (LSTM) and spatial attention mechanism, establish a load feature extraction model to extract load features from the load data in different regions;
[0083] The specific model structure diagram is as Figure 2 shown. The load data of Region 1, Region 2, Region 3,..., Region M in the input layer are respectively input into the corresponding spatial attention layers; N spatial attention layers and N long short-term memory network modules LSTM block constitute the high-dimensional feature extraction layer; the multi-region load data are input into the first long short-term memory network module LSTM Block1 after being processed by the first spatial attention layer. The output data of the first long short-term memory network module are concatenated with the multi-region load data and then input into the second spatial attention layer. The data processed by the second spatial attention layer and the output data of the first long short-term memory network module are input into the second long short-term memory network module together, and so on. The output data of the N-1th long short-term memory network module are concatenated with the multi-region load data and then input into the Nth spatial attention layer. The data processed by the Nth spatial attention layer and the output data of the N-1th long short-term memory network module are input into the Nth long short-term memory network module together; the data output by the Nth long short-term memory network module are the multi-region load feature data;
[0084] In the present invention, the output data of each spatial attention module include: the spatial dependence relationship of the multi-region load data and the dynamic weights of the load data of each region; the spatial dependence relationship of the multi-region load data is a kind of local attention feature. Each spatial attention module effectively models the potential associations between different regions, and can dynamically allocate the weights of the load in different regions, highlighting the contributions of important regions and suppressing irrelevant or redundant information at the same time;
[0085] Each long short-term memory network module is used to combine the time dependence relationship and spatial dependence relationship of the multi-region load data; at the same time, the output data of the n-1th long short-term memory network module are the input data of the nth long short-term memory network module. Through the series connection of multiple LSTM modules, each module not only retains the time dependence information, but also gradually integrates the multi-region global features extracted by the spatial attention module; the output data of the n-1th long short-term memory network module and the data concatenated with the multi-region load data obtain the global enhanced features in the nth spatial attention module. The global enhanced features are combined with the output data of the n-1th long short-term memory network module and then input into the nth long short-term memory network module. This design allows the load features of each region to not only fuse the time features of this region, but also comprehensively consider the influence of the features of other regions globally through the iterative update of each layer.
[0086] The multi-region load characteristics are used as the input data of the high-dimensional feature fusion layer, and the high-dimensional feature fusion layer includes multiple merge models; the load feature extraction model further includes an output layer, and the data output by the high-dimensional feature fusion layer is transmitted to the output layer, that is, the load feature Output-L.
[0087] Each space attention module has the same structure, which includes: a load data input unit, a linear mapping layer, and an activation function unit Softmax; the space attention module performs weighted summation on the load data of different regions to strengthen the spatial correlation and mine the local and global features of the multi-region load; the processing process of the m-th space attention module for the multi-region load data is as follows:
[0088] The multi-region load data is input to the linear mapping layer after passing through the load data input unit. The linear mapping layer includes 3 linear mapping units. The multi-region load data passes through the 3 linear mapping units respectively to obtain 3 weighted matrices, satisfying the following relational expressions:
[0089] Q = W Q ·X
[0090] K = W K ·X
[0091] V = W V ·X
[0092] In the formula, Q is the weighted sum matrix of the first weight matrix W Q output by the first linear mapping unit and the multi-region load data X, defined as the query matrix, and K is the weighted sum matrix of the second weight matrix W K output by the second linear mapping unit and the multi-region load data X, defined as the key matrix, and V m is the weighted sum matrix of the third weight matrix W V output by the third linear mapping unit and the multi-region load data X, defined as the value matrix;
[0093] The activation function unit Softmax obtains the output data of the m-th space attention module according to the following relational expression:
[0094]
[0095] Where Attention(Q, K, V) is the output data of the spatial attention module, and dk is the scaling factor;
[0096] QK T is the similarity matrix of the query matrix Q and the key matrix K. To prevent the similarity from being too large, the ratio of the similarity matrix QK T to the scaling factor dk is normalized by the Softmax function, thereby transforming the similarity matrix into an attention weight matrix. The weighted sum of the attention weight matrix and the value matrix V is used as the output data of the spatial attention module; among them, the value of the scaling factor dk is the number of columns of the query matrix Q and the key matrix K.
[0097] In the spatial domain, through the spatial attention mechanism, essentially a matrix is re-encoded into a matrix after considering global information through matrix operations, while considering both the global and local characteristics of multi-region load data, thereby strengthening the spatial correlation of the data. At the same time, as a recurrent neural network, LSTM mines the characteristics of the historical load data of each region and captures the change rules of the load at different time scales. The core idea of LSTM is to control the flow of information by introducing gating mechanisms, including the forget gate, input gate, and output gate. These gating mechanisms allow the LSTM network to effectively capture long-term dependencies and make it suitable for processing time series data of the load.
[0098] The load feature extraction model established in the present invention is a high-dimensional feature extraction module. The high-dimensional feature extraction module contains N networks of spatial attention and LSTM. Each network performs spatial and temporal feature extraction operations to mine the local and global characteristics of multi-region load data and more comprehensively capture the information at different time steps. The extracted high-dimensional feature sequence is used as the input of the spatial attention and LSTM network at the next time step, and so on. Finally, the output under N sampling points is calculated. This makes the high-dimensional features of multi-region load extracted by the model pay more attention to the temporal correlation of time series data, ensuring a more comprehensive understanding and extraction of the temporal features in the data, and enabling the model to better adapt to complex regional load data patterns.
[0099] Step 2.2, establish a meteorological feature extraction model based on the token to token model, Vision Transformer model, and Transformer model to extract meteorological features from multi-source meteorological data in different regions;
[0100] The multi-source meteorological data includes but is not limited to: wind velocity, wind direction, wind force, humidity, barometric pressure, temperature;
[0101] The specific model structure diagram is as follows Figure 3 shown. The multi-source meteorological data in the input layer (Input layer) is simultaneously input into the model; after passing through the first T2T module (token to token module), the multi-source meteorological data is first embedded into patches and then input into the first Vision Transformer module. The data output by the first Vision Transformer module is transmitted to the first Transformer module (Transformer block 1); the data output by the first Transformer module and the multi-source meteorological data are first embedded into patches after passing through the second T2T module and then input into the second Vision Transformer module. The data output by the second Vision Transformer module and the data output by the first Transformer module are transmitted to the second Transformer module (Transformer block 2) together; and so on. The data output by the (N - 1)-th Transformer module and the multi-source meteorological data are first embedded into patches after passing through the N-th T2T module and then input into the N-th Vision Transformer module. The data output by the N-th Vision Transformer module and the data output by the (N - 1)-th Transformer module are transmitted to the N-th Transformer module (Transformer block N) together; the data output by the N-th Transformer module (Transformer block N) is transmitted to the output layer (Output layer), that is, the meteorological feature (OutputT).
[0102] To fully exploit the hidden features in the NWP data of multiple regions, this section constructs a meteorological data feature processing model based on token to token and Vision Transformer to extract features from the multi-source meteorological data collected from multiple regions.
[0103] Each T2T module includes: f processing units (Token 1, Token 2, ……, Token f) and f - 1 transformation units (T2T transformer); the Tokens-to-Token module aims to overcome the limitations of simple tokenization in ViT. It gradually structures the image into tokens and models the local structural information, thereby gradually reducing the length of the tokens. Each T2T process has two steps: recombination and soft splitting. Given the T sequence from the previous Transformer layer, it will be transformed by the self-attention block, satisfying the following relationship:
[0104] T' = MLP(MSA(T))
[0105] Wherein, T' is the transformed sequence, T is the sequence, MSA() is the multi-layer normalized multi-head self-attention operation, and MLP() is the multi-layer self-attention operation;
[0106] Reshape T' in the spatial dimension into an image, satisfying the following relational expression:
[0107] I = Reshape(T')
[0108] Wherein, Reshape() is the image reshaping operation, and I is the reconstructed image of T' in the spatial dimension;
[0109] After obtaining the reconstructed image I, perform soft segmentation on it, model the local structure information, and reduce the length of the tokens. Specifically, to avoid information loss when generating Tokens from the reconstructed image, it is segmented into overlapping image patches. Therefore, each image patch is related to the surrounding image patches to establish a prior, that is, there should be a stronger correlation between the surrounding Tokens. The tokens in each segmented image patch are concatenated into one token, so that local information can be aggregated from the surrounding pixels and patches. When performing soft segmentation, the size of each image patch is k×k, there are s overlaps and p paddings on the image, where k - s is similar to the stride in the convolution operation. Therefore, for the reconstructed image i2r with size h×ω×c, the output token length after soft segmentation is:
[0110]
[0111] Wherein, l0 is the output token length after soft segmentation;
[0112] The size of each segmented patch is k×k×c; all the image patches in the spatial dimension are tiled into tokens. After soft segmentation, the output tokens are provided to the next T2T process.
[0113] The T2T module gradually reduces the length of the tokens and transforms the spatial structure of the image by iterating the above reconstruction and soft segmentation. The iterative process of the T2T module satisfies the following relational expression:
[0114] T i ' = MLP(MSA(T i ))
[0115] I i = Reshape(T i ')
[0116] T i+1 = SS(I i), i = 1....(n - 1)
[0117] In the formula, T i ′ is the transformed sequence in the i-th iteration, and T i is the sequence in the i-th iteration, I i is the reconstructed image in the i-th iteration, SS is the image transformation sequence operation, and T i+1 is the sequence in the (i + 1)-th iteration, and n is the number of iterations;
[0118] For the input image I0, soft segmentation is first applied to segment it into tokens. After the final iteration, the output token T of the T2T module n has a fixed length, so the backbone of T2T-vit can be on T n to model the global relationship.
[0119] Inside the T2T model, the Transformer structure is adopted, and the core mechanism is the self-attention mechanism. In this process, each Token will interact with other Tokens in the sequence, and according to their correlation, perform weighted summation to obtain a new Token representation. The specific formula is as follows:
[0120] y[i] = Sum_j(α[i,j] * X[j])
[0121] In the formula, y[i] is the i-th Token, X[j] is the j-th Token in the sequence, α[i,j] is the weight between the i-th Token and the j-th Token, and Sum_j() is the weighted summation; among them, the calculation formula of the weight is:
[0122] α[i,j] = exp(score[i,j]) / Sum_k(exp(score[i,j]))
[0123] In the formula, score[i,j] is the correlation degree between the i-th Token and the j-th Token, and Sum_k() is the summation;
[0124] In each layer of the Transformer, this process will be repeated. After obtaining the new representations of all Tokens in the final layer through the calculation of multiple layers of the Transformer, an unfolding operation is performed to map these representations back to each meteorological scalar again. The new representation is an N-dimensional vector, which is re-expanded into the Token sequences of the original temperature, humidity, wind speed, etc. through a specific mapping relationship. Finally, based on these high-dimensional features, subsequent decoding operations are performed. This process not only ensures that the correlation between meteorological variables in each dimension is fully considered, but also extracts richer features.
[0125] Step 3: Establish a multi-region load forecasting model, including: a multi-scale LSTM feature extraction network composed of multiple LSTM models with different scales, a multi-layer perceptron network, and a multi-region load forecasting network established based on the Transformer model; among them, the multi-scale LSTM feature extraction network extracts the load global feature and the load local feature from the load features, and the multi-layer perceptron network fuses the load global feature and the load local feature with the meteorological features to obtain the input data set of the multi-region load forecasting network; according to the input data set, the multi-region load forecasting network outputs the multi-region load forecasting result.
[0126] Specifically, Step 3 includes:
[0127] Step 3.1: Process the load features and meteorological features into the form of a confusion matrix;
[0128] In the embodiment, the high-dimensional features extracted from the load data and the high-dimensional features extracted from the meteorological data are reshaped into a shape suitable for the confusion model.
[0129] Step 3.2: Use multiple LSTM models with different scales to form a multi-scale LSTM feature extraction network; according to the load features in the form of a confusion matrix, the multi-scale LSTM feature extraction network outputs the global feature and the local feature of the multi-region load;
[0130] As Figure 4 shown, the load feature input data in the form of a confusion matrix is transmitted to the LSTM out layer of the multi-scale LSTM feature extraction network, and the multi-scale LSTM feature extraction network extracts the global feature and the local feature from the load feature input data; among them, the multi-scale LSTM feature extraction network includes N LSTM models with different scales, the load feature in the form of a confusion matrix is transmitted to LSTM block 1, LSTM block 1 outputs the first local feature of the load, the first local feature of the load is transmitted to LSTM block 2, LSTM block 2 outputs the second local feature of the load, and so on, LSTM block N outputs the Nth local feature of the load and the load global feature;
[0131] Step 3.3: Use a multi-layer perceptron to perform dimension matching and feature fusion on the meteorological features in the form of a confusion matrix with the global feature and the local feature of the multi-region load, so as to obtain the local fusion feature and the global fusion feature after fusion to form the input data set of the multi-region load forecasting network;
[0132] As Figure 4As shown, the multi-layer perceptron network Mixed layer includes N multi-layer perceptron layers Multilayerperceptron layer. In the first multi-layer perceptron layer, feature fusion is performed on the first local feature of the load to obtain the first local fusion feature, and the first local fusion feature matches the dimension of the meteorological feature in the form of a confusion matrix; in the second multi-layer perceptron layer, feature fusion is performed on the second local feature of the load to obtain the second local fusion feature, and the second local fusion feature matches the dimension of the meteorological feature in the form of a confusion matrix; and so on. In the Nth multi-layer perceptron layer, feature fusion is respectively performed on the Nth local feature of the load to obtain the Nth local fusion feature and on the global feature of the load to obtain the global fusion feature. Both the Nth local fusion feature and the global fusion feature match the dimension of the meteorological feature in the form of a confusion matrix. The first local fusion feature, the second local fusion feature, …, the Nth local fusion feature, and the global fusion feature constitute the input data set of the multi-region load prediction network.
[0133] In the embodiment, the multi-layer perceptron MLP learns complex mapping relationships through multiple levels of non-linear transformations, can perform dimensional transformations on the local and global features extracted from the loads in multiple regions, and can thus match the features of meteorological data. First, the input layer receives the original feature vector as input. Assume that the input layer has d neurons, and the input vector is x = [x1, x2, ..., x d . Secondly, since the MLP contains one or more hidden layers, each hidden layer contains several neurons. Each neuron is connected to all neurons in the previous layer and has weights and biases. The input weighted sum of the i-th neuron in the l-th layer is shown as follows:
[0134]
[0135] In the formula, is the input weighted sum of the i-th neuron in the input layer of the l-th layer, d (l-1) is the number of neurons in the input layer of the (l - 1)-th layer, is the weight between the j-th neuron in the input layer of the (l - 1)-th layer and the i-th neuron in the input layer of the l-th layer, is the bias of the i-th neuron in the input layer of the l-th layer, is the output of the j-th neuron in the input layer of the (l - 1)-th layer.
[0136] Then, non-linear transformation is performed through the activation function, and the specific operation is shown as follows:
[0137]
[0138] In the formula, is the non-linear transformation of the i-th neuron in the input layer of the l-th layer, and f() is the activation function;
[0139] Common activation functions include sigmoid, tanh, ReLU, etc.
[0140] The output layer is usually used to generate the final prediction result. Similar to the hidden layer, the weighted sum of each neuron in the output layer is as shown in the following formula:
[0141]
[0142] In the formula, is the weighted sum of the k-th neuron in the output layer of the L-th layer, d (L-1) is the number of neurons on the output layer of the (L - 1)-th layer, is the weight between the i-th neuron in the output layer of the (L - 1)-th layer and the k-th neuron in the output layer of the L-th layer, is the bias of the k-th neuron in the output layer of the L-th layer, is the output of the i-th neuron in the output layer of the (L - 1)-th layer.
[0143] Since the activation function of the output layer is for multi-classification problems, the output result is obtained through the softmax function. Finally, the output value processed by the MLP is merged with the last three layers in the transformer model. Considering both the global and local features of meteorological features and multi-region loads in the final Transformer model can effectively predict the load values in multiple regions under different meteorological conditions.
[0144] The present invention uses meteorological feature data to promote the improvement of the reliability of multi-region load prediction results. The key lies in realizing the interaction between meteorological data features and load data features through the MLP feature fusion model. It includes: the influence of meteorological features on regional loads is non-linear and heterogeneous, and different regions have different sensitivities to meteorological conditions (such as temperature, humidity, wind speed); introducing meteorological data into the Transformer model and jointly modeling it with load local fusion features, this design can capture the direct influence of meteorological conditions on the dynamic changes of loads; global information sharing: through the joint input of global features and meteorological features, capturing the influence of meteorological conditions on the overall load distribution of all regions, thereby enhancing the cooperation between regions.
[0145] Step 3.4, establish a multi-region load prediction network based on the transformer model,
[0146] Such as Figure 4As shown, the multi-region load forecasting network includes N Transformer models. The meteorological feature input data and the first local fusion feature are transmitted to the first Transformer model, Transformer block 1 together. The data output by the first Transformer model and the second local fusion feature are transmitted to the second Transformer model, Transformer block 2 together, and so on. The data output by the N - 1 Transformer model and the Nth local fusion feature and the global fusion feature are transmitted to the Nth Transformer model, Transformer block N together. The Nth Transformer model outputs the multi-region load forecasting result;
[0147] When the global and local features of the multi-region load extracted by the MLP and the NWP features are fused and input into the Transformer model for training and prediction, an early stopping strategy is set. When the loss of the validation set has not improved within 30 training epochs, stop training and restore the best model weights. Use the data of the training set to train the model, and at the same time use the data of the test set to validate the model.
[0148] According to the multi-region load forecasting result, use two common model evaluation metrics, Root Mean Square Error (RMSE) and Mean Absolute Error (MAE), to evaluate the reliability of the multi-region load forecasting model proposed in the present invention.
[0149] In the embodiment, the following three scenarios are set respectively for comparative analysis:
[0150] 1). Scenario 1: Prediction of the load when only global enhanced feature fusion is performed; Set the LSTM-out switch, use the LSTM model to perform a separate output of the load data, and this output is directly connected to the input of the classifier. The output of the model directly depends on the output of the last time step of the LSTM to perform the load prediction;
[0151] 2). Scenario 2: Prediction of the load when neither global enhancement nor local attention feature fusion is performed; Set the MLP-Mixer switch, and do not apply the MLP module after the extraction of meteorological and load features, so that the model does not perform transformation and fusion on the local and global features of meteorology and load, and then perform the load prediction of the model;
[0152] 3) Scenario 3: Prediction of load when fusing global enhancement and local attention features; Feature mining of load and meteorology through LSTM-Attention and Vit-Transformer, feature fusion processing through MLP, and then prediction through Transformer.
[0153] An ablation experiment was conducted on the prediction model by setting the local attention feature extraction of LSTM-Attention and the MLP fusion feature module. Without using the attention mechanism, a separate LSTM model was used to extract features from the load data. The model performance without local attention feature extraction was compared to evaluate the role of the local attention mechanism in the model. An MLP feature fusion mechanism was set. Whether to apply the MLP feature fusion module after mining global enhancement and local attention features enabled the model to perform deeper feature extraction and transformation on the features. To evaluate the performance difference of the model when using and not using the MLP feature fusion mechanism. The comparison curves of the prediction curves and the actual values for the three scenarios are as Figure 5 shown, Figure 5 where the abscissa is the sampling time Time / Sample, the ordinate is the value Y-value output by the model, the solid line is the predicted value Predicted value, and the dotted line is the real value Real value; The predicted load curve obtained by the model that mines the load and meteorology features through LSTM-Attention and Vit-Transformer, performs feature fusion processing through MLP, and then makes predictions through Transformer is closer to the real value. This is because this model combines the advantages of LSTM-Attention and Vision Transformer (ViT) and performs feature fusion processing through MLP. This architecture uses the LSTM-Attention model to capture long-term dependencies in time series data and focuses on key information through the attention mechanism. Then, the ViT model further extracts and processes the features of load and meteorology data, enhancing the model's ability to recognize complex patterns. After that, the MLP module transforms and fuses these features, improving the model's understanding of the relationships between features and the accuracy of predictions. Finally, the Transformer model uses its self-attention mechanism to effectively handle long-range dependency problems and make accurate load predictions. This multi-model fusion method not only improves the prediction accuracy and efficiency but also enhances the model's adaptability and generalization ability to data changes. Therefore, the prediction ability of the model in this paper is better than that of its ablation experiment model.
[0154] Table 1 Prediction result evaluation
[0155]
[0156] The comparison of the evaluation parameters of the prediction models for the three scenarios is shown in Table 1. It can be seen that, compared with the prediction models for Scenario 1 and Scenario 2, the multi-region load prediction model that fuses global enhancement and local attention features can effectively improve the accuracy of load prediction. The MAE values of the load prediction results of the method proposed in the present invention are reduced by 0.0051 and 0.0026 respectively; the RMSE are reduced by 0.0095 and 0.0094 respectively. Analyzing the above results, it can be known that both the MAE value and the RMSE value of the load prediction results have been reduced to a certain extent, verifying that the model proposed in the present invention has good applicability in load prediction.
[0157] The present invention also proposes a multi-region load prediction system that fuses global enhancement and local attention features, including:
[0158] A data processing module, configured to obtain load data and multivariate meteorological data of multiple regions; based on the long short-term memory network and the spatial attention mechanism, establish a load feature extraction model to extract load features from the load data of different regions; based on the token to token model, the Vision Transformer model, and the Transformer model, establish a meteorological feature extraction model to extract meteorological features from the multivariate meteorological data of different regions;
[0159] A load prediction module, configured to establish a multi-region load prediction model, including: a multi-scale LSTM feature extraction network composed of multiple different-scale LSTM models, a multi-layer perceptron network, and a multi-region load prediction network established based on the transformer model; wherein, the multi-scale LSTM feature extraction network extracts load global features and load local features from the load features, and the multi-layer perceptron network fuses the load global features and load local features with the meteorological features to obtain the input data set of the multi-region load prediction network; according to the input data set, the multi-region load prediction network outputs the multi-region load prediction result.
[0160] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0161] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0162] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0163] Computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A multi-region load forecasting method that fuses global enhancement and local attention features, characterized in that Including: Obtain load data and multi-source meteorological data of multiple regions; Based on the long short-term memory network and spatial attention mechanism, establish a load feature extraction model to extract load features from the load data of different regions. Among them, the load feature extraction model includes: an input layer, N spatial attention modules, N long short-term memory network modules, and an output layer. The multi-region load data is processed by the first spatial attention module and then input into the first long short-term memory network module. The output data of the first long short-term memory network module, the multi-region load data, and the data processed by the multi-region load data through the second spatial attention module are input into the second long short-term memory network module together, and so on. The output data of the N-1th long short-term memory network module, the multi-region load data, and the data processed by the multi-region load data through the Nth spatial attention module are input into the Nth long short-term memory network module together. The load features output by the Nth long short-term memory network module are transmitted to the output layer. Based on the token to token model, Vision Transformer model, and Transformer model, establish a meteorological feature extraction model to extract meteorological features from the multi-source meteorological data of different regions. Among them, the meteorological feature extraction model includes: an input layer, N T2T modules, N Vision Transformer modules, N Transformer modules, and an output layer. The multi-source meteorological data is first embedded with patches after passing through the first T2T module and then input into the first Vision Transformer module. The data output by the first Vision Transformer module is transmitted to the first Transformer module. The data output by the first Transformer module and the multi-source meteorological data are first embedded with patches after passing through the second T2T module and then input into the second Vision Transformer module. The data output by the second Vision Transformer module and the data output by the first Transformer module are transmitted to the second Transformer module together, and so on. The data output by the N-1th Transformer module and the multi-source meteorological data are first embedded with patches after passing through the Nth T2T module and then input into the Nth Vision Transformer module. The data output by the Nth Vision Transformer module and the data output by the N-1th Transformer module are transmitted to the Nth Transformer module together. The meteorological features output by the Nth Transformer module (Transformer block N) are transmitted to the output layer; A multi-region load forecasting model is established, including: a multi-scale LSTM feature extraction network composed of multiple LSTM models with different scales, a multi-layer perceptron network, and a multi-region load forecasting network established based on a transformer model; among them, the multi-scale LSTM feature extraction network extracts load global features and load local features from load features, and the multi-layer perceptron network fuses the load global features and load local features with meteorological features to obtain the input data set of the multi-region load forecasting network; according to the input data set, the multi-region load forecasting network outputs the multi-region load forecasting result.
2. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 1, characterized in that The multivariate meteorological data includes: wind speed, wind direction, wind force, humidity, atmospheric pressure, temperature.
3. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 2, characterized in that Processing the load data and multivariate meteorological data of multiple regions includes: Deleting outliers, filling in missing values, and performing time series smoothing on the data set after deleting outliers and filling in missing values; Among them, the moving average method is used to perform time series smoothing on the data set after deleting outliers and filling in missing values.
4. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 1, characterized in that The structure of each spatial attention module is the same, and each includes: a load data input unit, a linear mapping layer, and an activation function unit; The multi-region load data is input to the linear mapping layer after passing through the load data input unit. The linear mapping layer includes 3 linear mapping units. The multi-region load data passes through 3 linear mapping units to obtain 3 weighted matrices respectively, satisfying the following relational expression: In the formula, is the first weight matrix output by the first linear mapping unit and the weighted sum matrix of multi-region load data is defined as the query matrix. is the second weight matrix output by the second linear mapping unit and the weighted sum matrix of multi-region load data is defined as the key matrix. is the third weight matrix output by the third linear mapping unit and the weighted sum matrix of multi-region load data is defined as the value matrix. The activation function unit Softmax obtains the output data of the th spatial attention module according to the following relational expression: In the formula, is the output data of the spatial attention module, is the scaling factor, and the scaling factor has a value equal to the number of columns of the query matrix and the key matrix .
5. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 1, characterized in that Each T2T module includes: f processing units and f-1 conversion units; Within the conversion unit, the th element in the sequence is updated with the weighted sum of the th element and other elements in the sequence. After all elements are updated, a new sequence is obtained; Among them, the weights used for weighted summation satisfy the following relational expression: Wherein, is the weight between the -th element and the -th element, is the correlation degree between the -th element and the -th element, represents summation.
6. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 1, characterized in that The multi-scale LSTM feature extraction network extracts load global features and load local features from load features, including: Processing the load features and meteorological features into the form of a confusion matrix; The multi-scale LSTM feature extraction network includes N LSTM models with different scales; the first LSTM model takes the load features in the form of a confusion matrix as input and outputs the first local feature of the load; the second LSTM model takes the first local feature of the load as input and outputs the second local feature of the load; and so on, the Nth LSTM model takes the N-1th local feature of the load as input and outputs the Nth local feature and the load global feature.
7. The multi-region load forecasting method for fusing global enhancement and local attention features according to claim 6, characterized in that The multi-layer perceptron network fuses the load global features and load local features with meteorological features to obtain the input data set of the multi-region load forecasting network, including: The multi-layer perceptron network includes N multi-layer perceptron layers; In the first multi-layer perceptron layer, after the first local feature and the meteorological feature in the form of a confusion matrix are dimensionally matched, feature fusion is performed to obtain the first local fusion feature. In the second multi-layer perceptron layer, after the second local feature and the meteorological feature in the form of a confusion matrix are dimensionally matched, feature fusion is performed to obtain the second local fusion feature, and so on. In the Nth multi-layer perceptron layer, after the Nth local feature and the global feature are dimensionally matched with the meteorological feature in the form of a confusion matrix, feature fusion is performed to obtain the Nth local fusion feature and the global fusion feature respectively. The first local fusion feature, the second local fusion feature,..., the Nth local fusion feature and the global fusion feature constitute the input data set of the multi-region load prediction network.
8. The multi-region load prediction method for fusing global enhancement and local attention features according to claim 7, wherein According to the input data set, the multi-region load prediction network outputs multi-region load prediction results, including: The multi-region load prediction network includes N transformer models. The meteorological feature input data and the first local fusion feature are transmitted to the first transformer model together. The data output by the first transformer model and the second local fusion feature are transmitted to the second transformer model together, and so on. The data output by the N-1th transformer model and the Nth local fusion feature and the global fusion feature are transmitted to the Nth transformer model together, and the Nth transformer model outputs the multi-region load prediction results.
9. A multi-region load forecasting system integrating global enhancement and local attention features, which is used to implement the multi-region load forecasting method integrating global enhancement and local attention features according to any one of claims 1 to 8, characterized in that Including: A data processing module for obtaining load data and multivariate meteorological data of multiple regions; Based on the long short-term memory network and spatial attention mechanism, a load feature extraction model is established to extract load features from load data in different regions. Among them, the load feature extraction model includes: an input layer, N spatial attention modules, N long short-term memory network modules, and an output layer. The multi-region load data is processed by the first spatial attention module and then input into the first long short-term memory network module. The output data of the first long short-term memory network module, the multi-region load data, and the data of the multi-region load data processed by the second spatial attention module are input into the second long short-term memory network module together, and so on. The output data of the N-1th long short-term memory network module, the multi-region load data, and the data of the multi-region load data processed by the Nth spatial attention module are input into the Nth long short-term memory network module together. The load features output by the Nth long short-term memory network module are transmitted to the output layer. Based on the token to token model, Vision Transformer model, and Transformer model, a meteorological feature extraction model is established to extract meteorological features from multivariate meteorological data in different regions. Among them, the meteorological feature extraction model includes: an input layer, N T2T modules, N Vision Transformer modules, N Transformer modules, and an output layer. The multivariate meteorological data is first patched after passing through the first T2T module and then input into the first Vision Transformer module. The data output by the first Vision Transformer module is transmitted to the first Transformer module. The data output by the first Transformer module and the multivariate meteorological data are first patched after passing through the second T2T module and then input into the second Vision Transformer module. The data output by the second Vision Transformer module and the data output by the first Transformer module are transmitted to the second Transformer module together, and so on. The data output by the N-1th Transformer module and the multivariate meteorological data are first patched after passing through the Nth T2T module and then input into the Nth Vision Transformer module. The data output by the Nth Vision Transformer module and the data output by the N-1th Transformer module are transmitted to the Nth Transformer module together. The meteorological features output by the Nth Transformer module Transformer block N are transmitted to the output layer. A load forecasting module, which is used to establish a multi-region load forecasting model, including: a multi-scale LSTM feature extraction network composed of multiple LSTM models with different scales, a multi-layer perceptron network, and a multi-region load forecasting network established based on a transformer model; among them, the multi-scale LSTM feature extraction network extracts load global features and load local features from load characteristics, and the multi-layer perceptron network fuses the load global features and load local features with meteorological features to obtain the input data set of the multi-region load forecasting network; according to the input data set, the multi-region load forecasting network outputs the multi-region load forecasting result.
10. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Mid-term hour-level load probability prediction method based on time sequence fusion Transform model
CN115660161A
Medium and long term regional load prediction method based on cross-regional data fusion
CN116960962A