Power supply power prediction method based on bidirectional GRU power supply in combination with attention mechanism
By combining attention mechanism and bidirectional GRU network to process heterogeneous data in distributed power systems, the problems of low model accuracy and prediction accuracy in the prior art are solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202411871129.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The prior art cannot effectively handle the heterogeneity of data in distributed power systems, resulting in low model accuracy, while linear prediction models are difficult to capture the complex behavior of the system, resulting in low prediction accuracy.
The power supply power prediction method based on bidirectional GRU combined with attention mechanism is adopted to process real-time data through one-hot encoding and feature weighting mechanisms, and the important features are extracted using multi-layer attention mechanisms, and the complex relationship between history and future data is captured through a bidirectional GRU network with codec structures.
The model's processing capability of heterogeneous data is improved, the context-connected capture capability of time-series data is enhanced, and the prediction accuracy of power supply is significantly improved.
Smart Images

Figure CN120033659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a power supply prediction method, and in particular to a power supply prediction method based on a bidirectional GRU power supply combined with an attention mechanism. Background Art
[0002] Distributed energy has the advantages of high energy utilization, environmental friendliness, and high grid flexibility, and is an important part of the future new power system; however, distributed power sources rely heavily on renewable energy (such as solar energy and wind energy), and their power generation capacity has uncontrollable characteristics such as intermittent and volatile, resulting in unstable power quality of distributed power sources. With the increasing dependence of modern society on power supply, power outages or power quality fluctuations may have serious impacts on key areas such as industrial production and data centers; therefore, the demand for stability and reliability of power quality of distributed energy systems is increasing.
[0003] To meet this challenge, power quality prediction technology, as one of the effective means to ensure the stability of power quality, can achieve efficient prediction of future power quality by analyzing the fluctuation characteristics of distributed energy, and reasonably allocate power resources based on the prediction results, thereby improving the stability of power quality by reducing the imbalance between power supply and demand. Therefore, studying efficient power quality prediction methods is not only an important way to improve the operating stability of distributed energy systems, but also a necessary means to ensure the normal operation of key infrastructure and reduce economic losses caused by power outages.
[0004] The existing technology usually optimizes and screens multi-dimensional time series data to obtain the correlation between each characteristic parameter set and wind power; based on the characteristic correlation, the best characteristic parameter set is selected after comparison as the model feature input into the linear model for solution, thereby predicting the power generation.
[0005] Although this prediction method can improve the prediction accuracy by optimizing and screening multi-dimensional input parameters, it still has the following defects:
[0006] 1. Power quality is affected by a variety of factors. These data are not only diverse in dimensions and complex in sources, but also have different correlations and time series between different types of data. Existing technologies are often unable to handle the heterogeneity of data, resulting in low model accuracy.
[0007] 2. The operating state of the power system is affected by many factors such as load changes, equipment failures, and external environment, and has obvious dynamic characteristics. The fluctuation of power quality is often not linear. Simple linear prediction models are difficult to capture the complex behavior of the system, resulting in low prediction accuracy.
[0008] The information disclosed in this background technology section is only intended to increase the understanding of the overall background of the application, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to ordinary technicians in this field. Summary of the invention
[0009] The purpose of the present invention is to overcome the shortcomings of the prior art that the model accuracy is low due to the inability to handle the heterogeneity of data, and the linear prediction model is difficult to capture the complex behavior of the system, resulting in low prediction accuracy. A method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism is provided that can handle the heterogeneity of data to improve the model accuracy, and capture the complex behavior of the system through a multi-layer attention mechanism, so that the prediction accuracy is higher.
[0010] To achieve the above objectives, the technical solution of the present invention is:
[0011] A method for predicting power supply power based on a bidirectional GRU power supply combined with an attention mechanism, the prediction method comprising:
[0012] S1. Collect real-time data, convert text data in the real-time data into numerical vectors through one-hot encoding, calculate the weight of each data feature in the numerical vector through the feature weight mechanism, and combine the weight and the numerical vector to calculate a weighted numerical vector;
[0013] S2. The multi-layer attention mechanism is used to extract important features from weighted numerical vectors at two different time scales, namely, day and hour, to obtain data features extracted by the multi-layer attention mechanism.
[0014] S3. Based on the data features extracted by the multi-layer attention mechanism and the numerical data in the real-time data, a bidirectional GRU network based on the encoding and decoding structure is used to predict the power supply power of distributed power supply.
[0015] In S1, the real-time data includes numerical data and text data, the text data includes historical environmental data and estimated future environmental data, the historical environmental data and estimated future environmental data include meteorological conditions, time-related features, and power company discount plans, and the text data is converted into a numerical vector matrix using one-hot encoding, including:
[0016]
[0017] In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X;
[0018] The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as:
[0019]
[0020] In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism;
[0021] The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as:
[0022]
[0023] In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively;
[0024] Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours.
[0025] In S2, For the historical data of day t, let Indicates T on the tth day h Hours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour;
[0026] Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as:
[0027]
[0028] In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ;
[0029] Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as:
[0030]
[0031] The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network with a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as:
[0032]
[0033] In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and They are the biases in the first and second layers of the multi-layer attention mechanism, T is the data from the previous T days, and Th is the data for the Th hours in a day.
[0034] In S3, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as:
[0035] z u =σ(x u Wz +h u-1 U z +b z );
[0036] In the above formula, z u The larger the value, the more state information of the previous time step is retained in the current state. σ(·) is the sigmoid function, x u is the input of the GRU module, h u-1 is the hidden state of the time step u-1 of the output sequence, which contains the information of the time step, where W z With b z are the update gate network weights and biases, U z is the weight matrix from hidden state to gate;
[0037] The reset gate is used to control the influence of the previous hidden state in calculating the candidate hidden state, which is expressed as:
[0038] r u =σ(W r ·x u +U r ·h u-1 +b r );
[0039] In the above formula, r u is the output of the re-made gate network, W r With b r Reset the gate network weights and biases, U r is the weight matrix from hidden state to reconstruction gate;
[0040] Based on the reset gate, the candidate state of GRU The calculation formula includes:
[0041]
[0042] In the above formula, W h represents the network weight after re-gate adjustment, U h is the weight matrix from the hidden state adjusted by the reconstruction gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combined with the update gate, the hidden state h u It can be expressed as:
[0043]
[0044] The input sequence of the bidirectional GRU network is simultaneously input into two GRU networks in a forward and reverse manner to calculate the hidden state, including:
[0045]
[0046] In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state
[0047] The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h 1 ,h 2 ,…,h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as:
[0048]
[0049] In the above formula, and Represent the network weights and biases in the fully connected layer respectively.
[0050] In S3, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network;
[0051] The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as:
[0052] c u =F(r t,u ,c u-1 );
[0053] In the above formula, C u is the hidden state of the time step u of the input sequence, r t,u Represents the input sequence r t In the uth row, F is the encoding function;
[0054] The hidden states at all times are summarized to obtain the intermediate semantic matrix C = (c 1 ,c 2 ,…,c U ), by combining the data features λ extracted by the multi-layer attention mechanism in S2 t With β t,h And the intermediate semantic vector C to obtain the hierarchical attention vector where c t,h Represents the historical hidden state of the Encoder network at the hth hour on the tth day;
[0055] The level attention vector a t and future environmental dataX f As input to the Decoder network to obtain the output sequence, including:
[0056] h u =g(h u-1 ,X f ,C);
[0057] In the above formula, h u-1 represents the hidden state of the time step u-1 of the output sequence, and g represents the decoding function.
[0058] A power prediction system based on a bidirectional GRU power supply combined with an attention mechanism, the system is used to execute the power prediction method based on a bidirectional GRU power supply combined with an attention mechanism as described above, specifically comprising: a data processing module, a multi-layer attention module and a prediction module;
[0059] The data processing module is used to collect real-time data, convert text data in the real-time data into a numerical vector through one-hot encoding, calculate the weight of each data feature in the numerical vector through a feature weight mechanism, and combine the weight and the numerical vector to calculate a weighted numerical vector;
[0060] The multi-layer attention module is used to extract important features in the weighted numerical vector at two different time scales of day and hour through the multi-layer attention mechanism, and obtain data features extracted by the multi-layer attention mechanism;
[0061] The prediction module is used to predict the power supply of distributed power sources based on data features extracted by a multi-layer attention mechanism and numerical data in real-time data, and adopts a bidirectional GRU network based on a codec structure.
[0062] In the data processing module, the real-time data includes numerical data and text data, the text data includes historical environmental data and estimated future environmental data, the historical environmental data and estimated future environmental data include meteorological conditions, time-related features and power company discount plans, and the text data is converted into a numerical vector matrix using one-hot encoding, including:
[0063]
[0064] In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X;
[0065] The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as:
[0066]
[0067] In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism;
[0068] The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as:
[0069]
[0070] In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively;
[0071] Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours;
[0072] In the multi-layer attention module, For the historical data of day t, let Indicates T on the tth day h Hours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour;
[0073] Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as:
[0074]
[0075] In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ;
[0076] Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as:
[0077]
[0078] The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network with a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as:
[0079]
[0080] In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and are the biases in the first and second layers of the multi-layer attention mechanism, T is the data from the previous T days, and Th is the data for the Th hours in a day;
[0081] In the prediction module, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as:
[0082] z u =σ(x u Wz +h u-1 U z +b z );
[0083] In the above formula, z u The larger the value, the more state information from the previous time step is retained in the current state. σ(·) is the sigmoid function, and x u is the input to the GRU module, and h u-1 is the hidden state at time step u-1 of the output sequence, which contains the information of the time step. Among them, W z and b z are the update gate network weights and biases respectively, and U z is the weight matrix from the hidden state to the gate;
[0084] The reset gate is used to control the influence of the previous hidden state when calculating the candidate hidden state, which is expressed as:
[0085] r u =σ(W r ·x u +U r ·h u-1 +b r );
[0086] In the above formula, r u is the output of the reset gate network, and W r and b r are the reset gate network weights and biases respectively, and U r is the weight matrix from the hidden state to the reset gate;
[0087] Based on the reset gate, the candidate state of GRU has the following calculation formula:
[0088]
[0089] In the above formula, W h represents the network weight adjusted by the reset gate, and U h is the weight matrix from the hidden state adjusted by the reset gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combining with the update gate, the hidden state h u can be expressed as:
[0090]
[0091] The input sequence of the bidirectional GRU network is input into two GRU networks simultaneously in forward and reverse directions to calculate the hidden state, including:
[0092]
[0093] In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state
[0094] The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h 1 ,h 2 ,…,h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as:
[0095]
[0096] In the above formula, and Represent the network weights and biases in the fully connected layer respectively;
[0097] In the predicted module, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network;
[0098] The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as:
[0099] c u =F(r t,u ,c u-1 );
[0100] In the above formula, C u is the hidden state of the time step u of the input sequence, r t,u Represents the input sequence r t In the uth row, F is the encoding function;
[0101] The hidden states at all times are summarized to obtain the intermediate semantic matrix C = (c 1 ,c 2 ,…,c U ), by combining the data features λ extracted by the multi-layer attention mechanism in the multi-layer attention module t With β t,h And the intermediate semantic vector C to obtain the hierarchical attention vector where c t,h Represents the historical hidden state of the Encoder network at the hth hour on the tth day;
[0102] The level attention vector a t and future environmental dataX f As input to the Decoder network to obtain the output sequence, including:
[0103] h u =g(h u-1 ,X f ,C);
[0104] In the above formula, h u-1 represents the hidden state of the time step u-1 of the output sequence, and g represents the decoding function.
[0105] A bidirectional GRU power supply prediction device based on an attention mechanism, characterized in that it includes a memory and a processor, wherein the memory is used to store computer program code and transmit the computer program code to the processor;
[0106] The processor is used to execute the aforementioned bidirectional GRU power supply prediction method based on the attention mechanism according to the instructions in the computer program code.
[0107] A computer program product includes a computer program, characterized in that the computer program is executed by a processor to implement the aforementioned bidirectional GRU power supply prediction method combined with an attention mechanism.
[0108] A computer storable medium having a computer program stored therein, characterized in that the computer program is executed by a processor to implement the aforementioned method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism.
[0109] Compared with the prior art, the present invention has the following beneficial effects:
[0110] 1. In a power prediction method based on a bidirectional GRU power supply combined with an attention mechanism, after collecting real-time data, the text data in the real-time data is converted into a numerical vector through unique hot encoding, and the weight of each data feature in the numerical vector is calculated through a feature weight mechanism, and the importance of the feature is measured by combining the weight and the numerical vector, and the key features are extracted from a large amount of heterogeneous data, and the noise data is cleaned, so as to improve the accuracy of subsequent modeling, thereby improving the accuracy of the model. Therefore, this design can effectively improve the accuracy of the model by combining the weight and the numerical vector to measure the importance of the feature.
[0111] 2. In the power prediction method based on bidirectional GRU power supply combined with attention mechanism of the present invention, multi-dimensional extraction of important features in input data is realized through multi-layer attention mechanism, the model's ability to capture the contextual connection of time-series related data is enhanced, and the prediction of distributed power supply power is realized by using bidirectional GRU network based on codec structure. Therefore, this design can realize the prediction of distributed power supply power through bidirectional GRU network combined with attention mechanism, effectively improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0112] Figure 1 is a flow chart of the method of the present invention.
[0113] Figure 2 It is a planning idea diagram of the method described in the present invention.
[0114] Figure 3 It is a diagram showing the predicted power supply effect of the method of the present invention and the comparative scheme in Example 3.
[0115] Figure 4 It is a structural diagram of the system described in the present invention.
[0116] Figure 5 It is a structural diagram of the device described in the present invention. DETAILED DESCRIPTION
[0117] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0118] Embodiment 1:
[0119] See also Figure 1 to Figure 2 , a power prediction method based on a bidirectional GRU power supply combined with an attention mechanism, the prediction method comprising:
[0120] S1. Collect real-time data, convert text data in the real-time data into numerical vectors through one-hot encoding, calculate the weight of each data feature in the numerical vector through the feature weight mechanism, and combine the weight and the numerical vector to calculate a weighted numerical vector;
[0121] S2. The multi-layer attention mechanism is used to extract important features from weighted numerical vectors at two different time scales, namely, day and hour, to obtain data features extracted by the multi-layer attention mechanism.
[0122] S3. Based on the data features extracted by the multi-layer attention mechanism and the numerical data in the real-time data, a bidirectional GRU network based on the encoding and decoding structure is used to predict the power supply power of distributed power supply.
[0123] In S1, the real-time data includes numerical data and text data, the text data includes historical environmental data and estimated future environmental data, the historical environmental data and estimated future environmental data include meteorological conditions, time-related features, and power company discount plans, and the text data is converted into a numerical vector matrix using one-hot encoding, including:
[0124]
[0125] In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X;
[0126] The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as:
[0127]
[0128] In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism;
[0129] The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as:
[0130]
[0131] In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively;
[0132] Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours.
[0133] In S2, For the historical data of day t, let Indicates T on the tth day h Hours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour;
[0134] Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as:
[0135]
[0136] In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ;
[0137] Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as:
[0138]
[0139] The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network with a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as:
[0140]
[0141] In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and They are the biases in the first and second layers of the multi-layer attention mechanism, T is the data from the previous T days, and Th is the data for the Th hours in a day.
[0142] In S3, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as:
[0143] z u =σ(x u W z +h u-1 U z +b z );
[0144] In the above formula, z u The larger the value, the more state information of the previous time step is retained in the current state. σ(·) is the sigmoid function, x u is the input of the GRU module, h u-1 is the hidden state of the time step u-1 of the output sequence, which contains the information of the time step, where W z With b z are the update gate network weights and biases, U z is the weight matrix from hidden state to gate;
[0145] The reset gate is used to control the influence of the previous hidden state in calculating the candidate hidden state, which is expressed as:
[0146] r u =σ(W r ·x u +U r ·h u-1 +b r );
[0147] In the above formula, r u is the output of the re-made gate network, W r With b r Reset the gate network weights and biases, U r is the weight matrix from hidden state to reconstruction gate;
[0148] Based on the reset gate, the candidate state of GRU The calculation formula includes:
[0149]
[0150] In the above formula, W h represents the network weight after re-gate adjustment, U his the weight matrix from the hidden state adjusted by the reconstruction gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combined with the update gate, the hidden state h u It can be expressed as:
[0151]
[0152] The input sequence of the bidirectional GRU network is simultaneously input into two GRU networks in a forward and reverse manner to calculate the hidden state, including:
[0153]
[0154] In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state
[0155] The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h 1 ,h 2 ,…,h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as:
[0156]
[0157] In the above formula, and Represent the network weights and biases in the fully connected layer respectively.
[0158] In S3, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network;
[0159] The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as:
[0160] c u =F(r t,u ,c u-1 );
[0161] In the above formula, C uis the hidden state at time step u of the input sequence, r t,u represents the u-th row in the input sequence r t where F is the encoding function;
[0162] The hidden states at all times are aggregated to obtain the intermediate semantic matrix C = (c 1 , c 2 , …, c U ). By combining the data features λ t and β t,h extracted by the multi-layer attention mechanism in S2 and the intermediate semantic vector C, the hierarchical attention vector is obtained where c t,h represents the historical hidden state at the t-th day and h-th hour in the Encoder network;
[0163] The hierarchical attention vector a t and the future environmental data X f are used as inputs to the Decoder network to obtain the output sequence, including:
[0164] h u = g(h u-1 , X f , C);
[0165] In the above formula, h u-1 represents the hidden state at time step u - 1 of the output sequence, and g represents the decoding function.
[0166] Example 2:
[0167] A power supply power prediction system based on bidirectional GRU combined with an attention mechanism, which is used to execute the power supply power prediction method based on bidirectional GRU combined with an attention mechanism as described in Example 1, and specifically includes: a data processing module, a multi-layer attention module, and a prediction module;
[0168] The data processing module is used to collect real-time data, convert the text data in the real-time data into a numerical vector through one-hot encoding, calculate the weight of each data feature in the numerical vector through a feature weight mechanism, and calculate a weighted numerical vector by combining the weight and the numerical vector;
[0169] In the data processing module, the real-time data includes numerical data and text data, and the text data includes historical environmental data and estimated future environmental data. The historical environmental data and estimated future environmental data include meteorological conditions, time-related features, and power company discount plans. Using one-hot encoding to convert the text data into a numerical vector matrix, including:
[0170]
[0171] In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X;
[0172] The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as:
[0173]
[0174] In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism;
[0175] The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as:
[0176]
[0177] In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively;
[0178] Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours;
[0179] The multi-layer attention module is used to extract important features in the weighted numerical vector at two different time scales of day and hour through the multi-layer attention mechanism, and obtain data features extracted by the multi-layer attention mechanism;
[0180] In the multi-layer attention module, For the historical data of day t, let Indicates T on the tth day hHours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour;
[0181] Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as:
[0182]
[0183] In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ;
[0184] Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as:
[0185]
[0186] The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network with a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as:
[0187]
[0188] In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and are the biases in the first and second layers of the multi-layer attention mechanism, T is the data from the previous T days, and Th is the data for the Th hours in a day;
[0189] The prediction module is used to predict the power supply of distributed power sources based on data features extracted by a multi-layer attention mechanism and numerical data in real-time data, and adopts a bidirectional GRU network based on a codec structure.
[0190] In the prediction module, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as:
[0191] z u =σ(x u W z +h u-1 U z +b z );
[0192] In the above formula, z u The larger the value, the more state information of the previous time step is retained in the current state. σ(·) is the sigmoid function, x u is the input of the GRU module, h u-1 is the hidden state of the time step u-1 of the output sequence, which contains the information of the time step, where W z With b z are the update gate network weights and biases, U z is the weight matrix from hidden state to gate;
[0193] The reset gate is used to control the influence of the previous hidden state in calculating the candidate hidden state, which is expressed as:
[0194] r u =σ(W r ·x u +U r ·h u-1 +b r );
[0195] In the above formula, r u is the output of the re-made gate network, W r With b r Reset the gate network weights and biases, U r is the weight matrix from hidden state to reconstruction gate;
[0196] Based on the reset gate, the candidate state of GRU The calculation formula includes:
[0197]
[0198] In the above formula, W h represents the network weight after re-gate adjustment, U h is the weight matrix from the hidden state adjusted by the reconstruction gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combined with the update gate, the hidden state h u It can be expressed as:
[0199]
[0200] The input sequence of the bidirectional GRU network is simultaneously input into two GRU networks in a forward and reverse manner to calculate the hidden state, including:
[0201]
[0202] In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state
[0203] The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h 1 ,h 2 ,…,h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as:
[0204]
[0205] In the above formula, and Represent the network weights and biases in the fully connected layer respectively;
[0206] In the predicted module, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network;
[0207] The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as:
[0208] c u =F(rt,u ,c u-1 );
[0209] In the above formula, C u is the hidden state of the time step u of the input sequence, r t,u Represents the input sequence r t In the uth row, F is the encoding function;
[0210] The hidden states at all times are summarized to obtain the intermediate semantic matrix C = (c 1 ,c 2 ,…,c U ), by combining the data features λ extracted by the multi-layer attention mechanism in the multi-layer attention module t With β t,h And the intermediate semantic vector C to obtain the hierarchical attention vector where c t,h Represents the historical hidden state of the Encoder network at the hth hour on the tth day;
[0211] The level attention vector a t and future environmental dataX f As input to the Decoder network to obtain the output sequence, including:
[0212] h u =g(h u-1 ,X f ,C);
[0213] In the above formula, h u-1 represents the hidden state of the time step u-1 of the output sequence, and g represents the decoding function.
[0214] Embodiment 3:
[0215] This embodiment simulates the scheme described in Example 1 and the comparative scheme to illustrate the effect of this design. The comparative scheme is a conventional power supply prediction scheme based on a long short-term memory network.
[0216] The simulation data of this embodiment is based on the ISO-NE dataset, which is a power grid operating organization in the New England region of the northeastern United States, responsible for the reliable operation of the regional power system, the management of the power market, and energy planning;
[0217] The comparison scheme directly merges the digitized historical environmental data, future environmental data and historical power supply data as input, and uses a typical long short-term memory network to predict the future power supply power. Figure 3 It can be seen that compared with the comparative scheme, the power value predicted by this design is closer to the actual value. Therefore, the prediction accuracy performance of the proposed scheme is better than that of the comparative scheme.
[0218] Embodiment 4:
[0219] See also Figure 4 , a bidirectional GRU power supply prediction device based on an attention mechanism, characterized in that it includes a memory and a processor, the memory is used to store computer program code and transmit the computer program code to the processor;
[0220] The processor is used to execute the bidirectional GRU power supply prediction method based on the attention mechanism as described in Example 1 according to the instructions in the computer program code.
[0221] See also Figure 5 , a bidirectional GRU power supply prediction device based on an attention mechanism, characterized in that it includes a memory and a processor, the memory is used to store computer program code and transmit the computer program code to the processor;
[0222] The processor is used to execute the bidirectional GRU power supply prediction method based on the attention mechanism as described in Example 1 according to the instructions in the computer program code.
[0223] A computer program product includes a computer program, characterized in that the computer program is executed by a processor to implement the bidirectional GRU power supply prediction method combined with an attention mechanism as described in Example 1.
[0224] A computer storable medium having a computer program stored therein, characterized in that the computer program is executed by a processor to predict the power supply power based on a bidirectional GRU power supply combined with an attention mechanism as described in Example 1.
[0225] The above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed by the present invention should be included in the protection scope recorded in the claims.
Claims
1. A method for predicting power supply based on bidirectional GRU power supply combined with attention mechanism, characterized by: The prediction method comprises: S1. Collect real-time data, convert text data in the real-time data into numerical vectors through one-hot encoding, calculate the weight of each data feature in the numerical vector through the feature weight mechanism, and combine the weight and the numerical vector to calculate a weighted numerical vector; S2. The multi-layer attention mechanism is used to extract important features from weighted numerical vectors at two different time scales, namely, day and hour, to obtain data features extracted by the multi-layer attention mechanism. S3. Based on the data features extracted by the multi-layer attention mechanism and the numerical data in the real-time data, a bidirectional GRU network based on the encoding and decoding structure is used to predict the power supply power of distributed power supply.
2. According to claim 1, a method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism, characterized in that: In S1, the real-time data includes numerical data and text data, the text data includes historical environmental data and estimated future environmental data, the historical environmental data and estimated future environmental data include meteorological conditions, time-related features, and power company discount plans, and the text data is converted into a numerical vector matrix using one-hot encoding, including: In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X; The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as: In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism; The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as: In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively; Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours.
3. According to claim 1, a method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism, characterized in that: In S2, For the historical data of day t, let Indicates T on the tth day h Hours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour; Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as: In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ; Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as: The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network of a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as: In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and are the biases in the first and second layers of the multi-layer attention mechanism, respectively. T is the data from the previous T days, and Th is the data for the Th hours in a day.
4. According to claim 3, a method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism is characterized in that: In S3, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as: z u =σ(x u W z +h u-1 U z +b z ); In the above formula, z u The larger the value, the more state information of the previous time step is retained in the current state. σ(·) is the sigmoid function, x u is the input of the GRU module, h u-1 is the hidden state of the time step u-1 of the output sequence, which contains the information of the time step, where W z With b z are the update gate network weights and biases, U z is the weight matrix from hidden state to gate; The reset gate is used to control the influence of the previous hidden state in calculating the candidate hidden state, which is expressed as: r u =σ(W r ·x u +U r ·h u-1 +b r ); In the above formula, r u is the output of the re-made gate network, W r With b r Reset the gate network weights and biases, U r is the weight matrix from hidden state to reconstruction gate; Based on the reset gate, the candidate state of GRU The calculation formula includes: In the above formula, W h represents the network weight after re-gate adjustment, U h is the weight matrix from the hidden state adjusted by the reconstruction gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combined with the update gate, the hidden state h u It can be expressed as: The input sequence of the bidirectional GRU network is simultaneously input into two GRU networks in a forward and reverse manner to calculate the hidden state, including: In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h1, h2, ..., h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as: In the above formula, and Represent the network weights and biases in the fully connected layer respectively.
5. According to claim 4, a method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism, characterized in that: In S3, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network; The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as: c u =F(r t,u ,c u-1 ); In the above formula, C u is the hidden state of the time step u of the input sequence, r t,u Represents the input sequence r t In the uth row, F is the encoding function; The hidden states at all times are summarized to obtain the intermediate semantic matrix C = (c1, c2, ..., c U ), by combining the data features λ extracted by the multi-layer attention mechanism in S2 t With β t,h And the intermediate semantic vector C to obtain the hierarchical attention vector where c t,h Represents the historical hidden state of the Encoder network at the hth hour on the tth day; The level attention vector a t and future environmental dataX f As input to the Decoder network to obtain the output sequence, including: h u =g(h u-1 ,X f ,C); In the above formula, h u-1 represents the hidden state of the time step u-1 of the output sequence, and g represents the decoding function.
6. A bidirectional GRU power supply prediction system based on attention mechanism, characterized in that: The system is used to execute the power prediction method based on the bidirectional GRU power supply combined with the attention mechanism as described in any one of claims 1 to 5, specifically comprising: a data processing module, a multi-layer attention module and a prediction module; The data processing module is used to collect real-time data, convert text data in the real-time data into a numerical vector through one-hot encoding, calculate the weight of each data feature in the numerical vector through a feature weight mechanism, and combine the weight and the numerical vector to calculate a weighted numerical vector; The multi-layer attention module is used to extract important features in the weighted numerical vector at two different time scales of day and hour through the multi-layer attention mechanism, and obtain data features extracted by the multi-layer attention mechanism; The prediction module is used to predict the power supply of distributed power sources based on data features extracted by a multi-layer attention mechanism and numerical data in real-time data, and adopts a bidirectional GRU network based on a codec structure.
7. According to claim 6, a bidirectional GRU power supply prediction system combined with an attention mechanism is characterized in that: In the data processing module, the real-time data includes numerical data and text data, the text data includes historical environmental data and estimated future environmental data, the historical environmental data and estimated future environmental data include meteorological conditions, time-related features and power company discount plans, and the text data is converted into a numerical vector matrix using one-hot encoding, including: In the above formula, the numerical vector matrix includes M rows, each row includes N features, is the rth row of X; The feature weight mechanism comprises two layers of fully connected networks, wherein the first layer of the feature weight mechanism uses an exponential linear unit as an activation function to process the negative part of the input vector, which is expressed as: In the above formula, is the output of the first layer of the neural network of the feature weight mechanism, W 1 is the weight of the first layer of the neural network of the feature weight mechanism, α is the proportional coefficient, and b 1 is the bias of the first layer of the neural network for the feature weight mechanism; The output of the activation function Input to the second layer of the neural network of the feature weight mechanism, and use the softmax function to calculate the weight of each element in the input X, expressed as: In the above formula, γ j is the weight of the jth element, is the output y of the second layer neural network of the feature weight mechanism 2 =W 2 y 1 +b 2 The jth element in W 2 With b 2 They represent the weight and bias of the second layer neural network of the feature weight mechanism respectively; Multiply the weight vector γ by the input feature X to obtain a weighted numerical vector Expressed as Where ⊙ represents the matrix dot multiplication operation; It contains the historical environmental data of the previous T days and the estimated environmental data for the next day, where each day includes T h hours; In the multi-layer attention module, For the historical data of day t, let Indicates T on the tth day h Hours of historical environmental data, Indicates T in the next day h An estimate of the environmental data for one hour; Calculate the historical data for day t With Future DataX f The Pearson correlation coefficient and the correlation coefficient of the data features to quantify the two days are expressed as: In the above formula, and Respectively With X f The element in row i and column j in and Respectively With X f The mean of the elements in the jth column of ; Based on the correlation coefficient, the softmax function is further combined to calculate the feature weight of the historical data on the tth day, expressed as: The historical data is combined with the data of the same time period in the future environmental data, expressed as Then it is input into a two-layer fully connected network of a multi-layer attention mechanism to obtain the correlation coefficient, where the first layer uses ELU as the activation function and the second layer uses softmax as the activation function, expressed as: In the above formula, β t,h is the correlation coefficient between the attention coefficient of the hth hour on the tth day and the hth hour in the future, is the output of the second layer of the neural network of the multi-layer attention mechanism (t*T h +h) elements, among which and are the network weights in the first and second layers of the multi-layer attention mechanism, and are the biases in the first and second layers of the multi-layer attention mechanism, T is the data from the previous T days, and Th is the data for the Th hours in a day; In the prediction module, the numerical data includes historical power supply data, and the bidirectional GRU network based on the codec structure is used to capture the complex relationship between historical data and future data. The bidirectional GRU network is composed of two GRUs, and the GRU is used to capture the dependency between context information in the input sequence through a gating mechanism. The GRU includes an update gate and a reset gate. The update gate is used to determine how much information from the previous hidden state needs to be retained in the current hidden state, which is expressed as: z u =σ(x u W z +h u-1 U z +b z ); In the above formula, z u The larger the value, the more state information of the previous time step is retained in the current state. σ(·) is the sigmoid function, x u is the input of the GRU module, h u-1 is the hidden state of the time step u-1 of the output sequence, which contains the information of the time step, where W z With b z are the update gate network weights and biases, U z is the weight matrix from hidden state to gate; The reset gate is used to control the influence of the previous hidden state in calculating the candidate hidden state, which is expressed as: r u =σ(W r ·x u +U r ·h u-1 +b r ); In the above formula, r u is the output of the re-made gate network, W r With b r Reset the gate network weights and biases, U r is the weight matrix from hidden state to reconstruction gate; Based on the reset gate, the candidate state of GRU The calculation formula includes: In the above formula, W h represents the network weight after re-gate adjustment, U h is the weight matrix from the hidden state adjusted by the reconstruction gate to the candidate hidden state, is the hyperbolic tangent activation function, which is used to calculate the candidate hidden state. Combined with the update gate, the hidden state h u It can be expressed as: The input sequence of the bidirectional GRU network is simultaneously input into two GRU networks in a forward and reverse manner to calculate the hidden state, including: In the above formula, and Represent the forward hidden state and the reverse hidden state respectively, and integrate the forward hidden state and the reverse hidden state to obtain the final hidden state The hidden states at all times are summarized to obtain the intermediate semantic matrix H = (h1, h2, ..., h U ), using the fully connected layer with Softplus activation function to convert hidden information into future T h Power supply forecast for one hour predict , expressed as: In the above formula, and Represent the network weights and biases in the fully connected layer respectively; In the predicted module, the bidirectional GRU network based on the codec structure includes an Encoder network and a Decoder network, and the Encoder network and the Decoder network are both composed of a bidirectional GRU network; The input of the Encoder network is the integration of historical environmental data and historical power supply data. r t Input into the Encoder network, the Encoder network encodes the input sequence into a context vector of fixed length, and its calculation process can be expressed as: c u =F(r t,u ,c u-1 ); In the above formula, C u is the hidden state of the time step u of the input sequence, r t,u Represents the input sequence r t In the uth row, F is the encoding function; The hidden states at all times are summarized to obtain the intermediate semantic matrix C = (c1, c2, ..., c U ), by combining the data features λ extracted by the multi-layer attention mechanism in the multi-layer attention module t With β t,h And the intermediate semantic vector C to obtain the hierarchical attention vector where c t,h Represents the historical hidden state of the Encoder network at the hth hour on the tth day; The level attention vector a t and future environmental dataX f As input to the Decoder network to obtain the output sequence, including: h u =g(h u-1 ,X f ,C); In the above formula, h u-1 represents the hidden state of the time step u-1 of the output sequence, and g represents the decoding function.
8. A power prediction device based on a bidirectional GRU power supply combined with an attention mechanism, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store computer program code and transmit the computer program code to the processor; The processor is used to execute the bidirectional GRU power supply prediction method based on the attention mechanism according to the instructions in the computer program code as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that The computer program is executed by a processor to implement the method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism as described in any one of claims 1 to 5.
10. A computer storable medium, wherein a computer program is stored in the computer storable medium, characterized in that: The computer program is executed by a processor as described in any one of claims 1 to 5 as a method for predicting power supply based on a bidirectional GRU power supply combined with an attention mechanism.
Citation Information
Patent Citations
Encrypted network traffic classification method based on space-time attention mechanism
CN115348215A
Two-way gating circulation unit power distribution network net power prediction method and device based on attention mechanism
CN118446363A
Text-to-speech synthesis method, device, computer apparatus, and non-volatile computer readable storage medium
WO2020147404A1
Attention weight calculation method and apparatus based on convolutional neural network, and device
WO2021068528A1