Renewable energy output prediction method and device, equipment and medium
Through a two-stage method, combining data augmentation, attention mechanism coding, comparative learning and LSTM network, the problem of low prediction accuracy of renewable energy in the existing technology is solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510075155.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-27
AI Technical Summary
The existing renewable energy prediction methods have problems such as coding sensitive to coding and easy loss of high-dimensional information, resulting in low prediction accuracy.
The two-stage method is adopted, firstly, the data representation learning performance is improved through the data enhancement method of segmented smoothing and segmented jittering. Then, the encoder model is obtained through attention mechanism encoding and comparative learning network training, and a multi-layer jump connection long and short-term memory network is constructed for prediction.
Overcome the problem of information missing during the encoding process, enhance high-dimensional information extraction, and significantly improve the accuracy and reliability of renewable energy output prediction.
Smart Images

Figure CN120218296A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of energy management, and particularly relates to a method, device, equipment and medium for predicting the output of renewable energy. Background Art
[0002] With the transformation of the world's energy structure towards clean and low-carbon, a new power system mainly based on renewable energy such as wind power will be gradually built. It is estimated that by 2030, the installed capacity ratio of renewable energy in China will reach 41%, and the power generation ratio will reach 22%, which are 1.7 times and 2.3 times that in 2020 respectively. The continuous increase in the penetration rate of renewable energy will significantly increase the uncertainty of power grid operation. As an effective means to cope with this uncertainty, accurate and reliable day-ahead prediction of renewable energy power is of great significance and is widely used in scenarios such as unit commitment, spinning reserve optimization, and maintenance planning, becoming the cornerstone of the safe and stable operation of the power system.
[0003] Traditional renewable energy prediction methods adopt an "end-to-end" prediction method, that is, after preprocessing the input wind power, photovoltaic, and meteorological data, it is directly predicted by time series or deep learning methods such as LSTM, autoregression, Transformer, and GNN. Although the "end-to-end" prediction method is intuitive and easy to operate, there are problems such as the prediction result being sensitive to coding and high-dimensional information being easily lost, ultimately resulting in low accuracy of renewable energy prediction.
[0004] In the field of renewable energy power prediction, there has not yet been a relatively mature time series prediction method, and data preprocessing is relatively simple and rough, making it difficult to correctly grasp the effective key information / high-dimensional information in historical data. Summary of the Invention
[0005] In view of the above problems, the present disclosure provides a method, device, equipment and medium for predicting the output of renewable energy.
[0006] In a first aspect, a method for predicting the output of renewable energy, the method includes:
[0007] Normalize the renewable energy data, and perform data augmentation on the processed data in different ways to obtain two sets of augmented data;
[0008] Process the two sets of augmented data respectively through the encoding of the attention mechanism to obtain two weight matrices;
[0009] Train based on the two weight matrices through a contrastive learning network to obtain a trained encoder model;
[0010] Using the LSTM network as the main body and the trained encoder model as the pre-stage of the LSTM network, a multi-layer skip-connection long short-term memory network for prediction is constructed; the real-time data of renewable energy is input into the multi-layer skip-connection long short-term memory network to obtain the renewable energy power prediction result.
[0011] Specifically, an encoder is obtained from the training result of the contrast learning network, and this encoder is used as the pre-stage of the LSTM network. After the data is encoded by the encoder, it is predicted by the LSTM module. After the training stage of the contrast learning network is completed, an encoder module is obtained, and the encoder module is connected to the LSTM network to obtain a renewable energy encoding-prediction module.
[0012] Furthermore, the normalization process of renewable energy data includes:
[0013] The renewable energy data is represented as a sequence of length m, denoted as x0 = [y T-m+ 1,..., y T , where T is the number of input data timestamps;
[0014] The historical weather forecast information, denoted as w s = [w T-m+1 ,..., w T ;
[0015] The weather forecast information for the next 24 hours is obtained from the meteorological department, denoted as w′ s = [w T+1 ,..., w T+24 ;
[0016] The time series of renewable energy data, historical weather forecast information, and weather forecast information for the next 24 hours are jointly used as the input data for representation learning data preprocessing and time series prediction, and the input data is normalized by reversible instance normalization; the output is the renewable energy output y = [y T+1 ,…, y T+24 for the next 24 hours, and the prediction task is to construct the prediction for each moment t.
[0017] Furthermore, data augmentation includes: random noise, random masking, scaling, and smoothing.
[0018] Furthermore, data augmentation includes: segment jitter and segment smoothing;
[0019] Each feature dimension in the multi-dimensional time series is regarded as an independent single-channel time series. If the sequence length is s and the given segment length is l, each dimension sequence is divided into non-overlapping segments; assume the number of augmented segments is M, randomly select M segments from all N segments and then augment them; assume represents the data of the i-th feature channel at the j-th moment, and the segment jitter and segment smoothing are expressed as follows:
[0020] Jitter:
[0021] Smoothing:
[0022] where ξ is the scale factor of the Gaussian variable, and the smoothing function Smooth() is expressed as follows:
[0023]
[0024] where k is the smoothing window length.
[0025] Furthermore, through the encoding of the attention mechanism, two groups of weight matrices are obtained by processing two groups of augmented data respectively, including:
[0026] Use the vanilla Transformer encoder as the encoder module, and use two layers as the optimal hyperparameter for the stacking layers of the encoder module to obtain the optimal temporal latent space information.
[0027] Use the Transformer encoder as the encoder module, denoted as f; dynamically allocate the attention weights of different input parts, and denote the input as where n is the length of the sequence, and d model is the dimension of the input vector. The attention mechanism first generates query, key, and value matrices by linearly transforming the input X:
[0028]
[0029] where, is the trainable weight matrix, and d k is the dimension of each attention head;
[0030] Calculate the dot product of the query and the key, and obtain the attention weights through the scaling factor and the softmax function:
[0031]
[0032] In the multi-head attention mechanism, h different attention heads are calculated simultaneously. The query, key, and value matrices of each head are linearly transformed with different weight matrices respectively, and then the attention scores are calculated independently and weighted and summed:
[0033]
[0034] where is the weight matrix of the i-th attention head;
[0035] The outputs of all attention heads are concatenated and integrated into the final output through a linear transformation:
[0036] MultiHead(Q, K, V) = Concat(head1,..., head h )W O
[0037] where is the final output weight matrix, and Concat represents concatenating the outputs of multiple attention heads.
[0038] Furthermore, through the contrastive learning network, based on two sets of weight matrices, a trained encoder model is obtained, including:
[0039] In contrastive learning, two identical encoders are used for the same input data to process two sets of weight networks respectively; the encoders share weights during training to ensure consistent transformation of the input; the data after segmented jittering is denoted as x1, and the data after segmented smoothing is denoted as x2, then the outputs are expressed as z1 = f(x1) and z2 = f(x2);
[0040] A predictor h is introduced, and the outputs after the predictor h are denoted as p1 = h(z1) and p2 = h(z2). The similarity of the two predicted outputs is maximized through the contrastive learning network, and the loss function is expressed as follows:
[0041]
[0042] where Sim() represents negative cosine similarity calculation, which is expressed as follows:
[0043]
[0044] The loss function symmetrically calculates the average similarity of two data under different branches, and its minimum optimization result is -1;
[0045] Combined with the stop-gradient loss function, it is expressed as follows:
[0046]
[0047] where r1 and r2 are the representations of two augmented data after passing through the encoder f, stopgrad() represents operating with the stop-gradient loss function, and q1 and q2 are the outputs of r1 and r2 after passing through the predictor h respectively.
[0048] Specifically, under the condition of ensuring the generality of the module, any predictor h is introduced. The predictor h is a prior art, and its specific structure is not part of the protection scope of this disclosure.
[0049] Furthermore, an LSTM network is used as the main body, and a trained encoder model is used as the pre-stage of the LSTM network to construct a multi-layer skip-connected long short-term memory network for prediction. The real-time data of renewable energy is input into the multi-layer skip-connected long short-term memory network to obtain the renewable energy power prediction result, including:
[0050] Using LSTM as the main part of the network, a multi-layer skip-connected long short-term memory network is constructed to learn the relationship between the pre-processed input time series data and the time series data of the next moment as the label, and predict the future output state information of renewable energy;
[0051] LSTM is used as the prediction network of the system. Its input data comes from the output of the encoder trained under the contrast learning framework. The output data of LSTM is the final renewable energy prediction result, including the future output state information of renewable energy;
[0052] The network includes long short-term memory units and skip connection structures. The forward propagation formula of long short-term memory is:
[0053]
[0054] i t = σ(W i · [h t-1 , x t + b i )
[0055] f t = σ(W f · [h t-1 , x t + b f )
[0056] o t = σ(W o [h t-1 , x t + b o )
[0057] Among them, C t is the cell state; is the intermediate variable of the cell state; i t is the input gate; f t is the forget gate; o t is the output gate; σ represents the Sigmoid activation function, acting as a gating signal; b f , b i , bo are the bias amounts of the forget gate, input gate, and output gate respectively; W f , W i , W c , W o are the weight matrices of the forget gate, input gate, cell state, and output gate respectively;
[0058] The forward propagation formula of the skip connection structure is:
[0059]
[0060] where, is the input feature vector of the k-th layer skip connection LSTM at the t-th sampling moment; is the hidden layer information of the k-th layer skip connection LSTM from the previous sampling moment at the t-th sampling moment; H k is the simplified transformation function of the k-th layer LSTM;
[0061] The hidden layer vector of the last layer of the skip connection long short-term memory network is used as the network output. After connecting with the fully connected layer and the Softmax activation layer, the prediction result is obtained. After reversibly instance denormalizing the prediction result, the predicted output of the future output of renewable energy is obtained, which is the prediction result.
[0062] In a second aspect, a renewable energy output prediction device includes: a data enhancement unit, a high-dimensional feature extraction unit, a contrast learning unit, and a prediction unit;
[0063] The data enhancement unit is configured to perform normalization processing on renewable energy data, and perform data enhancement on the processed data in different ways to obtain two sets of enhanced data;
[0064] The high-dimensional feature extraction unit is configured to process the two sets of enhanced data respectively through the encoding of the attention mechanism to obtain two sets of weight matrices;
[0065] The contrast learning unit is configured to train based on the two sets of weight matrices through a contrast learning network to obtain a trained encoder model;
[0066] The prediction unit is configured to use the LSTM network as the main body, use the trained encoder model as the pre-stage of the LSTM network, and construct a multi-layer skip connection long short-term memory network for prediction; input the real-time data of renewable energy into the multi-layer skip connection long short-term memory network to obtain the renewable energy power prediction result.
[0067] In a third aspect, an electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0068] A memory that stores a computer program;
[0069] A processor, when executing the computer program stored in the memory, implements the above-mentioned method for predicting the output of renewable energy.
[0070] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for predicting the output of renewable energy.
[0071] This disclosure at least includes the following beneficial effects:
[0072] The two-stage method of this disclosure is used for predicting the output of renewable energy, improving the comprehensive prediction accuracy. It has the following advantages:
[0073] In the first stage, the data augmentation methods of piecewise smoothing and piecewise jittering are used to improve the subsequent representation learning performance. After encoding the augmented data, asymmetric processing is adopted to prevent overfitting during the training of the encoder. Two data augmentation methods are used, and the same encoder is used for the augmented data, that is, asymmetric processing is adopted for the augmented and encoded data.
[0074] In the second stage, the encoder trained in the first stage is adopted, and a multi-layer skip connection network is used to predict the output of renewable energy, overcoming the information loss brought about during the encoding process, while enhancing the high-dimensional information, and finally improving the accuracy of the prediction result.
[0075] Other features and advantages of this disclosure will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing this disclosure. The objectives and other advantages of this disclosure can be achieved and obtained through the structures pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of this disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0077] Figure 1 It is a schematic flowchart of the prediction method according to the embodiment of this disclosure;
[0078] Figure 2 It is a schematic diagram of the principle of the prediction method according to the embodiment of this disclosure;
[0079] Figure 3 It is a schematic structural diagram of the prediction device according to the embodiment of this disclosure;
[0080] Figure 4 Schematic diagram of the structure of the electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0081] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0082] As Figure 1 shown, a renewable energy output prediction method, the method includes:
[0083] S101, performing normalization processing on renewable energy data, and performing data enhancement on the processed data in different ways to obtain two sets of enhanced data;
[0084] S102, respectively processing the two sets of enhanced data through the encoding of the attention mechanism to obtain two sets of weight matrices;
[0085] S103, training based on the two sets of weight matrices through a contrastive learning network to obtain a trained encoder model;
[0086] S104, using an LSTM network as the main body, using the trained encoder model as the front stage of the LSTM network, and constructing a multi-layer skip-connected long short-term memory network for prediction; inputting the real-time renewable energy data into the multi-layer skip-connected long short-term memory network to obtain the renewable energy power prediction result.
[0087] Specifically in implementation, it is introduced as follows:
[0088] The two-stage time series prediction can more deeply understand the time information between clue facts at different timestamps and the structural information between concurrent clue facts by constructing the first stage and the second stage, so as to improve the accuracy and reliability of the prediction. Generally, in the first stage, by preprocessing the system, extracting features, and training the model, historical information / high-dimensional information related to the prediction target is found. In the second stage: on the basis of the first stage, the second stage further processes these candidate information and uses the time information and structural information in the time series to make a more refined prediction.
[0089] As Figure 2 shown, the definition and processing method of renewable energy data includes the data definition of renewable energy prediction, taking reversible instance normalization to process the input renewable energy data, and performing data reconstruction on the processed data.
[0090] An encoding method based on the attention mechanism, which includes inputting a dataset after data augmentation, adopting a Transformer as an encoder, and designing an encoder module for the representation learning of renewable energy time-series data.
[0091] A high-dimensional feature extraction method for representation learning, which includes obtaining their latent space representations through the encoder module respectively, and adopting an asymmetric subsequent processing method to design a representation learning framework for renewable energy time-series data.
[0092] Predict the processed high-dimensional information, make predictions based on the obtained representation learning data preprocessing results, retain the trained encoder model for the prediction training in the second stage, including constructing a skip connection long short-term memory network, and finally obtain the renewable energy power prediction results.
[0093] Data definition for renewable energy prediction, which adopts invertible instance normalization to process the input renewable energy data and reconstruct the processed data.
[0094] The day-ahead wind power interval prediction problem can be formulated as follows: given a nominal confidence level of 100(1 - β)%, the renewable energy data is represented as a sequence of length m, denoted as x0 = [y T-m+1 ,..., y T , where T is the number of input data timestamps; historical weather forecast information, denoted as w s = [w T-m+1 ,..., w T ; at the same time, obtain the weather forecast information for the next 24 hours from the meteorological department, denoted as w′ s = [w T+1 ,..., w T+24 , and use the above time series as the input for representation learning data preprocessing and time series prediction. Take the renewable energy output y = [y T+1 ,..., y T+24 for the next 24 hours as the output, and the prediction task is to construct a prediction for each moment t. Invertible instance normalization has learnable affine transformations and also has normalization and denormalization methods. By normalizing the input data, it is a simple and effective method to significantly reduce the distribution shift problem between the training times and the test time series data.
[0095] Data augmentation is the first step in the preprocessing of representation learning data. Common time series data augmentation methods include random noise, random masking, scaling, and smoothing. These augmentation methods are widely used in time series representation learning, but they all have some drawbacks. The globally added noise easily contaminates the overall information of the original time series; the overall scaling changes the scale of the original sequence, but does not enhance the information about the periodic trends and other aspects of the time series at all; the random masking method adds random masks to the time series, but in slowly changing sequences, it is easy to infer the masked values from the values of the previous and subsequent timestamps without a high-level understanding of the entire sequence, which deviates from the goal of learning the important abstract representation of the entire time series; time series smoothing removes the high-frequency noise part in the data and retains the more important low-frequency part in the time series, but the smoothing window is a key hyperparameter that is difficult to grasp, and if selected improperly, it is easy to cause the loss of important information in the original time series.
[0096] For the dataset composed of renewable energy data, historical weather forecast information, and future 24-hour weather forecast information, it is used as the input data. At the same time, for the existing data missing and other problems, two coding methods are combined to enhance the data segment by segment with segment jitter and segment smoothing. In representation learning, the segment jitter and segment smoothing methods are used to enhance the original sequence to obtain a pair of positive representations of the original sequence. Segment jitter adds noise information to the original sequence, while segment smoothing filters out the high-frequency noise in the sequence and retains the main information of the data. The contrastive enhancement method is suitable for contrastive learning.
[0097] Specifically, each feature dimension in the multi-dimensional time series is regarded as an independent single-channel time series. If the sequence length is s and the given segment length is l, then each dimension sequence can be divided into non-overlapping segments. Assuming the number of enhanced segments is M, it means that M segments need to be randomly selected from all N segments and then enhanced. Assuming represents the data of the i-th feature channel at the j-th moment, the defined jitter and smoothing enhancement methods are:
[0098] Jitter:
[0099] Smoothing:
[0100] where ξ is the proportionality factor of the Gaussian variable, and the definition of the smoothing function Smooth() is:
[0101]
[0102] where k is the smoothing window length.
[0103] Input the dataset after data augmentation, and use Transformer as the encoder to design an encoder module for the representation learning of renewable energy time series data.
[0104] Since the self-attention mechanism introduced by the Transformer model can effectively capture the long-range dependencies in the input dataset sequence information and allows the model to perform parallel computing, showing strong capabilities in the time series field. Use the vanilla Transformer encoder as the encoder module, and use two layers as the optimal hyperparameter for the stacking layer of the encoder module to obtain the optimal time series latent space information.
[0105] Take the Transformer encoder as the encoder module, denoted as f. The attention mechanism is the core element of the Transformer. By dynamically allocating the attention weights of different input parts, the network can focus on important subsets of the input or features, so as to more effectively capture the key information in the given sequence. Denote the input as where n is the length of the sequence, and d model is the dimension of the input vector. The attention mechanism first generates query (Query), key (Key), and value (Value) matrices by linearly transforming the input X:
[0106]
[0107] where is the trainable weight matrix, and d k is the dimension of each attention head.
[0108] Next, calculate the dot product of the query and the key, and obtain the attention weights through the scaling factor and the softmax function:
[0109]
[0110] In the multi-head attention mechanism, h different attention heads are calculated simultaneously. The query, key, and value matrices of each head are linearly transformed with different weight matrices, and then the attention scores are calculated independently and weighted and summed:
[0111]
[0112] where is the weight matrix of the i-th attention head.
[0113] Connect the outputs of all attention heads and integrate them into the final output through a linear transformation:
[0114] MultiHead(Q,K,V) = Concat(head1,...,head h )W O
[0115] where is the final output weight matrix, and Concat means concatenating the outputs of multiple attention heads.
[0116] After passing through the encoder module, their latent space representations are obtained respectively, and an asymmetric subsequent processing method is adopted to design a representation learning framework for renewable energy time series data.
[0117] The representation learning framework consists of a contrastive learning framework and an encoder-decoder framework. When renewable energy prediction data is input, a pair of positive sample pairs required for contrastive learning are obtained through piecewise smoothing and piecewise jittering, and then their latent space representations are obtained respectively after passing through the encoder module f. Contrastive learning uses two identical encoders for the same input data, and processes two different enhanced versions of the same data respectively. The encoders share weights during the training process to ensure consistent transformation of the input. Denote the data after piecewise jittering as x1 and the data after piecewise smoothing as x2, then the outputs are expressed as z1 = f(x1) and z2 = f(x2).
[0118] At the same time, a predictor module h is introduced, and the outputs after passing through the predictor are denoted as p1 = h(z1) and p2 = h(z2). The similarity of the two prediction outputs is maximized through the contrastive learning network, and the loss function is defined as:
[0119]
[0120] where Sim() represents the negative cosine similarity calculation, defined as follows
[0121]
[0122] The loss function symmetrically calculates the average similarity of two data under different branches, and its minimum optimization result is -1.
[0123] At the same time, in order to prevent the network output from collapsing to a constant, the contrastive learning network introduces a stop gradient operation. During the backpropagation process, this mechanism stops the gradient flow of one branch to ensure that only meaningful gradients are used for parameter updates. Further, the loss function can be modified as:
[0124]
[0125] where r1 and r2 are the representations of the two augmented data after passing through the encoder f, and q1 and q2 are the outputs of r1 and r2 after passing through the predictor h respectively.
[0126] Perform prediction based on the obtained preprocessing results of representation learning data, and retain the trained encoder model for the prediction training in the second stage.
[0127] In the second stage, use LSTM as the main part of the network to construct a multi-layer skip connection network, learn the relationship between the preprocessed input time series data and the time series data of the next moment as the label, and predict the future output state information of renewable energy. The network includes long short-term memory units and skip connection structures. The forward propagation formula of the long short-term memory is as follows:
[0128]
[0129] i t = σ(W i · [h t-1 , x t + b i )
[0130] f t = σ(W f · [h t-1 , x t + b f )
[0131] o t = σ(W o [h t-1 , x t + b o )
[0132] where C t is the cell state; is the intermediate variable of the cell state; i t is the input gate; f t is the forget gate; o t is the output gate; σ represents the Sigmoid activation function and acts as a gating signal; b f , b i , b o are the bias amounts of the forget gate, input gate, and output gate respectively; W f , W i , W c , W o are the weight matrices of the forget gate, input gate, cell state, and output gate respectively.
[0133] The forward propagation formula of the skip connection structure is as follows:
[0134]
[0135] where refers to the input feature vector of the k-th layer skip connection LSTM at the t sampling moment; refers to the hidden layer information of the k-th layer skip connection LSTM at the t sampling moment from the previous sampling moment; H k refers to the simplified transformation function of the k-th layer LSTM.
[0136] The hidden layer vector of the last layer of the skip connection long short-term memory network is used as the network output. After connecting with the fully connected layer and the Softmax activation layer, the prediction result is obtained. After reversibly instance-inverse normalizing the prediction result, the predicted output of the future output of renewable energy can be obtained.
[0137] The present disclosure overcomes the problem of high-dimensional information loss in end-to-end prediction. While adopting a two-stage strategy, it adopts the method of training the attention mechanism encoder, avoiding the subjective problem of directly encoding and then predicting, and at the same time can better extract key information from historical data. In the first stage, data augmentation methods of piecewise smoothing and piecewise jittering are adopted and combined with the contrastive learning framework, ensuring that the information of the original data does not get lost and avoiding the collapse during the process of training the encoder.
[0138] As Figure 3 shown, a renewable energy output prediction device includes: a data augmentation unit 301, a high-dimensional feature extraction unit 302, a contrastive learning unit 303, and a prediction unit 304;
[0139] The data augmentation unit is used to normalize the renewable energy data and perform data augmentation on the processed data in different ways to obtain two sets of augmented data;
[0140] The high-dimensional feature extraction unit is used to process the two sets of augmented data respectively through the encoding of the attention mechanism to obtain two sets of weight matrices;
[0141] The contrastive learning unit is used to train based on the two sets of weight matrices through the contrastive learning network to obtain a trained encoder model;
[0142] The prediction unit is used to use LSTM as the main part of the network, construct a multi-layer skip connection long short-term memory network, and output the prediction result of the renewable energy power.
[0143] As Figure 4 shown, the present disclosure provides an electronic device, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404;
[0144] The memory 403 stores a computer program;
[0145] The processor 401 is configured to implement the above-mentioned method when executing the computer program stored in the memory 403.
[0146] The present disclosure provides a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the above-mentioned method.
[0147] The computer-readable storage medium may be included in the device / apparatus described in the above embodiments; or may exist alone without being assembled into the device / apparatus. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0148] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, including but not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0149] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A renewable energy output prediction method, characterized in that: The method comprises: The renewable energy data is normalized and the processed data is enhanced in different ways to obtain two sets of enhanced data; Through the encoding of the attention mechanism, two sets of enhanced data are processed separately to obtain two sets of weight matrices; By contrasting the learning network, training is performed based on two sets of weight matrices to obtain a trained encoder model; The LSTM network is used as the main body, and the trained encoder model is used as the front stage of the LSTM network to construct a multi-layer skip connection long short time memory network for prediction; the real-time data of renewable energy is input into the multi-layer skip connection long short time memory network to obtain the renewable energy power prediction result.
2. A renewable energy output prediction method according to claim 1, characterized in that: Normalize renewable energy data, including: Renewable energy data is represented by a sequence of length m, denoted as x0 = [y T-m+1 ,...,y T ], where T is the number of input data timestamps; Historical weather forecast information, denoted as w s =[w T-m+1 ,...,w T ]; Get the weather forecast information for the next 24 hours from the meteorological department, denoted as w′ s =[w T+1 ,...,w T+24 ]; The renewable energy data, historical weather forecast information and the weather forecast information time series for the next 24 hours are used as the input data for representation learning data preprocessing and time series prediction. The input data is normalized by reversible instance normalization. The renewable energy output y = [y T+1 ,…,y T+24 ] as output, the prediction task is to construct a prediction for each time instant t.
3. A renewable energy output prediction method according to claim 1, characterized in that: Data augmentation, including: random noise, random masking, scaling, and smoothing.
4. A renewable energy output prediction method according to claim 1, characterized in that: Data enhancement, including: piecewise jitter and piecewise smoothing; Each feature dimension in the multidimensional time series is regarded as an independent single-channel time series. If the sequence length is s and the segment length is l, each dimensional sequence is divided into non-overlapping segments; assuming that the number of enhanced segments is M, randomly select M segments from all N segments and then enhance them; assuming Represents the data of the feature channel of the i-th dimension at time j, and the piecewise jitter and piecewise smoothing are expressed as follows: Jitter: smooth: Among them, ξ is the scale factor of the Gaussian variable, and the smoothing function Smooth() is expressed as follows: Where k is the smoothing window length.
5. A renewable energy output prediction method according to claim 1, characterized in that: Through the encoding of the attention mechanism, two sets of enhanced data are processed separately to obtain two sets of weight matrices, including: The vanilla Transformer encoder is used as the encoder module, and two layers are used as the optimal hyperparameter for the number of stacked layers of the encoder module to obtain the optimal temporal latent space information; The Transformer encoder is used as the encoder module, denoted as f; the attention weights of different input parts are dynamically allocated, and the input is denoted as Where n is the length of the sequence, d model The dimension of the input vector is x. The attention mechanism first transforms the input X through a linear transformation to generate the query, key, and value matrices: in, is the trainable weight matrix, d k is the dimension of each attention head; Compute the dot product of the query and the key and multiply by the scaling factor And the softmax function to obtain the attention weights: In the multi-head attention mechanism, h different attention heads are calculated at the same time. The query, key, and value matrices of each head are linearly transformed with different weight matrices, and then the attention scores are calculated independently and weighted summed: in is the weight matrix of the i-th attention head; The outputs of all attention heads are connected and integrated into the final output through a linear transformation: MultiHead(Q,K,V)=Concat(head1,...,head h )W O in is the final output weight matrix, and Concat means concatenating the outputs of multiple attention heads.
6. A renewable energy output prediction method according to claim 1, characterized in that: By contrasting the learning network, training is performed based on two sets of weight matrices to obtain a trained encoder model, including: Contrastive learning uses two identical encoders for the same input data, processing two sets of weight networks separately; the encoders share weights during training to ensure consistent conversion of the input; the data after segmented jitter is recorded as x1, and the data after segmented smoothing is recorded as x2, then the output is expressed as z1 = f(x1), z2 = f(x2); Introduce predictor h, and the output after predictor h is recorded as p1=h(z1), p2=h(z2). The similarity of the two predicted outputs is maximized through contrast learning network. The loss function is expressed as follows: Among them, Sim() represents the negative cosine similarity calculation, which is expressed as follows: The loss function symmetrically calculates the average similarity of two data in different branches, and its minimum optimization result is -1; Combined with the stop gradient loss function, it is expressed as follows: Among them, r1 and r2 are the representations of the two enhanced data after passing through the encoder f, stopgrad() means using the stop gradient loss function operation, and q1 and q2 are the outputs of r1 and r2 after passing through the predictor h.
7. A renewable energy output prediction method according to claim 1, characterized in that: The LSTM network is used as the main body, and the trained encoder model is used as the front stage of the LSTM network to construct a multi-layer skip connection long short time memory network for prediction; The real-time data of renewable energy is input into the multi-layer skip connection long short-term memory network to obtain the renewable energy power prediction results, including: Using LSTM as the main part of the network, a multi-layer skip connection long short time memory network is constructed to learn the relationship between the time series data input after preprocessing and the time series data as the label at the next moment, and predict the future output status information of renewable energy; LSTM is used as the prediction network of the system. Its input data comes from the encoder output trained under the contrastive learning framework. The LSTM output data is the final renewable energy prediction result, which includes the future output status information of renewable energy. The network contains long short-term memory units and skip connection structures, where the forward propagation formula of long short-term memory is: i t =σ(W i ·[h t-1 ,x t ]+b i ) f t =σ(W f ·[h t-1 ,x t ]+b f ) the t =σ(W o [h t-1 ,x t ]+b o ) Among them, C t is the cell state; is the intermediate variable of cell state; i t is the input gate; f t For the forget gate; t is the output gate; σ represents the Sigmoid activation function, which acts as a gating signal; b f , b i , b o are the biases of the forget gate, input gate, and output gate respectively; W f , W i , W c , W o They are the weight matrices of the forget gate, input gate, cell state, and output gate respectively; The forward propagation formula of the skip connection structure is: in, is the input feature vector of the k-th layer skip-connected LSTM at sampling time t; H is the hidden layer information of the k-th layer skip connection LSTM from the previous sampling time at sampling time t; k Simplify the transformation function for the k-th layer LSTM; The last hidden layer vector of the skip-connected long short-term memory network is used as the network output, and is connected with the fully connected layer and the Softmax activation layer to obtain the prediction result. The reversible instance of the prediction result is denormalized to obtain the prediction output of the future output of renewable energy.
8. A renewable energy output prediction device, characterized in that: include: Data enhancement unit, high-dimensional feature extraction unit, contrastive learning unit and prediction unit; A data enhancement unit is used to normalize the renewable energy data and enhance the processed data in different ways to obtain two sets of enhanced data; A high-dimensional feature extraction unit is used to process two sets of enhanced data respectively through the encoding of the attention mechanism to obtain two sets of weight matrices; A contrastive learning unit, used for training based on two sets of weight matrices through a contrastive learning network to obtain a trained encoder model; The prediction unit is used to adopt the LSTM network as the main body and the trained encoder model as the front stage of the LSTM network to construct a multi-layer skip connection long short time memory network for prediction; the real-time data of renewable energy is input into the multi-layer skip connection long short time memory network to obtain the renewable energy power prediction result.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; a memory storing a computer program; A processor is used to implement a renewable energy output prediction method according to any one of claims 1 to 7 when executing a computer program stored in a memory.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a renewable energy output prediction method according to any one of claims 1 to 7 is implemented.