Short-term photovoltaic power generation power prediction method and related device
By simplifying the structure of the Transformer model, a photovoltaic power prediction model including the input layer, attention mechanism layer, addition layer, residual connection and layer normalization layer, forward propagation layer, full connection layer and output layer is proposed, which solves the problem of too long training time caused by excessive training parameters of the traditional model, and achieves more efficient training and more accurate prediction.
Patent Information
- Application Number
- CN202510156407.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional photovoltaic power prediction models have too many training parameters in short-sequence tasks, resulting in too long training time and unable to effectively improve prediction accuracy.
By simplifying the structure of the Transformer model, a photovoltaic power prediction model including the input layer, attention mechanism layer, addition layer, residual connection and layer normalization layer, forward propagation layer, fully connected layer and output layer is proposed, which reduces training parameters and shortens training time.
This method can effectively reduce training parameters, shorten training time, and improve the accuracy of photovoltaic power generation prediction.
Smart Images

Figure CN120090171A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of photovoltaic power generation prediction, and particularly to a short-term photovoltaic power generation prediction method and related devices. Background Technique
[0002] China has a vast territory and high levels of sunlight, providing unique geographical conditions for the field of photovoltaic power generation. Although the research and application of photovoltaic power generation in China started later than in foreign countries, with the continuous improvement of relevant national policies and core technology research, the installed capacity of photovoltaic power generation in China has now steadily increased and ranks first in the world. However, the large-scale grid connection of photovoltaic power generation remains a problem in China. An effective way to solve this problem is to effectively improve the prediction accuracy of photovoltaic power generation and the coordinated dispatching among new energy power generation systems, which helps the photovoltaic power generation industry in China to transform from the early distributed off-grid photovoltaic power generation mode to the large-scale grid-connected photovoltaic power generation mode.
[0003] Compared with traditional energy sources, photovoltaic power generation is affected by the alternation of day and night on the earth and has characteristics such as periodic intermittency and random volatility. The large-scale grid connection of photovoltaic power stations has a certain impact on the stable operation of the power grid. Predicting the photovoltaic output in advance and improving the predictability of the photovoltaic output can enable the power grid dispatching department to formulate dispatching plans in advance, ensure the stable operation of the power grid, and promote the consumption of photovoltaic power, which is crucial for maintaining the reliability of the photovoltaic power generation system and improving the power quality.
[0004] In the prediction task of photovoltaic power generation, due to the characteristics of the model structures of traditional LSTM and its series of variant models, although the phenomenon of gradient explosion during the propagation process can be reduced by introducing a gating mechanism or adding a residual structure, the training parameters of the model also increase exponentially. The traditional Transformer model performs well when applied to long-sequence tasks such as translation tasks, but when applied to short-sequence tasks, there are too many training parameters, resulting in too long training time. Summary of the Invention
[0005] The purpose of the present application is to provide a short-term photovoltaic power generation prediction method and related devices, which can reduce the training parameters and shorten the training time.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] In the first aspect, the present application provides a short-term photovoltaic power generation prediction method, and the short-term photovoltaic power generation prediction method includes:
[0008] Obtain the feature sequence data for each unit time period in the historical time period; the feature sequence data includes the values of each feature at each moment within the unit time period, the duration of the unit time period is less than the preset duration, and the features include irradiance, humidity, temperature, and wind speed;
[0009] Using the feature sequence data of each unit time period in the historical time period as input, and using the trained photovoltaic power prediction model to determine the power sequence data of the next unit time period in the historical time period; the power sequence data includes the values of the photovoltaic power at each moment within the unit time period;
[0010] Among them, the trained photovoltaic power prediction model is a model obtained by simplifying the structure of the Transformer model. The trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer, which are connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer.
[0011] In a second aspect, the present application provides a short-term photovoltaic power prediction device, and the short-term photovoltaic power prediction device includes:
[0012] A data acquisition module, configured to obtain the feature sequence data for each unit time period in the historical time period; the feature sequence data includes the values of each feature at each moment within the unit time period, the duration of the unit time period is less than the preset duration, and the features include irradiance, humidity, temperature, and wind speed;
[0013] A power prediction module, configured to use the feature sequence data of each unit time period in the historical time period as input, and use the trained photovoltaic power prediction model to determine the power sequence data of the next unit time period in the historical time period; the power sequence data includes the values of the photovoltaic power at each moment within the unit time period;
[0014] Among them, the trained photovoltaic power prediction model is a model obtained by simplifying the structure of the Transformer model. The trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer, which are connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer.
[0015] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-mentioned short-term photovoltaic power prediction method.
[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-mentioned short-term photovoltaic power prediction method.
[0017] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned short-term photovoltaic power prediction method.
[0018] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0019] The present application provides a short-term photovoltaic power prediction method and related devices. The trained photovoltaic power prediction model is a model obtained by simplifying the structure of the Transformer model. At this time, the trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer. Directly using the feature sequence data of each unit time period in the historical time period as input, and using the trained photovoltaic power prediction model to determine the power sequence data of the next unit time period in the historical time period. The present application simplifies the structure of the Transformer model to obtain a trained photovoltaic power prediction model, and uses this trained photovoltaic power prediction model to predict the photovoltaic power, which can reduce the training parameters and shorten the training time. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is an application environment diagram of a short-term photovoltaic power prediction method provided in Embodiment 1 of the present application.
[0022] Figure 2Schematic flowchart of a short-term photovoltaic power prediction method provided in Embodiment 1 of this application.
[0023] Figure 3 Schematic diagram of the network structure of the trained photovoltaic power prediction model provided in Embodiment 1 of this application.
[0024] Figure 4 Schematic flowchart of the calculation process of Gate-attention provided in Embodiment 1 of this application.
[0025] Figure 5 Schematic flowchart of the calculation process of Conv-attention provided in Embodiment 1 of this application.
[0026] Figure 6 Schematic diagram of the features in the calculation process of Conv-attention provided in Embodiment 1 of this application.
[0027] Figure 7 Schematic diagram of the comparison of loss values of different models provided in Embodiment 1 of this application.
[0028] Figure 8 Schematic diagram of the comparison of prediction results of different models in spring provided in Embodiment 1 of this application.
[0029] Figure 9 Schematic diagram of the comparison of prediction results of different models in summer provided in Embodiment 1 of this application.
[0030] Figure 10 Schematic diagram of the comparison of prediction results of different models in autumn provided in Embodiment 1 of this application.
[0031] Figure 11 Schematic diagram of the comparison of prediction results of different models in winter provided in Embodiment 1 of this application.
[0032] Figure 12 Schematic diagram of the functional modules of a short-term photovoltaic power prediction device provided in Embodiment 2 of this application.
[0033] Figure 13 Schematic diagram of the structure of a computer device provided in Embodiment 3 of this application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0035] Example 1
[0036] The short - term photovoltaic power generation prediction method provided by the embodiment of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send a prediction request to be processed (for requesting the prediction of the photovoltaic power station power) to the server. After receiving the prediction request to be processed, for the prediction request to be processed, the server obtains the feature sequence data of each unit time period in the historical time period. The feature sequence data includes the value of each feature at each moment within the unit time period. The duration of the unit time period is less than the preset duration, and the features include irradiance, humidity, temperature, and wind speed. Using the feature sequence data of each unit time period in the historical time period as the input, the trained photovoltaic power generation prediction model is used to determine the power sequence data of the next unit time period in the historical time period. The power sequence data includes the value of the photovoltaic power generation at each moment within the unit time period. The server can feedback the prediction result for the prediction request to the terminal.
[0037] In addition, in some embodiments, the short - term photovoltaic power generation prediction method can also be implemented independently by the server or the terminal. For example, the terminal can directly process the prediction request to be processed, or the server can obtain the prediction request to be processed from the data storage system and process the prediction request to be processed.
[0038] Among them, the terminal can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in - vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head - mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0039] In an exemplary embodiment, as Figure 2 shown, a short - term photovoltaic power generation prediction method is provided. This method is executed by a computer device, and specifically can be executed independently by a computer device such as a terminal or a server, or jointly executed by the terminal and the server. In the embodiment of the present application, taking this method applied to the Figure 1 server as an example for illustration, it includes the following steps.
[0040] Step S1: Obtain the feature sequence data for each unit time period in the historical time period; the feature sequence data includes the values of each feature at each moment within the unit time period, the duration of the unit time period is less than the preset duration, and the features include irradiance, humidity, temperature, and wind speed.
[0041] Step S2: Use the feature sequence data of each unit time period in the historical time period as input, and utilize the trained photovoltaic power prediction model to determine the power sequence data for the next unit time period in the historical time period; the power sequence data includes the values of the photovoltaic power at each moment within the unit time period.
[0042] As Figure 3 shown, the trained photovoltaic power prediction model is a model obtained by simplifying the structure of the Transformer model. The trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer, which are connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer.
[0043] Implementing the above Steps S1 to S2, in this embodiment, the structure of the Transformer model is simplified to obtain the trained photovoltaic power prediction model. At this time, the trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer, which are connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer. Subsequently, use the feature sequence data of each unit time period in the historical time period as input, and utilize the trained photovoltaic power prediction model to determine the power sequence data for the next unit time period in the historical time period to complete the prediction of the photovoltaic power. Since the structure of the trained photovoltaic power prediction model is simple and there are few training parameters, the training time can be shortened, the training efficiency can be improved, and it has been experimentally shown that the prediction accuracy is also improved.
[0044] In this embodiment, the number of unit time periods included in the historical time period can be set according to user requirements. Generally, the last unit time period included in the historical time period is the current unit time period, and the next unit time period in the historical time period is the next unit time period of the current unit time period. The duration of the unit time period can be 15 minutes, so as to achieve the purpose of short-term photovoltaic power prediction.
[0045] The calculation process of the attention layer in the traditional Transformer model is as follows:
[0046] This attention layer first maps the input data into three vectors through three linear mapping matrices, and the calculation formula is as follows:
[0047]
[0048] X ∈ R N×T (2)
[0049]
[0050] In the above formula, Q, K, and V are the query vector, key vector, and value vector respectively; X is the input data, which represents the feature data of a unit time period in this embodiment; W q 、W k 、W v are the weight matrices for mapping the query vector, key vector, and value vector respectively; N is the dimension of the features in the input data; D q 、D k 、D v are the dimensions of the mapped query vector, key vector, and value vector respectively; T is the dimension of the time moment in the input sequence.
[0051] Through the above formula (1), the query vector Q, key vector K, and value vector V can be calculated. Subsequently, the attention scores are calculated, and the calculation formula is as follows:
[0052]
[0053] In formula (4), h n is the attention score; Softmax represents an operation function used to summarize all calculations; d k is the dimension of the key vector K.
[0054] To improve the prediction accuracy, in this embodiment, after simplifying the structure of the Transformer model, the attention layer in the improved Transformer model is improved. Specifically, two improvement methods are provided as follows:
[0055] (1) The first improvement method
[0056] This embodiment proposes an improved Transformer model without feature processing. This improved Transformer model is a variant model of the Transformer model obtained by adding a gating mechanism and redesigning the attention scoring mechanism (which can be called Gate-attention or Gate-ATT). It is a short-order feature prediction model, such as Figure 4As shown, in order to make the model more suitable for short sequence prediction tasks, a gating mechanism is introduced in the attention mechanism layer to reduce the computational complexity of self-Attention. Compared with the attention layer in the traditional Transformer model, the attention mechanism layer becomes more concise and faster in propagation. More importantly, the model accuracy no longer overly depends on the accuracy of attention score calculation. In the defined attention mechanism layer, instead of mapping the input data into three types of vectors with three different weight matrices, it directly propagates and calculates through a fully connected layer, and the calculation formula is as follows:
[0057] O(X) = φ(W O X)(5)
[0058] In Equation (5), O(X) is the propagation vector; φ is the activation function; W O is the weight matrix for mapping the propagation vector; X is the feature sequence data of the unit time period.
[0059] Although this method improves the operation efficiency, its ability to filter features is not strong. Therefore, a gating unit is added to the attention mechanism layer to further improve the model ability. The improved calculation formulas for the query vector Q and the key vector K in the attention mechanism layer are as follows:
[0060]
[0061] In Equation (6), Q is the query vector; is a trainable hyperparameter used to control the overall gating forgetting degree. The calculation formula for the key vector K is the same as that for the query vector Q, only the value of the trainable hyperparameter is different.
[0062] After completing the vector calculation, the calculation formula for the attention score is as follows:
[0063] A = αV(7)
[0064] In Equation (7), A is the first attention score; α is the intermediate parameter; V is the value vector.
[0065] α = relu(Q * K)(8)
[0066] In Equation (8), relu() is the activation function; K is the key vector.
[0067] The final output of the attention mechanism layer is output, and the calculation formula is as follows:
[0068] output = αV + V(9)
[0069] At this time, in this embodiment, the attention mechanism layer includes: a second fully connected layer, a gating unit, a first attention score calculation unit, and a third addition layer. The first output end of the second fully connected layer is respectively connected to the first input end of the first attention score calculation unit and the first input end of the third addition layer. The second output end of the second fully connected layer is connected to the input end of the gating unit. The output end of the gating unit is connected to the second input end of the first attention score calculation unit. The output end of the first attention score calculation unit is connected to the second input end of the third addition layer.
[0070] Among them, the second fully connected layer is used to calculate a value vector and a propagation vector based on the feature sequence data in a unit time period. The calculation formula of the value vector is the above formula (1), and the calculation formula of the propagation vector is the above formula (5).
[0071] Among them, the gating unit is used to calculate a query vector and a key vector based on the propagation vector. The calculation formula of the query vector is the above formula (6), and the calculation formula of the key vector is the same as that of the query vector, only the value of the trainable hyperparameter needs to be changed.
[0072] Among them, the first attention score calculation unit is used to calculate the first attention score corresponding to a unit time period based on the query vector, the key vector, and the value vector. The calculation formula of the first attention score is the above formula (7) and the above formula (8).
[0073] Among them, the third addition layer is used to sum the first attention score and the value vector to obtain an output vector.
[0074] (II) The second improvement method
[0075] In the prediction task of photovoltaic power generation, the traditional model models the task as a conventional time series prediction task, and the form of introducing features is also linearly introduced in time series. This ignores the data noise problem caused by external forces and the overfitting of the traditional model to noise data, and also produces the problem of being difficult to learn the correlation between various features. In view of the above existing problems, in order to improve the prediction accuracy, solve the time misalignment situation and data noise problem between the collected feature data and the real feature data of the data features due to comprehensive factors, and the problems such as too many training parameters of the traditional Transformer model and inability to notice the variability in the time series, this embodiment introduces a multi-scale convolution module, a feature difference module, and a time series prediction module to solve the above problems.
[0076] This embodiment proposes an improved Transformer model processed by features (which can be called Conv-attention or Conv-ATT), which is a short-term photovoltaic power prediction model based on a short-sequence multi-scale convolutional attention mechanism. In view of the time misalignment in traditional feature data collection, the way of collecting feature data is reconsidered, and then the correlation between feature data is reconstructed through a multi-scale convolutional module to obtain cross-physical features. The feature difference module is used to further enhance the differential effect of the model in learning different features. Finally, the processed features are input into the time series prediction module for prediction. It is an improved prediction model proposed for the problems of many parameters and information loss when the traditional Transformer model processes short-sequence features. It not only combines the advantages of the attention mechanism in the Transformer model to make the model more suitable for the prediction of photovoltaic power, but also uses methods such as multi-scale convolution to further fuse features, so as to improve the correlation between various features of photovoltaic power, help the model fully explore the correlation changes of real physical features, and finally obtain high-quality and consistent short-term photovoltaic power prediction results. As Figure 5 and Figure 6 shown, the attention mechanism layer includes multiple multi-layer perceptrons, multiple multi-scale convolutional modules and a second attention score calculation unit (i.e., the feature difference module), and the time series prediction module is Figure 3 the subsequent part of the attention mechanism layer in Figure 3 . The value of the same feature within a unit time period is taken as an input separately and passed through a multi-layer perceptron, enabling the model to learn multiple inputs of multiple features. This method not only solves the misalignment phenomenon during feature collection but also reduces the risk of overfitting of the model. At the same time, a multi-scale convolutional module including convolutional kernels of different scales is considered to extract multi-source input features to obtain cross-physical features, thereby reconstructing the correlation between different features. The obtained cross-physical features are input into the feature difference module for differential learning. Since the traditional attention layer cannot solve the problem of feature differences in the time dimension, the vector calculation structure of the attention layer is improved to fully simulate the differential effects of various features affecting photovoltaic power. After obtaining sufficient physical features, the time series features are incorporated through the time series prediction module to further help the model complete the prediction.
[0077] Due to different collected features and different sensors used, measurement delay problems occur due to hardware, network and other issues, resulting in a misalignment phenomenon in the time of the values of each feature. To solve this problem, it is considered to re-model the problem and solve it by re-establishing the input method of features. The photovoltaic power prediction task is defined as a time series prediction task of multi-feature variables. The time length is set to T and the number of features is N. At the same time, the feature data within the unit time period recorded by the sensor is denoted as:
[0078] X = {x 1 , x 2 ,......, x T} ∈ R T*N (10)
[0079] Based on the above characteristic data, predict the photovoltaic power generation within a certain period in the future (i.e., the next unit time period of the historical time period).
[0080] The sequence data of a single feature within the time length T (i.e., the value of a single feature at each moment within the time length T) can be directly mapped to an input vector, and its calculation formula is as follows:
[0081] v i = Emb(X)(11)
[0082] In formula (11), v i represents the input vector mapped from the sequence data of feature i, and i ∈ [1, 4]; Emb() is a multi-layer perceptron applicable to all features.
[0083] Mapping multiple features to input vectors respectively solves the problem of time asynchrony, but at the same time introduces new problems: the correlation between features is separated more thoroughly, that is, although the non-alignment phenomenon of features is solved, the correlation between features will be weakened. Therefore, it is necessary to introduce a multi-scale convolution module to re-establish the connection between features, and finally obtain cross-physical features. The multi-scale convolution module means that different-sized convolution kernels are used to calculate the input features, and their sizes are 1*N, 2*N, and 3*N respectively. The reason for choosing the number of features N as the convolution width is that there is a correlation between the external features before and after within a certain time range. The step size of the convolution operation is set to 1, and cross-physical features are obtained after passing through this multi-scale convolution module. One-dimensional convolution based on the number of features N can adaptively adjust the time-asymmetric situation inside the convolution and reduce noise, which is also the advantage of the convolution network for feature extraction.
[0084] Since the sequence features are short and the global correlation is weak, the encoding part of the sequence in the traditional Transformer model is discarded, which reduces the computational memory of the model. The attention layer of the traditional Transformer model calculates the approximate scoring situation between the query point and the key, and does not fully utilize the information before and after in time series (i.e., local context information). Therefore, a feature difference module is proposed, and each cross-physical feature is used as a new calculation node and calculated through the attention mechanism. The following introduces the calculation process of the query vector Q and the key vector K:
[0085] It is assumed that after the feature data of n unit time periods are calculated by the multi-scale convolution module, n cross-physical features X are obtained n represents the number of unit time periods.
[0086] If the mapping of the weight matrix in the traditional attention layer is directly applied to the query vector Q and the key vector K, it is difficult to obtain a vector representation with the same dimension as the value vector V. Therefore, it is necessary to perform reasonable operations on multiple features within a certain previous step length in advance to transform them into vectors with the same dimension as the value vector V. At this time, it is necessary to effectively fuse the convolution calculation results in the multi-scale convolution module and perform attention operations on each cross-physical feature. The calculation formula is as follows:
[0087]
[0088] In Equation (12), W q is the weight matrix for mapping the query vector; Conv is a one-dimensional convolution operation; When all cross-physical features are sorted in chronological order, it is the sequence composed of all cross-physical features between and ; is the i-th cross-physical feature, where i is equal to the difference between j and the sequence length; is the j-th cross-physical feature, where j = 1, 2,..., n, and n is the number of cross-physical features; W k is the weight matrix for mapping the key vector.
[0089] In Equation (12), i represents the subscript of the starting node of the convolution calculation, j represents the subscript of the ending node of the convolution calculation, and also the subscript of the target node of the convolution calculation. The convolution sampling step is j - i + 1. When the target node is the head node, the number of nodes in front of the target node is less than the convolution step. At this time, 0 values are selected for padding. After the above steps, the operation of the attention mechanism is no longer a point-like calculation, but a joint calculation of the current target feature and the previous features.
[0090] At this time, in this embodiment, the attention mechanism layer includes multiple multi-layer perceptrons, multiple multi-scale convolution modules, and a second attention score calculation unit. The number of multiple multi-layer perceptrons, the number of multiple multi-scale convolution modules, and the number of unit time periods in the historical time period are the same. The multi-layer perceptrons and the multi-scale convolution modules are in one-to-one correspondence. The output end of the multi-layer perceptron is connected to the input end of the multi-scale convolution module, and the output ends of all multi-scale convolution modules are connected to the input end of the second attention score calculation unit.
[0091] Among them, the multi-layer perceptron is used to calculate the feature vector of each feature by taking the sequence data of each feature in the feature sequence data of the unit time period as the input respectively. The sequence data of the feature includes the values of the feature at each moment within the unit time period. The calculation formula of the feature vector is the above Equation (11).
[0092] Among them, the multi-scale convolution module is used to perform feature extraction and feature fusion on the feature vectors of all features using convolution kernels of multiple scales, and obtain cross-physical features.
[0093] The multi-scale convolution module includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a splicing layer. The first convolutional layer, the second convolutional layer, and the third convolutional layer are connected in parallel. The output ends of the first convolutional layer, the second convolutional layer, and the third convolutional layer are all connected to the input end of the splicing layer. Among them, the convolutional kernel size of the first convolutional layer is 1*N, the convolutional kernel size of the second convolutional layer is 2*N, and the convolutional kernel size of the third convolutional layer is 3*N, where N is the number of features.
[0094] Among them, the second attention score calculation unit is used to calculate the second attention score corresponding to each unit time period in the historical time period based on the cross-physical features output by all multi-scale convolution modules. The calculation formula for the second attention score is the above formula (4), the calculation formulas for the query vector Q and the key vector K are the above formula (12), and the calculation formula for the value vector V is the above formula (1).
[0095] When building the model, the misalignment of different feature data in the time scale during the reading of the data set is considered. The method of short sequence single feature is used to reduce the deviation between the collected features and the real features, and finally achieve the effect of reducing data noise. In order to re-establish the correlation between different features and reduce the model complexity, a one-dimensional convolutional neural network is used to extract multi-source input features, where the sizes of the convolutional kernels are 1*N, 2*N, and 3*N respectively. The reason for choosing the feature quantity N as the convolutional width is that the external features before and after within a certain time range are considered to have a correlation relationship. Through the above method, the error brought by the data itself is reduced and cross-physical features are obtained. The introduction of the feature difference module is essentially to change the mapping form of the attention mechanism, so that the scoring mechanism of the model incorporates temporal information, and this structure avoids information leakage.
[0096] Before using the feature sequence data of each unit time period in the historical time period as the input to determine the power sequence data of the next unit time period in the historical time period using the trained photovoltaic power prediction model, the short-term photovoltaic power prediction method of this embodiment further includes:
[0097] (1) Calculate the Pearson correlation coefficient between each influencing factor and the photovoltaic power, and screen the features based on all the Pearson correlation coefficients.
[0098] The influencing factors include multiple factors that affect the photovoltaic power, such as the tilt angle of the photovoltaic panel, the incident angle of the radiation surface, temperature, humidity, wind speed, and irradiance.
[0099] Perform a correlation analysis on multiple influencing factors to select appropriate features. Specifically, collect data related to the power generation of a photovoltaic power station (i.e., data of various influencing factors and data of photovoltaic power generation at the same moment), and select features with strong correlation through Pearson correlation analysis. The Pearson correlation coefficient is the most widely used correlation analysis method currently. By using the Pearson correlation coefficient to reduce the dimensionality of influencing factors, redundant information and computational complexity can be reduced, and overfitting can be prevented. The Pearson correlation coefficient measures the strength and direction of the linear relationship between two variables through the regularity of data distribution. The value range of its calculation result is [-1, 1]. 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no linear correlation. The calculation formula of the Pearson correlation coefficient is as follows:
[0100]
[0101] In Equation (13), ρ m,n is the Pearson correlation coefficient between variable m and variable n; m is the photovoltaic power generation; n is the influencing factor; cov is the covariance; σ m and σ n are the standard deviations of variable m and variable n respectively; E is the expectation; m i is the expectation of variable m; n i is the expectation of variable n. In this embodiment, influencing factors with |ρ m,n | > 0.2 are selected as the input features of the model.
[0102] (2) Construct a total data set based on all features. The total data set includes data sets for each season. The data set includes multiple samples and the labels corresponding to each sample. The sample is the feature sequence data of each unit time period in the sample historical time period, and the label is the power sequence data of the next unit time period in the sample historical time period.
[0103] Due to the seasonal volatility of photovoltaic power generation, considering the strong correlation between the power generation of a photovoltaic power station and irradiance, not only the influence of seasonal factors on photovoltaic power generation needs to be considered, but also the influence of other features on photovoltaic power generation and the correlation between features need to be considered. Therefore, irradiance, humidity, temperature, wind speed, and actual photovoltaic power generation data every 15 minutes in different seasons of a certain photovoltaic power station are intercepted as the total data set. Specifically, irradiance, humidity, temperature, wind speed, and actual photovoltaic power generation data every 15 minutes from 7:00 to 19:00 every day in March, June, September, and November of a certain photovoltaic power station are intercepted as the total data set.
[0104] Dividing the total data set by season is beneficial to reducing the influence of seasonal factors on the data.
[0105] (3) Use the total dataset to train the initial photovoltaic power prediction model to obtain a trained photovoltaic power prediction model.
[0106] To highlight the advantages of this embodiment, two additional traditional neural network models are selected as comparison models, and the prediction results are evaluated using two error evaluation metrics: Mean Absolute Error (MAE) and Root Mean Square Error (RMSE). The core of calculating the error evaluation metrics is to analyze the gap between the true value and the predicted value and display the overall error level in a reasonable normalized manner. The calculation formulas for the two commonly used error evaluation metrics are as follows:
[0107]
[0108] In Equation (14), n is the number of samples; y i is the true value of the i-th sample; is the predicted value of the i-th sample. The smaller the MAE and RMSE, the higher the accuracy of the model prediction and the better the fitting effect.
[0109] In this embodiment, two traditional Long Short-Term Memory neural networks, LSTM and ResidualLSTM, are selected for comparative analysis with the two Transformer variant models based on the attention mechanism provided in this embodiment. First, compare the loss values of the above four models during the training process, as Figure 7 shown. It can be seen that, on the one hand, the training loss of Cate-ATT and Conv-ATT decreases faster than that of LSTM and ResidualLSTM. This is due to the fact that both Attention and CNN can be trained in parallel and parameter sharing can be achieved, while LSTM can only be trained linearly. On the other hand, there is an obvious gap in the training accuracy in the later stage, which fully demonstrates the advantages of the design of this embodiment. Moreover, the loss of ResidualLSTM fluctuates greatly, indicating that the model is too sensitive to some data, or some data belongs to noise data. The Conv-ATT method effectively avoids this situation by combining CNN to extract features from the data in a short time span in the early stage and obtains good results.
[0110] Figures 8 - 11Shows the comparison of the power generation power prediction effects of four models in different seasons. Real represents the real power change curve. It can be clearly seen that the predictions of Cate-ATT and Conv-ATT for the power generation power are smoother. Although there are certain errors, the overall slope change trend of the prediction results is similar to that of the real power change curve. In addition, referring to Table 1, it can be known that Cate-ATT and Conv-ATT have obtained the best values for the evaluation indicators in the short-term photovoltaic power generation power prediction tasks in the four seasons, which fully demonstrates the comprehensive learning effect of Cate-ATT and Conv-ATT on the fine-grained and multi-source cross of various features. In the prediction results of other models, when there are turning points where the power drops or rises sharply, the predicted values still have gaps compared with the predicted values of the Cate-ATT and Conv-ATT models, indicating that other models overfit the training data and do not fully explore the relationships between features.
[0111] Table 1 Performance Index Table of Each Model in Four Seasons
[0112]
[0113]
[0114] The present application also provides an application scenario, which applies the above-mentioned short-term photovoltaic power generation power prediction method. Specifically, the short-term photovoltaic power generation power prediction method provided in this embodiment can be applied in a photovoltaic grid connection scenario. The photovoltaic grid connection scenario includes a power prediction link and a grid connection link. The power prediction link is used to predict the power sequence data for each unit time period in the prediction time period, and the grid connection link is used to achieve photovoltaic grid connection based on the power sequence data for each unit time period in the prediction time period. The short-term photovoltaic power generation power prediction method provided in this embodiment belongs to the power prediction link.
[0115] Embodiment 2
[0116] Based on the same inventive concept, the embodiment of the present application also provides a short-term photovoltaic power generation power prediction device for implementing the above-mentioned short-term photovoltaic power generation power prediction method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the short-term photovoltaic power generation power prediction device provided below can refer to the limitations on the short-term photovoltaic power generation power prediction method in the above text, and will not be repeated here.
[0117] In an exemplary embodiment, as Figure 12 shown, a short-term photovoltaic power generation power prediction device is provided. The short-term photovoltaic power generation power prediction device includes:
[0118] A data acquisition module M1 is used to acquire the characteristic sequence data of each unit time period in a historical time period; the characteristic sequence data includes the values of each characteristic at each moment within the unit time period, the duration of the unit time period is less than a preset duration, and the characteristics include irradiance, humidity, temperature, and wind speed.
[0119] A power prediction module M2 is used to take the characteristic sequence data of each unit time period in a historical time period as input, and use a trained photovoltaic power prediction model to determine the power sequence data of the next unit time period in the historical time period; the power sequence data includes the values of the photovoltaic power at each moment within the unit time period.
[0120] Among them, the trained photovoltaic power prediction model is a model obtained by simplifying the structure of the Transformer model. The trained photovoltaic power prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and layer normalization layer, a first fully connected layer, and an output layer, which are connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and layer normalization layer.
[0121] Embodiment 3
[0122] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 13 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a short-term photovoltaic power prediction method.
[0123] Those skilled in the art can understand, Figure 13The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0124] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the short-term photovoltaic power prediction method in Embodiment 1 is implemented.
[0125] Embodiment 4
[0126] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the short-term photovoltaic power prediction method in Embodiment 1 is implemented.
[0127] Embodiment 5
[0128] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the short-term photovoltaic power prediction method in Embodiment 1 is implemented.
[0129] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0131] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on this application.
Claims
1. A short-term photovoltaic power generation prediction method, characterized in that: The short-term photovoltaic power generation power prediction method comprises: Acquire feature sequence data of each unit time period in the historical time period; the feature sequence data includes the value of each feature at each moment in the unit time period, the length of the unit time period is less than the preset length, and the features include irradiance, humidity, temperature and wind speed; Taking the characteristic sequence data of each unit time period in the historical time period as input, the trained photovoltaic power generation power prediction model is used to determine the power sequence data of the next unit time period of the historical time period; the power sequence data includes the value of the photovoltaic power generation power at each moment in the unit time period; Among them, the trained photovoltaic power generation prediction model is a model obtained by structurally simplifying the Transformer model. The trained photovoltaic power generation prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and a layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and a layer normalization layer, a first fully connected layer and an output layer connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and the layer normalization layer.
2. The short-term photovoltaic power generation prediction method according to claim 1, characterized in that: The attention mechanism layer includes: a second fully connected layer, a gating unit, a first attention score calculation unit and a third addition layer, the first output end of the second fully connected layer is respectively connected to the first input end of the first attention score calculation unit and the first input end of the third addition layer, the second output end of the second fully connected layer is connected to the input end of the gating unit, the output end of the gating unit is connected to the second input end of the first attention score calculation unit, and the output end of the first attention score calculation unit is connected to the second input end of the third addition layer; The second fully connected layer is used to calculate the value vector and propagation vector based on the feature sequence data of the unit time period; The gating unit is used to calculate the query vector and the key vector based on the propagation vector; The first attention score calculation unit is used to calculate the first attention score corresponding to the unit time period based on the query vector, the key vector and the value vector.
3. The short-term photovoltaic power generation prediction method according to claim 2, characterized in that: The formula for calculating the propagation vector is: O(X)=φ(W O X); Among them, O(X) is the propagation vector; φ is the activation function; W O is the weight matrix of the mapping propagation vector; X is the characteristic sequence data of the unit time period; The calculation formula of the query vector is: Where Q is the query vector; is a trainable hyperparameter; The calculation formula for the key vector is the same as that for the query vector, and only the values of the trainable hyperparameters need to be changed; The calculation formula for the first attention score is: A = αV; Among them, A is the first attention score; α is the intermediate parameter; V is the value vector; α = relu(Q*K); Among them, K is the key vector.
4. The short-term photovoltaic power generation prediction method according to claim 1, characterized in that: The attention mechanism layer includes multiple multi-layer perceptrons, multiple multi-scale convolution modules and a second attention score calculation unit. The number of the multiple multi-layer perceptrons and the number of the multiple multi-scale convolution modules are the same as the number of unit time periods in the historical time period. The multi-layer perceptrons and the multi-scale convolution modules correspond one to one. The output end of the multi-layer perceptron is connected to the input end of the multi-scale convolution module, and the output ends of all the multi-scale convolution modules are connected to the input end of the second attention score calculation unit. The multilayer perceptron is used to take the sequence data of each feature in the feature sequence data of the unit time period as input, and calculate the feature vector of each feature; the feature sequence data includes the value of the feature at each moment in the unit time period; The multi-scale convolution module is used to extract and fuse the feature vectors of all features using convolution kernels of multiple scales to obtain cross-physical features; The second attention score calculation unit is used to calculate the second attention score corresponding to each unit time period in the historical time period based on the cross-physical features output by all multi-scale convolution modules.
5. The short-term photovoltaic power generation prediction method according to claim 4, characterized in that: The multi-scale convolution module includes a first convolution layer, a second convolution layer, a third convolution layer and a splicing layer. The first convolution layer, the second convolution layer and the third convolution layer are connected in parallel. The output ends of the first convolution layer, the second convolution layer and the third convolution layer are all connected to the input end of the splicing layer. The convolution kernel size of the first convolution layer is 1*N, the convolution kernel size of the second convolution layer is 2*N, and the convolution kernel size of the third convolution layer is 3*N, where N is the number of features. The calculation formula for the second attention score is: Among them, h n is the second attention score; Q is the query vector; K is the key vector; d k is the dimension of the key vector; V is the value vector; Among them, W q is the weight matrix for mapping query vector; To sort all cross-physical features in chronological order, arrive The sequence of all intersecting physical features between them; is the i-th crossover physical feature, i is equal to the difference between j and the sequence length; is the jth cross-physical feature, j = 1, 2, ..., n, n is the number of cross-physical features; W k is the weight matrix of the mapping key vector.
6. The short-term photovoltaic power generation prediction method according to claim 1, characterized in that: Before taking the characteristic sequence data of each unit time period in the historical time period as input and using the trained photovoltaic power generation power prediction model to determine the power sequence data of the next unit time period in the historical time period, the short-term photovoltaic power generation prediction method further includes: Calculate the Pearson correlation coefficient between each influencing factor and photovoltaic power generation, and screen and obtain features based on all the Pearson correlation coefficients; A total data set is constructed based on all the features; the total data set includes a data set for each season, the data set includes multiple samples and a label corresponding to each sample, the sample is the feature sequence data of each unit time period in the sample historical time period, and the label is the power sequence data of the next unit time period of the sample historical time period; The total data set is used to train the initial photovoltaic power generation prediction model to obtain a trained photovoltaic power generation prediction model.
7. A short-term photovoltaic power generation prediction device, characterized in that: The short-term photovoltaic power generation prediction device comprises: A data acquisition module is used to acquire feature sequence data of each unit time period in the historical time period; the feature sequence data includes the value of each feature at each moment in the unit time period, the length of the unit time period is less than the preset length, and the features include irradiance, humidity, temperature and wind speed; The power prediction module is used to use the characteristic sequence data of each unit time period in the historical time period as input, and use the trained photovoltaic power generation power prediction model to determine the power sequence data of the next unit time period of the historical time period; the power sequence data includes the value of the photovoltaic power generation power at each moment in the unit time period; Among them, the trained photovoltaic power generation prediction model is a model obtained by structurally simplifying the Transformer model. The trained photovoltaic power generation prediction model includes an input layer, an attention mechanism layer, a first addition layer, a first residual connection and a layer normalization layer, a forward propagation layer, a second addition layer, a second residual connection and a layer normalization layer, a first fully connected layer and an output layer connected in sequence. The input end of the first addition layer is also connected to the output end of the input layer, and the input end of the second addition layer is also connected to the output end of the first residual connection and the layer normalization layer.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the short-term photovoltaic power generation prediction method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the short-term photovoltaic power generation prediction method described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the short-term photovoltaic power generation prediction method described in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Photovoltaic power prediction method
CN121256366A