Photovoltaic power generation power prediction method based on deep reinforcement learning
By constructing a dual data stream CNN-LSTM encoder and dual attention mechanism, combined with an adaptive reward compensator and A3C algorithm, the problems of spatiotemporal and spatial characteristics decoupling and insufficient meteorological sensitivity in photovoltaic power generation prediction are solved, and high-precision short-term prediction is achieved.
Patent Information
- Application Number
- CN202510482920.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-15
AI Technical Summary
The existing short-term photovoltaic power prediction methods have problems such as spatiotemporal and spatial characteristics decoupling and insufficient meteorological sensitivity, resulting in a decrease in prediction accuracy and poor robustness under complex meteorological conditions.
Using a method based on deep reinforcement learning, we use the dual data flow CNN-LSTM encoder to extract meteorological space features, combine the dual attention mechanism and adaptive reward compensator, dynamically adjust the reward weight, and use the A3C algorithm to conduct parallel training of multiple agents to optimize the prediction strategy.
The prediction accuracy and robustness of photovoltaic power generation under complex meteorological conditions are significantly improved, and the model determination coefficient R2 reaches 0.96468, which is better than the existing model.
Smart Images

Figure CN120494156A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and reinforcement learning, specifically to a short-term photovoltaic power generation prediction method based on deep reinforcement learning, and more particularly to a high-precision prediction technology that combines spatiotemporal feature fusion with a dynamic reward mechanism. Background Art
[0002] With the rapid development of renewable energy worldwide, the proportion of solar photovoltaic power generation in the energy structure has increased year by year. However, photovoltaic power is significantly affected by meteorological conditions and has strong volatility and uncertainty, posing challenges to the stable operation of the power grid. Existing short-term photovoltaic power forecasting methods face two major technical bottlenecks:
[0003] First, the problem of spatiotemporal decoupling: Traditional models (such as LSTM, CNN, etc.) usually process time series or spatial features independently, and it is difficult to effectively capture the coupling relationship between cloud movement (space) and diurnal changes in irradiance (time), resulting in a decrease in prediction accuracy under complex meteorological conditions.
[0004] Second, there is insufficient meteorological sensitivity: existing models have limited ability to dynamically assign weights to key meteorological variables such as irradiance and temperature, and lack an adaptive adjustment mechanism for the confidence level of numerical weather forecasts. This results in poor robustness in extreme weather conditions (such as sudden thunderstorms and heavy rainfall). Furthermore, most research focuses on deep learning models themselves, with limited integration of the dynamic decision-making advantages of reinforcement learning into the forecasting process. This makes it difficult to guide models to optimize long-term forecasting strategies through reward mechanisms.
[0005] Although previous studies have attempted to improve forecasting performance through hybrid models (such as LSTM-BNN), their deep integration of spatiotemporal features and dynamic response to meteorological variables remain insufficient. Therefore, a new method that can efficiently integrate spatiotemporal information and adaptively adjust forecasting strategies is urgently needed to meet the high-precision forecasting needs of high-proportion photovoltaic grid access. Summary of the Invention
[0006] The purpose of this invention is to address the problems of spatiotemporal feature decoupling and insufficient meteorological sensitivity in the existing technology, and to provide a photovoltaic power generation prediction method based on deep reinforcement learning. Through spatiotemporal feature encoding, dual attention mechanism and adaptive reward compensation, the prediction accuracy and robustness under complex meteorological conditions are improved.
[0007] The present invention can be achieved through the following technical solutions:
[0008] A photovoltaic power generation power prediction method based on deep reinforcement learning specifically includes the following steps:
[0009] S1. Integrate photovoltaic power station power generation data and meteorological data, and process key parameters such as ambient temperature and irradiance.
[0010] S2. Construct a dual-stream CNN-LSTM encoder, extract meteorological spatial features through convolutional neural networks, and combine gated recurrent units to capture time series patterns.
[0011] S3. Using a dual attention mechanism, the time module focuses on key time points, and the weather module dynamically weights variables; the reward weight is adjusted according to the confidence of the weather forecast through an adaptive compensator.
[0012] S4. Use the A3C algorithm to implement multi-agent parallel training and maximize long-term cumulative rewards to optimize the prediction strategy.
[0013] S5. Verify on a real photovoltaic power station dataset.
[0014] The process of integrating photovoltaic power station power generation data and meteorological data and processing key parameters such as ambient temperature and irradiance in step S1 includes the following steps:
[0015] S11. Dataset Construction: The dataset is derived from a public dataset available at https: / / www.kaggle.com / code / pythonafroz / solar-power-generation-forecast / input. We integrated PV plant power generation data (including DC power, AC power, and daily power generation) with meteorological data (including ambient temperature, module temperature, and irradiance) and aligned them by timestamps to create a dataset containing 68,778 data points.
[0016] S12. Conventional data loading and merging.
[0017] S13. Feature Engineering: Classify and encode non-numerical features (such as sensor ID) and convert them into numerical features (SOURCE_KEY_NUMBER) to facilitate model processing.
[0018] S14. Data segmentation and scaling: The dataset is divided into training and test sets in an 8:2 ratio. The minimum-maximum normalization method is used to scale the features to the [0, 1] interval to avoid data leakage and improve model training efficiency.
[0019] Furthermore, the calculation formula for conventional data loading in step S12 is as follows:
[0020]
[0021] Among them, T i is the timestamp, DC i is DC power, AC i is the AC power, DY i is the daily power generation, TY iis the total power generation, T i is the ambient temperature, AT i is the module temperature, MT i is the irradiation intensity, I i is the number of data points.
[0022] Furthermore, the calculation formula for conventional data merging in step S12 is as follows:
[0023]
[0024] This means joining two datasets based on their timestamp (T).
[0025] Furthermore, the calculation formula of the feature engineering in step S13 is as follows:
[0026] SK i →SKN i ∈{1,2,...,K}
[0027] Among them, SK i Represents the SOURCE_KEY column, and there are (K) different sensor IDs. The categorical encoding maps each sensor ID to an integer. SKN i The value of the SOURCE_KEY_NUMBER column.
[0028] Furthermore, the calculation formula for data segmentation in step S14 is as follows:
[0029] D train =D[1:r·N]
[0030] D test =D[r·N:N]
[0031] Among them, D train is the training data set, D test For the test data set, r is 0.8.
[0032] Furthermore, the calculation formula for data scaling in step S14 is as follows:
[0033]
[0034] Here, the pixels are scaled to the range [0, 1]. X is each eigenvalue, and X' is the scaling value.
[0035] The process of constructing the dual-stream CNN-LSTM encoder in step S2 includes the following steps:
[0036] S21. Spatial feature extraction: Convolutional neural network (CNN) is used to extract local features of meteorological data. The original 6-dimensional features are expanded into 128-dimensional abstract spatial features through two convolutional layers (convolution kernel size 3×1, output channels 64 and 128 respectively) to capture the short-term correlation between adjacent time steps.
[0037] S22. Time Series Modeling: The spatial features of CNN output are processed in the time dimension through the long short-term memory network (LSTM), the long-term dependencies in the time series are learned, and the hidden state of each time step (dimension 128) is output to provide input for the subsequent attention mechanism.
[0038] Furthermore, the calculation formula for spatial feature extraction in step S21 is as follows:
[0039] X norm ∈R B×F×T
[0040] F2=CNN(X norm )
[0041] Among them, R B×F×T is the original data set, and CNN() is the convolutional neural network.
[0042] Furthermore, the calculation formula for time series modeling in step S22 is as follows:
[0043] h t =LSTM(F2)
[0044] Among them, LSTM() is a long short-term memory network.
[0045] The process of dual attention mechanism and adaptive reward design in step S3 includes the following steps:
[0046] S31, Temporal Attention Module: Automatically learn the importance weights of different time steps through the self-attention mechanism, focusing on key time nodes that have a significant impact on the prediction results (such as the moment of irradiance mutation).
[0047] S32, Meteorological feature attention module: Dynamically weighted impact of meteorological variables such as irradiance and temperature on photovoltaic power.
[0048] S33, Adaptive Reward Compensator: Designs a reward function based on relative error and dynamically adjusts reward weights based on the confidence level of the numerical weather forecast. When the relative error of the forecast exceeds a threshold, linear and quadratic penalties are applied to incentivize the model to reduce forecast bias.
[0049] Furthermore, the calculation formula of the temporal attention module in step S31 is as follows:
[0050] αt =Softmax(W a tanh(W h h t +b h ))
[0051]
[0052] Among them, h t is the hidden state of LSTM at time step t, α t is the temporal attention weight, and z is the time-weighted feature vector.
[0053] Furthermore, the calculation formula of the meteorological feature attention module in step S32 is as follows:
[0054] β f =Softmax(W b tanh(W z z+b z ))
[0055]
[0056] Among them, β f is the feature attention weight, and s is the state vector after feature weighting.
[0057] Furthermore, the calculation formula of the adaptive reward compensator in step S33 is as follows:
[0058]
[0059] Among them, ∈(t) is the relative error, α is the reward scaling factor, ω1,ω2 are penalty factors, and β is the quadratic penalty coefficient.
[0060] The process of parallel training of the A3C algorithm in step S4 includes the following steps:
[0061] S41, environment setting, defines the environment for power prediction of photovoltaic power station.
[0062] S42. Define A3C network
[0063] S43. Define A3C work process
[0064] S44, multi-agent parallel training
[0065] S45. Prediction and Evaluation
[0066] Furthermore, the process of setting the environment in step S41 is as follows:
[0067] Define the photovoltaic power prediction environment, including state space, action space and reward function.
[0068] Furthermore, the process of defining the A3C network in step S42 is as follows:
[0069] Construct an A3C network including a CNN-LSTM encoder and a dual attention module to output action probability and state value functions.
[0070] Furthermore, the process of defining the A3C work process in step S43 is as follows:
[0071] A3C defines the behavior of each worker process. Each worker process has its own local model and initially loads the parameters of the global model. Workers interact in the environment, collect experience (state, action, reward, etc.), calculate the return and advantage function, and update the parameters of the global model.
[0072] Furthermore, the process of defining the A3C network in step S44 is as follows:
[0073] Start multi-agent parallel training, each process independently interacts with the environment, and updates the global model parameters to maximize the cumulative reward.
[0074] Furthermore, the prediction and evaluation formulas in step S45 are as follows:
[0075]
[0076] Here, MAE refers to the average of the absolute errors between all predicted values and the true values.
[0077]
[0078] Here, MSE refers to the average of the squared errors between all predicted values and the true values.
[0079]
[0080] Among them, R 2 It is used to measure the degree to which the model explains the data, and its value range is [0, 1]. 2 The closer it is to 1, the better the model fits the data.
[0081] The process of performing verification on a real photovoltaic power station dataset in step S5 includes the following steps:
[0082] The model was validated on a real photovoltaic power station dataset, using the PyTorch framework for model training and GPUP100 for accelerated computing. The mean absolute error (MAE), mean square error (MSE), and coefficient of determination (R 2 ) is the evaluation indicator.
[0083] In addition, the present invention proposes a photovoltaic power generation power prediction method based on deep reinforcement learning, comprising:
[0084] A collection module is used to obtain relevant parameters of photovoltaic power generation;
[0085] Loading and merging modules, used to load and merge data sets;
[0086] The split and scale module splits the dataset into training and test sets and scales features to improve the training efficiency and generalization ability of the model;
[0087] The dual-stream encoder module extracts meteorological spatial features through convolutional neural networks and combines gated recurrent units to capture time series patterns.
[0088] Dual attention mechanism module: the time module focuses on key time points, and the weather module dynamically weights variables;
[0089] The reward module adjusts the reward weight according to the confidence of the weather forecast through an adaptive compensator;
[0090] The reinforcement learning module introduces the A3C algorithm. Combined with the above modules, the A3C algorithm uses multi-agent parallel training to continuously optimize the parameters of the neural network model to maximize the long-term cumulative reward, thereby improving the accuracy of photovoltaic power plant power prediction;
[0091] The analysis module evaluates the prediction performance by calculating the mean absolute error (MAE), mean square error (MSE) and R2 score between the predicted power and the actual power;
[0092] Compared with the prior art, the present invention has the following beneficial effects:
[0093] The results showed that the model R 2 It reaches 0.96468, which is significantly better than the comparison model (such as LSTM-BNN's R 2 The R of DRL-TD3 is 0.90758. 2 is 0.94274), which verifies the feasibility of high-precision short-term prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] In order to more clearly illustrate the technical solutions of the embodiments of the present application, a brief introduction to the drawings in the embodiments will be given below. The following drawings only show certain embodiments of the present application and are therefore not to be construed as limiting the scope. Scholars and technicians in this field can obtain other relevant drawings based on these drawings, and make various modifications and supplementary innovations to the described embodiments.
[0095] Figure 1 Flowchart for deep reinforcement learning photovoltaic power generation time series prediction;
[0096] Figure 2 To compare the prediction results;
[0097] Figure 3 is the evaluation index of all models;
[0098] Figure 4 It is a schematic diagram of the module; DETAILED DESCRIPTION
[0099] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0100] It should be noted that like reference numerals and letters denote like items in the following drawings, and therefore, once an item is defined in one drawing, it does not require further definition or explanation in subsequent drawings.
[0101] The present invention relates to the technical field of renewable energy photovoltaic power prediction, and specifically to a solar power prediction method and system that combines a spatiotemporal feature attention mechanism with an asynchronous advantage actor-critic (A3C) algorithm.
[0102] Based on the above situation, the present invention proposes a photovoltaic power generation power prediction method based on deep reinforcement learning. The present invention is described in detail below with reference to the embodiments and drawings.
[0103] like Figure 1 FIG. 1 is a flow chart of a photovoltaic power generation prediction method based on deep reinforcement learning according to the present invention, which may include the following steps:
[0104] S1. Integrate photovoltaic power station power generation data and meteorological data, and process key parameters such as ambient temperature and irradiance.
[0105] S2. Construct a dual-stream CNN-LSTM encoder, extract meteorological spatial features through convolutional neural networks, and combine gated recurrent units to capture time series patterns.
[0106] S3. Using a dual attention mechanism, the time module focuses on key time points, and the weather module dynamically weights variables; the reward weight is adjusted according to the confidence of the weather forecast through an adaptive compensator.
[0107] S4. Use the A3C algorithm to implement multi-agent parallel training and maximize long-term cumulative rewards to optimize the prediction strategy.
[0108] S5. Verify on a real photovoltaic power station dataset.
[0109] The process of integrating photovoltaic power station power generation data and meteorological data and processing key parameters such as ambient temperature and irradiance in step S1 includes the following steps:
[0110] S11. Dataset construction.
[0111] S12. Conventional data loading and merging.
[0112] S13. Feature Engineering: Classify and encode non-numerical features (such as sensor ID) and convert them into numerical features (SOURCE_KEY_NUMBER) to facilitate model processing.
[0113] S14. Data segmentation and scaling.
[0114] The dataset in step S12 is derived from a public dataset, available at https: / / www.kaggle.com / code / pythonafroz / solar-power-generation-forecast / input. PV power plant power generation data (including DC power, AC power, daily power generation, etc.) and meteorological data (including ambient temperature, module temperature, irradiance, etc.) are integrated and timestamp aligned to form a dataset containing 68,778 data points.
[0115] In step S12, the power generation data (D_G) and the meteorological data (D_W) are merged into a unified data set D based on the timestamp (DATE_TIME field) using the Python pandas library. The calculation formula for conventional data loading is as follows:
[0116]
[0117] Among them, T i is the timestamp, DC i is DC power, AC i is the AC power, DY i is the daily power generation, TY i is the total power generation, T i is the ambient temperature, AT i is the module temperature, MT i is the irradiation intensity, I i is the number of data points.
[0118] The calculation formula for conventional data merging in step S12 is as follows:
[0119]
[0120] This means joining two datasets based on their timestamp (T).
[0121] The calculation formula of the feature engineering in step S13 is as follows:
[0122] SK i →SKN i ∈{1,2,...,K}
[0123] Among them, SK i Represents the SOURCE_KEY column, and there are (K) different sensor IDs. The categorical encoding maps each sensor ID to an integer. SKN i The value of the SOURCE_KEY_NUMBER column.
[0124] The calculation formula for data segmentation in step S14 is as follows:
[0125] D train =D[1:r·N]
[0126] D test =D[r·N:N]
[0127] Among them, D train is the training data set, D test For the test data set, r is 0.8.
[0128] The calculation formula for data scaling in step S14 is as follows:
[0129]
[0130] Here, the pixels are scaled to the range [0, 1]. X is each eigenvalue, and X' is the scaling value.
[0131] The process of constructing the dual-stream CNN-LSTM encoder in step S2 includes the following steps:
[0132] S21. Spatial feature extraction.
[0133] S22. Time series modeling.
[0134] The spatial feature extraction in step S21 is to extract local features of meteorological data using a convolutional neural network (CNN). The original 6-dimensional features are expanded into 128-dimensional abstract spatial features through two convolutional layers (the convolution kernel size is 3×1, and the output channels are 64 and 128 respectively), capturing the short-term correlation between adjacent time steps. The calculation formula for spatial feature extraction is as follows:
[0135] X norm ∈R B×F×T
[0136] F2=CNN(X norm )
[0137] Among them, R B×F×T is the original data set, and CNN() is the convolutional neural network.
[0138] The time series modeling in step S22 is to process the spatial features output by the CNN in the time dimension through the long short-term memory network (LSTM), learn the long-term dependencies in the time series, and output the hidden state (dimension 128) of each time step to provide input for the subsequent attention mechanism. The calculation formula of the time series modeling is as follows:
[0139] h t =LSTM(F2)
[0140] Among them, LSTM() is a long short-term memory network.
[0141] The process of dual attention mechanism and adaptive reward design in step S3 includes the following steps:
[0142] S31, temporal attention module.
[0143] S32, meteorological feature attention module.
[0144] S33, adaptive reward compensator.
[0145] The time attention module in step S31 automatically learns the importance weights of different time steps through the self-attention mechanism, focusing on key time nodes that have a significant impact on the prediction results (such as the moment of irradiance mutation). The calculation formula of the time attention module is as follows:
[0146] α t =Softmax(W a tanh(W h h t +b h ))
[0147]
[0148] Among them, h t is the hidden state of LSTM at time step t, α t is the temporal attention weight, and z is the time-weighted feature vector.
[0149] The meteorological feature attention module in step S32 includes the impact of meteorological variables such as dynamic weighted irradiance and temperature on photovoltaic power. The calculation formula of the meteorological feature attention module is as follows:
[0150] β f =Softmax(W b tanh(W z z+b z))
[0151]
[0152] Among them, β f is the feature attention weight, and s is the state vector after feature weighting.
[0153] The adaptive reward compensator in step S33 designs a reward function based on relative error and dynamically adjusts the reward weight in combination with the confidence level of the numerical weather forecast. When the relative error of the forecast exceeds the threshold, linear and quadratic penalty terms are applied to incentivize the model to reduce the forecast deviation. The calculation formula of the adaptive reward compensator is as follows:
[0154]
[0155] Among them, ∈(t) is the relative error, α is the reward scaling factor, ω1,ω2 are penalty factors, and β is the quadratic penalty coefficient.
[0156] The process of parallel training of the A3C algorithm in step S4 includes the following steps:
[0157] S41, environment setting, defines the environment for power prediction of photovoltaic power station.
[0158] S42. Define A3C network
[0159] S43. Define A3C work process
[0160] S44, multi-agent parallel training
[0161] S45. Prediction and Evaluation
[0162] The process of setting the environment in step S41 is as follows:
[0163] Define the photovoltaic power prediction environment, including state space, action space and reward function.
[0164] The process of defining the A3C network in step S42 is as follows:
[0165] Construct an A3C network including a CNN-LSTM encoder and a dual attention module to output action probability and state value functions.
[0166] The process of defining the A3C work process in step S43 is as follows:
[0167] A3C defines the behavior of each worker process. Each worker process has its own local model and initially loads the parameters of the global model. Workers interact in the environment, collect experience (state, action, reward, etc.), calculate the return and advantage function, and update the parameters of the global model.
[0168] The process of defining the A3C network in step S44 is as follows:
[0169] Start multi-agent parallel training, each process independently interacts with the environment, and updates the global model parameters to maximize the cumulative reward.
[0170] The prediction and evaluation formulas in step S45 are as follows:
[0171]
[0172] Here, MAE refers to the average of the absolute errors between all predicted values and the true values.
[0173]
[0174] Here, MSE refers to the average of the squared errors between all predicted values and the true values.
[0175]
[0176] Among them, R 2 It is used to measure the degree to which the model explains the data, and its value range is [0, 1]. 2 The closer it is to 1, the better the model fits the data.
[0177] The process of performing verification on a real photovoltaic power station dataset in step S5 includes the following steps:
[0178] The model was validated on a real photovoltaic power station dataset, using the PyTorch framework for model training and GPUP100 for accelerated computing. The mean absolute error (MAE), mean square error (MSE), and coefficient of determination (R 2 ) is the evaluation indicator.
[0179] Comparison of experimental results:
[0180] All models used the same parameters and were developed with the help of the PyTorch library. The configuration used was a GPU P100 and Windows 11. The agent's learning rate (LR) was set to 0.0001, the discount factor (GAMMA) was set to 0.99, and the MAX_GLOBAL_STEPS was set to 100,000. The first 80% of the dataset was used for training, and the last 20% was used for validation and analysis. All predictions were performed over a 15-minute period.
[0181] like Figure 2 As shown in the figure, the prediction results of the photovoltaic power prediction method based on deep reinforcement learning of the present invention are compared with the existing mainstream model and the actual value of power generation, including:
[0182] The red solid line represents actual generated power (Actual Power), the yellow solid line represents the predicted power of the photovoltaic power prediction method based on deep reinforcement learning (DRL-TSFP Predicted Power), the green solid line represents the predicted power of the DRL-TD3 model (DRL-TD3Predicted Power), and the purple solid line represents the predicted power of the LSTM-BNN model (LSTM-BNN PredictedPower). The horizontal axis represents the time step, and the vertical axis represents the generated power.
[0183] like Figure 3 As shown in the figure, the photovoltaic power generation prediction method based on deep reinforcement learning of the present invention and the evaluation indicators of the current mainstream model include:
[0184] The mean absolute error (MAE) between all predicted values and the true values in S45, the mean square error (MSE) between all predicted values and the true values, and the coefficient of determination (R2).
[0185] like Figure 4 The figure shows a module diagram of the photovoltaic power prediction method based on deep reinforcement learning of the present invention, which includes:
[0186] A collection module is used to obtain relevant parameters of photovoltaic power generation;
[0187] Loading and merging modules, used to load and merge data sets;
[0188] The split and scale module splits the dataset into training and test sets and scales features to improve the training efficiency and generalization ability of the model;
[0189] The dual-stream encoder module extracts meteorological spatial features through convolutional neural networks and combines gated recurrent units to capture time series patterns.
[0190] Dual attention mechanism module: the time module focuses on key time points, and the weather module dynamically weights variables;
[0191] The reward module adjusts the reward weight according to the confidence of the weather forecast through an adaptive compensator;
[0192] The reinforcement learning module introduces the A3C algorithm. Combined with the above modules, the A3C algorithm uses multi-agent parallel training to continuously optimize the parameters of the neural network model to maximize the long-term cumulative reward, thereby improving the accuracy of photovoltaic power plant power prediction;
[0193] The analysis module evaluates the prediction performance by calculating the mean absolute error (MAE), mean square error (MSE) and R2 score between the predicted power and the actual power;
[0194] The above are only preferred embodiments of the present invention and are not intended to limit this application. Scholars and technicians in this field may make changes and innovations thereto. Any modification, replacement, improvement, etc. within the principles of this application shall be included in the scope of protection of this application.
Claims
1. A photovoltaic power generation prediction method based on deep reinforcement learning, characterized in that: The specific steps include: S1. Integrate photovoltaic power station power generation data and meteorological data, and process key parameters such as ambient temperature and irradiance. S2. Construct a dual-stream CNN-LSTM encoder, extract meteorological spatial features through convolutional neural networks, and combine gated recurrent units to capture time series patterns. S3. Using a dual attention mechanism, the time module focuses on key time points, and the weather module dynamically weights variables; the reward weight is adjusted according to the confidence of the weather forecast through an adaptive compensator. S4. Use the A3C algorithm to implement multi-agent parallel training and maximize long-term cumulative rewards to optimize the prediction strategy. S5. Verify on a real photovoltaic power station dataset.
2. The photovoltaic power generation prediction method based on deep reinforcement learning according to claim 1 is characterized in that: The process of integrating photovoltaic power station power generation data and meteorological data and processing key parameters such as ambient temperature and irradiance in step S1 includes the following steps: S11. Extract the dataset. The dataset is from a public dataset available at: https: / / www.kaggle.com / code / pythonafroz / solar-power-generation-forecast / input. The dataset contains 68,778 data points. These data include key metrics such as DC power, AC power, daily power generation, total power generation, ambient temperature, module temperature, and irradiance. S12. General data loading and merging: The code uses the pandas library to load power generation data and meteorological data and merges them using the DATE_TIME field. Among them, T i is the timestamp, DC i is DC power, AC i is the AC power, DY i is the daily power generation, TY i is the total power generation, T i is the ambient temperature, AT i is the module temperature, MT i is the irradiation intensity, I i is the number of data points. This means joining two datasets based on their timestamp (T). S13. Perform categorical encoding on SOURCE_KEY and convert it to SOURCE_KEY_NUMBER. Converting non-numeric features into numerical features enables the model to better process them. EN i →SKN i ∈{1,2,...,K} Among them, SK i Represents the SOURCE_KEY column, and there are (K) different sensor IDs. The categorical encoding maps each sensor ID to an integer. SKN i The value of the SOURCE_KEY_NUMBER column. S14. Data Split and Scaling: Split the dataset into training and test sets and scale the features. This improves model training efficiency and generalization. However, scaling is performed only after the data has been split. This method avoids data leakage. D train =D[1:r·N] D test =D[r·N:N] Among them, D train is the training data set, D test For the test data set, r is 0.
8. The features are then scaled to the range of [0, 1]. For each feature X, the scaled value X' can be expressed as:
3. The photovoltaic power generation prediction method based on deep reinforcement learning according to claim 2 is characterized in that: The process of constructing a dual-stream CNN-LSTM encoder in step S2, extracting meteorological spatial features through a convolutional neural network, and capturing time series patterns in combination with a gated recurrent unit includes the following steps: S21. Use CNN layers to extract local features, capture short-term correlations between adjacent time steps through convolution kernels, and expand from 6 original features to 128 abstract features. X norm ∈R B×F×T F2=CNN(X norm ) Among them, R B×F×T is the original data set, and CNN() is the convolutional neural network. S22. Use LSTM to model time series and learn long-term patterns in time series. The hidden state of each time step is used for subsequent attention mechanism. h t =LSTM(F2) Among them, LSTM() is a long short-term memory network.
4. The photovoltaic power generation prediction method based on deep reinforcement learning according to claim 3 is characterized in that: In step S3, a dual attention mechanism is used, where the time module focuses on key time points and the weather module dynamically weights variables; and the process of adjusting the reward weight according to the confidence of the weather forecast by the adaptive compensator includes the following steps: S31. A temporal attention mechanism is introduced to automatically learn different time steps. α t =Softmax(W a ·tanh(W h h t +b h )) Among them, h t is the hidden state of LSTM at time step t, α t is the temporal attention weight, and z is the time-weighted feature vector. S32. Introduce a weather feature attention mechanism to automatically learn different weather features, which enables the model to pay more attention to weather features that have a significant impact on the prediction results. β f =Softmax(W b ·tanh(W z z+b z )) Among them, β f is the feature attention weight, and s is the state vector after feature weighting. S33. Design a reward function based on relative error. This reward function can better reflect the accuracy of the prediction and encourage the model to make more accurate predictions. Among them, ∈(t) is the relative error, α is the reward scaling factor, ω1,ω2 are penalty factors, and β is the quadratic penalty coefficient.
5. The photovoltaic power generation prediction method based on deep reinforcement learning according to claim 4 is characterized in that: The process of implementing multi-agent parallel training using the A3C algorithm in step S4 and maximizing long-term cumulative rewards to optimize the prediction strategy includes the following steps: S41. Environment setup defines the environment for PV power plant prediction. It includes an observation space and an action space, and implements the reset and step methods. The reset method resets the environment state, while the step method returns information such as the next state, reward, and whether the process is complete based on the agent's action. S42. Define the A3C network, which includes the CNN layer of S21, the LSTM layer of S22, the dual attention mechanism of S31 and S32, and the Actor and Critic layers. The Actor layer outputs action probabilities, and the Critic layer outputs state value functions. S43. Define A3C worker processes. A3C defines the behavior of each worker process. Each worker process has its own local model and initially loads the parameters of the global model. The worker process interacts in the environment, collects experience (state, action, reward, etc.), calculates the return and advantage function, and updates the parameters of the global model. S44: Multi-agent parallel training. Multiple A3C_Worker processes are created and started. Each worker process independently interacts with the environment, collecting experience and updating the global model in parallel. Training ends when the maximum number of global steps is reached. S45, prediction and evaluation, calculate the mean absolute error (MAE), mean square error (MSE) and R2 score between the predicted power and the actual power to evaluate the prediction performance. MAE is the average of the absolute errors between all predicted values and the true values. MSE is the average of the squared errors between all predicted values and the true values. R 2 It is used to measure the degree to which the model explains the data, and its value range is [0, 1]. 2 The closer it is to 1, the better the model fits the data.
6. The photovoltaic power generation prediction method based on deep reinforcement learning according to claim 5, characterized in that: The process of performing verification on a real photovoltaic power station dataset in step S5 includes the following steps: This model was trained with the help of the PyTorch library. The configuration used was a P100 GPU and Windows 11. The agent's learning rate (LR) was set to 0.0001, and the discount factor (GAMMA) was set to 0.
99. Specific power generation data, irradiation intensity, ambient temperature, and module temperature were used as input. The power generation for the next 15 minutes was predicted based on the previous 6 hours of data. 80% of the dataset was used for training, and 20% for validation and analysis.
7. A photovoltaic power generation prediction method based on deep reinforcement learning, characterized in that: include: A collection module is used to obtain relevant parameters of photovoltaic power generation; Loading and merging modules, used to load and merge data sets; The split and scale module splits the dataset into training and test sets and scales features to improve the training efficiency and generalization ability of the model; The dual-stream encoder module extracts meteorological spatial features through convolutional neural networks and combines gated recurrent units to capture time series patterns. Dual attention mechanism module: the time module focuses on key time points, and the weather module dynamically weights variables; The reward module adjusts the reward weight according to the confidence of the weather forecast through an adaptive compensator; The reinforcement learning module introduces the A3C algorithm. Combined with the above modules, the A3C algorithm uses multi-agent parallel training to continuously optimize the parameters of the neural network model to maximize the long-term cumulative reward, thereby improving the accuracy of photovoltaic power plant power prediction; The analysis module evaluates the prediction performance by calculating the mean absolute error (MAE), mean square error (MSE) and R2 score between the predicted power and the actual power.
Citation Information
Cited By
Photovoltaic power generation power cross-substation prediction method based on self-attention mechanism
CN121072883A