Photovoltaic power generation power prediction method and system
By using a dual-channel neural network model with dynamic time constants in the prediction of photovoltaic power generation, the problem that models in the prior art are difficult to capture complex patterns and laws of photovoltaic power data is solved, and higher prediction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510128688.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
The existing photovoltaic power prediction methods are limited in capturing complex data patterns and rules, and are difficult to adapt to sudden fluctuations and periodic changes in photovoltaic power data. The model structure is not flexible enough and has low applicability.
A two-channel neural network model containing dynamic time constants is adopted, combined with a recurrent neural network and a gated cyclic unit, and the adaptability and prediction accuracy of the model are improved through data preprocessing and hyperparameter optimization.
By introducing dynamic time constants, the model can better capture the timing characteristics of photovoltaic power data, improve prediction accuracy and model adaptability, and reduce the consumption of computing resources.
Smart Images

Figure CN120069194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photovoltaic power generation prediction, and particularly to a photovoltaic power generation prediction method and system. Background Art
[0002] Minute-level photovoltaic power generation prediction is of great significance for planning power grid dispatching, improving energy utilization efficiency and economic benefits. Existing prediction methods mainly adopt traditional machine learning models such as support vector machines (SVM), or deep learning technologies such as long short-term memory networks (LSTM), etc. These methods can better predict the short-term fluctuations and long-term periodic changes of photovoltaic power by learning the non-linear features and time series relationships in photovoltaic power data.
[0003] The existing prediction methods have the following deficiencies: First, these methods usually assume that the time constant is fixed, that is, the default recurrent neural network or gated recurrent unit structure takes 1 as the time constant, and it is unable to adapt to the sudden fluctuations and periodic changes in photovoltaic power data. Second, the existing methods are relatively limited in feature extraction ability, and it is difficult to capture the complex patterns and rules in photovoltaic power data, and the improvement space of minute-level prediction accuracy is limited. In addition, the machine learning models in the existing methods usually adopt a pre-constructed method, and the hyperparameters will be optimized, but the structure usually will not be adjusted according to different scenarios, resulting in a reduced applicability of the model. At the same time, the existing technology does not consider the deployment and update strategies of the photovoltaic power generation prediction model in the immediate application scenario.
[0004] The above problems limit the accuracy and practicality of the existing methods in photovoltaic power generation prediction, and it is difficult to meet the requirements of prediction accuracy and model adaptability in actual applications.
[0005] Therefore, it is very necessary to propose a photovoltaic power generation prediction method and system to improve the accuracy of photovoltaic power generation prediction. Summary of the Invention
[0006] The purpose of the present invention is to provide a photovoltaic power generation prediction method and system to improve the accuracy of wind power prediction.
[0007] To achieve the above technical purpose, the present invention provides a photovoltaic power generation prediction method, which is characterized by including:
[0008] Data preprocessing: cleaning, enhancing and feature screening of photovoltaic power generation data;
[0009] Constructing a dual-channel neural network model including a dynamic time constant, the dual-channel neural network model includes:
[0010] The first channel, a recurrent neural network including a dynamic time constant,
[0011] The second channel, including a gated recurrent unit with a dynamic time constant;
[0012] Optimize the hyperparameters of the dual-channel neural network model;
[0013] Perform model training;
[0014] Use the trained model for prediction and obtain the predicted value.
[0015] Optionally, the steps of the data preprocessing include:
[0016] Collect environmental factors and component data at a preset time step and combine them to form an original data set;
[0017] Fill in missing values and process outliers in the original data set;
[0018] Perform feature screening based on mutual information analysis and correlation analysis.
[0019] Optionally, the first channel adds a time constant τ to the recurrent neural network t , the unit state is represented as n t , and is connected to the input x through the weight matrix W τ and is expressed as: t s t-1 τ
[0020] τ t = τ(concat(x t , s t-1 )W τ );
[0021] In the formula, concat represents the concatenation function, which directly concatenates the input vector x t and the previous state vector s t-1 to form a new vector, and the length is the sum of the concatenated vectors;
[0022] The update of the current state s t is expressed as:
[0023]
[0024] In the formula, s t represents the current state, s t-1 represents the previous state, and n t is the unit state.
[0025] Optionally, the second channel adds a time constant τ that changes with the unit state and input to the gated recurrent unit, and retains the reset gate r of the GRU (Gated Recurrent Unit) t and the update gate zt Calculation of, retaining candidate hidden states Calculation of;
[0026] Among them, the calculation of adding a dynamic time constant is expressed as:
[0027] τ t = τ(x t W τ + s t-1 U τ + b τ );
[0028] In the formula, W τ and U τ are the weight matrices connected to the current time step x t and the previous step state s t respectively, and b t-1 is the bias value for time constant update; τ
[0029] Meanwhile, the time constant participates in the update of the new state s t :
[0030]
[0031] Optionally, the process of optimizing the hyperparameters of the dual-channel neural network model includes:
[0032] Using the Bayesian optimization algorithm to optimize the hyperparameters, where the hyperparameters include the number of hidden layer units, the number of residual blocks, the learning rate, and the L2 regularization factor of DTC-GRU (Dynamic Time Constant Gated Recurrent Unit) and DTC-RNN (Dynamic Time Constant Recurrent Neural Network).
[0033] Optionally, the process of model training includes:
[0034] Pre-encoding the input data through DTC-RNN and DTC-GRU respectively;
[0035] Feature extraction from the pre-encoded features;
[0036] Concatenating the features extracted from the two channels;
[0037] The concatenated features pass through the fully connected layer and the fitting layer to obtain the predicted value;
[0038] Updating the weight parameters.
[0039] Optionally, after prediction, model deployment is also included, including cloud platform deployment or local deployment methods.
[0040] Optionally, after deployment, model update is also included:
[0041] Regularly collect new data and perform distribution change detection, where the distribution change detection includes analyzing the mean, variance, and distribution map of the data;
[0042] If the detected data distribution change amplitude is less than a preset threshold, use incremental learning to fine-tune the model;
[0043] If the detected data distribution change amplitude is greater than or equal to the preset threshold, add the new data to the original training set and re-perform Bayesian optimization and training;
[0044] Use blue-green deployment or rolling update strategy for model update.
[0045] The present invention also provides a photovoltaic power prediction system, including:
[0046] A data acquisition module for acquiring data related to photovoltaic power generation;
[0047] A preprocessing module for data cleaning, enhancement, and feature screening;
[0048] A prediction module including a dual-channel neural network model;
[0049] A model training module for implementing model training;
[0050] An optimization module for implementing hyperparameter optimization.
[0051] Optionally, the dual-channel neural network model includes:
[0052] An input and pre-coding area for receiving preprocessed input data and performing pre-coding through DTC-GRU and DTC-RNN respectively to obtain a first initial feature vector and a second initial feature vector;
[0053] A DTC-GRU residual channel connected to the input and pre-coding area for receiving the first initial feature vector and outputting a first feature vector through N Residual-GRU (Residual Gated Recurrent Unit) residual blocks each containing a DTC-GRU and a batch normalization structure;
[0054] The DTC-RNN residual channel, connected to the input and pre-coding area, is used to receive the second initial feature vector and output a second feature vector through N Residual-RNN (Residual Recurrent Neural Network) residual blocks each containing a DTC-RNN and a batch normalization structure;
[0055] The splicing and output area, connected to the DTC-GRU residual channel and the DTC-RNN residual channel respectively, is used to splice the first feature vector and the second feature vector to obtain a fused feature vector, and process the fused feature vector through a fully connected layer and a fitting layer to obtain a photovoltaic power prediction value.
[0056] Optionally, both the DTC-GRU residual channel and the DTC-RNN residual channel include a residual block structure, and the residual block includes: a batch normalization layer, a dynamic time constant neural network layer, and a skip connection structure;
[0057] The input end of the skip connection structure is respectively connected to the input end of the batch normalization layer and the output end of the dynamic time constant neural network layer;
[0058] The output end of the skip connection structure serves as the output end of this residual block and is connected to the input end of the next residual block or the splicing and output area.
[0059] Optionally, it further includes an update module, and the update module is used for:
[0060] Regularly collect new data for distribution change detection;
[0061] Select an update method of incremental learning or re-training according to the detection result;
[0062] Execute model update and deployment.
[0063] Optionally, it further includes a deployment module for realizing the deployment of the model in an actual scenario.
[0064] Compared with the prior art, the present invention has at least the following beneficial effects:
[0065] By introducing a dynamic time constant into the recurrent neural network and the gated recurrent unit, the sensitivity of the model to sudden fluctuations in photovoltaic power is enhanced, enabling the model to better capture the temporal characteristics of the data.
[0066] Using Bayesian optimization to dynamically adjust the model structure enables the model depth to match the complexity of different scenarios. By optimizing parameters such as the number of hidden layer units and the number of residual blocks, the computational efficiency is improved while ensuring the prediction accuracy, saving computational resources.
[0067] A complete deployment plan and update strategy are provided, and the maintainability and scalability of the system are improved through modular design. The model can continuously adapt to the changing characteristics of data and has strong engineering value in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a flowchart of the steps of the photovoltaic power prediction method in the embodiment of the present invention;
[0069] Figure 2A It is a schematic diagram of the correlation analysis ranking in the embodiment of the present invention;
[0070] Figure 2B It is a schematic diagram of the mutual information analysis ranking in the embodiment of the present invention;
[0071] Figure 3A It is a schematic diagram of the comparison between the DTC-RNN prediction value and the true value in the embodiment of the present invention;
[0072] Figure 3B It is a schematic diagram of the comparison between the prediction value and the true value of the Dynamic Time Constant RNN and Gated Recurrent Unit (DTC-RNNGRU) in the embodiment of the present invention;
[0073] Figure 3C It is a schematic diagram of the comparison between the prediction value and the true value of the Dynamic Time Constant RNN Networks (DTC-RNNets) in the embodiment of the present invention;
[0074] Figure 3D It is a schematic diagram of the comparison between the LSTM prediction value and the true value in the embodiment of the present invention;
[0075] Figure 3E It is a schematic diagram of the comparison between the GRU prediction value and the true value in the embodiment of the present invention;
[0076] Figure 3F It is the Categorical Boosting (Catboost) in the embodiment of the present invention
[0077] A schematic diagram of the comparison between the prediction value and the true value;
[0078] Figure 4 It is a schematic diagram of the comparison between the predictions of the LSTM, DTC-RRNets and DTC-RNN models in the embodiment of the present invention;
[0079] Figure 5 It is a schematic diagram of the structure of the dual-channel neural network model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0080] The following will describe a photovoltaic power prediction method and system of the present invention in more detail with reference to the accompanying drawings, in which the preferred embodiments of the present invention are shown. It should be understood that those skilled in the art can modify the present invention described herein while still achieving the advantageous effects of the present invention. Therefore, the following description should be understood as a broad guidance for those skilled in the art and not as a limitation to the present invention.
[0081] In the following paragraphs, the present invention will be described more specifically by way of example with reference to the accompanying drawings. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the objectives of the embodiments of the present invention.
[0082] The photovoltaic power prediction method of the present invention includes four main steps: data preprocessing, model construction, optimization training, and prediction deployment.
[0083] Specifically, as Figure 1 shown, the present invention provides a photovoltaic power prediction method, including the following steps:
[0084] S1. Data preprocessing: cleaning, enhancing, and feature screening of photovoltaic power generation data;
[0085] S2. Construct a dual-channel neural network model including a dynamic time constant, and the dual-channel neural network model includes:
[0086] The first channel, a recurrent neural network including a dynamic time constant,
[0087] The second channel, a gated recurrent unit including a dynamic time constant;
[0088] S3. Optimize the hyperparameters of the dual-channel neural network model;
[0089] S4. Conduct model training;
[0090] S5. Use the trained model for prediction and obtain the predicted value.
[0091] Specifically, in step S1, the data preprocessing step specifically includes: collecting environmental factor data at a 5-minute step length, including temperature, humidity, wind speed, irradiance, etc., and at the same time collecting component data through a power conditioning system (PCS) inverter, including current, power, voltage, etc. Using power as the dataset label and other data as features, a raw dataset is formed by combination.
[0092] The data resolution can be adjusted to 1 - 15 minutes according to the actual situation.
[0093] For the missing data in the original dataset, the method of filling with the adjacent mean value is adopted for processing. The specific calculation formula is:
[0094]
[0095] Among them, x i is the value to be replaced, n is the number of samples, and j represents the position of the adjacent non-empty numerical value relative to x i .
[0096] Subsequently, the 3σ principle is used to identify outliers, and the outliers are also replaced by the adjacent mean value method.
[0097] Furthermore, the data preprocessing step specifically includes data augmentation: extracting the month, day, hour, and minute in the timestamp separately as new features.
[0098] Furthermore, the data preprocessing step specifically includes feature screening: simultaneously using mutual information analysis and correlation analysis rankings, and each forming a feature ranking with a weight of 50%, and eliminating the features with low feature rankings.
[0099] Please refer to Figure 2A - Figure 2B , in a specific example, after cleaning the data, the dependence relationship between variables is explored through mutual information analysis, and the correlation degree between variables is measured in combination with the Pearson correlation coefficient. Both of these analysis methods are given a weight of 50% in the ranking, and finally the five features with the lowest rankings are eliminated: hail accumulation, rainfall, minute, month, and wind direction. The remaining features include radiation, hour, two thermosensitive temperature data, ambient temperature, number of days, air pressure, wind speed, and maximum wind speed.
[0100] In step S2, the construction of the dual-channel neural network model includes:
[0101] The first channel uses a recurrent neural network with a dynamic time constant (DTC-RNN), where:
[0102] The calculation of the time constant τ t is expressed as:
[0103] τ t = τ(concat(x t , s t-1 )W τ );
[0104] In the formula, concat represents the concatenation function, which directly concatenates the input vector x t and the previous step state vector s t-1 to form a new vector, and the length is the sum of the vectors to be concatenated.
[0105] The current state s tThe update is represented as:
[0106]
[0107] In the formula, s t represents the current state, s t-1 represents the previous state, and n t is the unit state.
[0108] The second channel uses a gated recurrent unit with a dynamic time constant (DTC-GRU), which retains the original reset gate r t and update gate z t as well as the calculation of the candidate hidden state , and at the same time adds a dynamic time constant:
[0109] τ t = τ(x t W τ + s t-1 U τ + b τ );
[0110] In the formula, W τ and U τ are the weight matrices connecting τ t to the current time step x t and the previous state s t-1 respectively, and b τ is the bias value for time constant update.
[0111] At the same time, the time constant participates in the update of the new state s t :
[0112]
[0113] Furthermore, the optimization of the model uses the Bayesian optimization algorithm to optimize the hyperparameters.
[0114] It should be noted that hyperparameters are parameters that need to be set in advance before training a machine learning model. They are used to control the learning process and overall architecture of the model. Different from the parameters automatically learned by the model during training, hyperparameters need to be set manually or determined through optimization algorithms. Hyperparameters directly affect the performance and training effect of the model. Through the optimal combination of hyperparameters, better model performance can be obtained.
[0115] In this embodiment, the hyperparameters include the number of hidden layer units, the number of residual blocks, the learning rate, and the L2 regularization factor of DTC-GRU and DTC-RNN. The optimized hyperparameters are shown in Table 1:
[0116] Table 1 Bayesian optimization parameter space and optimal results
[0117]
[0118] As shown in Table 1, the most suitable model structure parameters are obtained through Bayesian optimization.
[0119] Furthermore, in step S3, the model training process specifically includes:
[0120] S31. Perform pre - encoding on the input data through the DTC - RNN and DTC - GRU;
[0121] S32. Extract features through residual blocks;
[0122] S33. Concatenate the features of the two channels;
[0123] S34. Obtain the predicted value through the fully - connected layer and the fitting layer;
[0124] S35. Use the Adam (Adaptive Moment Estimation) optimizer to update the weight parameters.
[0125] It should be noted that in step S33, concat represents the concatenation function, which refers to the operation of merging the feature vectors of different layers or channels in a neural network.
[0126] Specifically, in the input and pre - encoding area, the input is the pre - processed data set after removing the power sequence, and is pre - encoded into feature vectors of two channels with specified dimensions (equal to their respective number of hidden units) through the DTC - GRU and DTC - RNN respectively.
[0127] Please refer to Figure 5 , the input of each channel undergoes deep feature extraction through a number of residual blocks, and the number N of residual blocks is determined by Bayesian optimization as a structural parameter. The residual blocks of the two channels each output feature vectors, which are directly concatenated and connected to the fitting regression layer through the fully - connected layer.
[0128] The above is the information flow of the data in a single learning. During actual training, the model undergoes several learning processes and backpropagation corrections of the weights (automatically performed by the Adam optimizer) to obtain the optimal parameter weights and complete the training.
[0129] Furthermore, use the trained model to make predictions and obtain the predicted values.
[0130] To verify the robustness and feature extraction ability of the model, weather clustering was not performed on the dataset to simulate the process of continuous dynamic prediction in actual use. Multiple benchmark models were used for comparison, including the Dynamic Time Constant Recurrent Neural Network (DTC-RNN), LSTM, GRU, Catboost, and a persistence model used only for calculating SS-RMSE. In addition, a two-channel Dynamic Time Constant Recurrent Neural Network model without residual blocks (TCR-RNNGRU) was used as a comparison model. Table 2 shows the performance of the DTC-RRNets proposed in the present invention and five other comparison models (including DTC-GRU, DTC-RNN, LSTM, Catboost, and DTC-RNNGRU) in the same test set. From the data, it can be seen that DTC-RRNets has obvious advantages in all indicators and is the best-performing model among all models.
[0131] Please refer to Figure 3A - Figure 3F , DTC-RRNets outperforms the benchmark models in all evaluation indicators. Specifically, its MAE is significantly reduced to 11.58, showing an obvious advantage compared with other models. At the same time, its SS-RMSE is improved by 23% compared with the persistence model. However, due to overfitting, the prediction performance of the benchmark models LSTM and GRU lags behind and is closer to the persistence model, resulting in poor performance in SS-RMSE. To further explore the performance of the model in prediction, data for two consecutive days were selected for detailed comparison, as Figure 3A - Figure 3F shown. When comparing the three models with dynamic time constants, the prediction performance of DTC-RNN is poor, especially in the case of low values, and the fluctuations and deviations between the predicted values and the actual values are particularly significant. DTC-RNNGRU fails to improve the problems of DTC-RNN and even shows a phenomenon of low-value leakage, and the peak fluctuations at night are more obvious. The models without dynamic time constants (i.e., LSTM and GRU) fit well with the actual values according to the curve fitting effect, but from the analysis of the SS-RMSE evaluation index, it is found that these two models overfit the power value of the previous moment, resulting in RMSE being almost the same as that of the persistence model. The Catboost model performs stably but has slightly lower accuracy. By adding residual blocks, DTC-RRNets effectively solves the problem of low-value fluctuations, and its predicted values can closely follow the actual value curve, avoiding overfitting of the power value of the previous moment and achieving a significant improvement compared with the persistence model.
[0132] Furthermore, the stability of the model was analyzed. Data for two consecutive days, one week after the model was trained, were used to compare the performance of LSTM, DTC-RNN, and DTC-RRNets. The data are shown in Table 3 and Figure 4 shown.
[0133] Figure 4The prediction curves of three models are shown. Among them, the LSTM network shows obvious peak loss and accuracy decline, and serious distortion occurs when tracking power fluctuations. Due to the introduction of a dynamic time constant, the performance of DTC-RNN is better than that of LSTM. DTC-RRNets shows the best prediction performance, demonstrates excellent fluctuation tracking ability, and its long-term stability is significantly better than that of LSTM and DTC-RNN.
[0134] It can be found that DTC-RRNets keeps close to the actual value curve and can well track the power fluctuations. Although the evaluation index of the LSTM model used for comparison is close to that of the model of the present invention in the complete test set, the peak information is obviously lost in the prediction curve here, and there are obvious deviations, so it cannot track the power fluctuations.
[0135] Table 2: Comparison table of mean absolute error, root mean square error, root mean square error of skill score, and coefficient of determination between DTC-RRNets and other comparison models in the test set
[0136]
[0137] Table 3: Comparison table of the prediction performance of DTC-RRNets and other comparison models for two consecutive days 7 days after the prediction starting point
[0138]
[0139] Furthermore, after step S5, there is also step S6, model deployment.
[0140] In a specific example, the deployment method can be selected as the cloud platform method or the local deployment method.
[0141] The cloud platform deployment process specifically includes: selecting a cloud service provider and creating a virtual machine instance, configuring network security groups and firewall rules, packing the dependent environment with Docker containers and uploading it, configuring the API gateway and conducting tests, and controlling resource occupancy with monitoring tools.
[0142] The local deployment process specifically includes: selecting engineering hardware with a CPU or GPU, setting necessary security measures, converting the model into the TensorFlow Lite format, configuring the model and ensuring data circulation, and setting up a monitoring system to track the running status.
[0143] Furthermore, after step S6, there is also step S7, model update.
[0144] The update strategy of the model is: parameter update is performed in the upper computer with a CPU / GPU at a predetermined time interval.
[0145] In another specific example, if the conditions are not met, the update is performed through the cloud service method.
[0146] Furthermore, during the update, distribution change detection needs to be performed to analyze whether there are significant changes in the data distribution. The distribution change detection includes analyzing the mean, variance, and distribution graph of the data.
[0147] Specifically, in step S7, the distribution change detection includes: calculating the change in statistical characteristics of the newly collected data relative to the original data set, including the changes in the mean, variance, and distribution graph.
[0148] The preset thresholds for data distribution changes include: the relative change threshold of the mean is in the range of 5% - 20%, preferably 10%; the relative change threshold of the variance can be set in the range of 10% - 25%, preferably 15%; the KL divergence threshold of the distribution graph can be set in the range of 0.3 - 0.8, preferably 0.5.
[0149] When the change in any statistical characteristic exceeds the corresponding threshold, it is determined that a significant change has occurred in the data distribution. If there is no obvious change, the incremental learning method is used for fine-tuning. If there is a significant change, Bayesian optimization and training are performed again. That is, if the change detection determines that there is no obvious distribution change, the incremental learning method is used to complete the fine-tuning in the host computer; if the change detection finds a significant distribution change, the original training data set is combined with the newly collected data to form a new data set for the Bayesian optimization process; the blue-green deployment or rolling update strategy is used to deploy and replace the new DTC-RRNets.
[0150] In practical applications, appropriate thresholds can be selected within the above ranges according to specific scenarios and requirements.
[0151] In a specific example, the preset thresholds for data distribution changes include: the relative change threshold of the mean is set to 10%, the relative change threshold of the variance is set to 15%, and the KL divergence threshold of the distribution graph is set to 0.5.
[0152] Where:
[0153] 1. The relative change of the mean is calculated by (new data mean - original data mean) / original data mean;
[0154] 2. The relative change of the variance is calculated by (new data variance - original data variance) / original data variance;
[0155] 3. The change in the distribution graph is measured by calculating the KL divergence between the probability distributions of the old and new data.
[0156] When the change of any statistical feature exceeds the corresponding threshold, it is determined that the data distribution has changed significantly, and new data needs to be added to the original training set to re - perform Bayesian optimization and training; if the changes of all statistical features do not exceed the corresponding thresholds, an incremental learning method is used for model fine - tuning.
[0157] The prediction method provided by the present invention can better adapt to the sudden fluctuations and periodic changes of photovoltaic power data by introducing a dynamic time constant, significantly improving the prediction accuracy. The depth of the model can be dynamically optimized according to the scene complexity, improving the calculation efficiency and saving computing resources while ensuring the accuracy. A complete deployment plan and update strategy are provided, and engineers can regularly update the model without imposing an extra burden on users during the usage phase. The system realizes the integrated operation of data pre - processing, model adjustment, optimization, and prediction.
[0158] Embodiment 2
[0159] The photovoltaic power generation prediction system provided by the present invention includes the following modules:
[0160] The data acquisition module is used to collect data related to photovoltaic power generation.
[0161] Specifically, it includes environmental data such as temperature, humidity, wind speed, and irradiance collected by environmental sensors, as well as component data such as current, power, and voltage transmitted by the PCS inverter.
[0162] The pre - processing module is used for data cleaning, enhancement, and feature screening.
[0163] This module first fills in missing values and processes outliers for the collected original data, and then performs feature screening based on mutual information analysis and correlation analysis.
[0164] The prediction module contains a dual - channel neural network model, as Figure 5 shown, specifically including:
[0165] The input and pre - coding area is used to receive the pre - processed input data, and perform pre - coding through DTC - GRU and DTC - RNN respectively to obtain the first initial feature vector and the second initial feature vector.
[0166] The DTC - GRU residual channel is connected to the input and pre - coding area, used to receive the first initial feature vector, and output the first feature vector through N Residual - GRU residual blocks containing DTC - GRU and batch normalization structures.
[0167] The DTC-RNN residual channel, connected to the input and pre-coding area, is used to receive the second initial feature vector and output the second feature vector through N Residual-RNN residual blocks each containing a DTC-RNN and a batch normalization structure.
[0168] The concatenation and output area, respectively connected to the DTC-GRU residual channel and the DTC-RNN residual channel, is used to concatenate the first feature vector and the second feature vector to obtain a fused feature vector, and process the fused feature vector through a fully connected layer and a fitting layer to obtain the photovoltaic power prediction value.
[0169] Among them, the residual block structure in the DTC-GRU residual channel and the DTC-RNN residual channel includes:
[0170] A batch normalization layer, used to perform standardized processing on the input features.
[0171] A dynamic time constant neural network layer, used for feature transformation.
[0172] A skip connection structure, used to fuse the input features and the transformed features.
[0173] The model training module is used to implement model training. This module realizes the following functions:
[0174] Obtain the initial feature vector through pre-coding, extract features through residual blocks, concatenate the dual-channel features and output the prediction value, and implement backpropagation to update the model parameters.
[0175] The optimization module is used to implement hyperparameter optimization. This module adopts the Bayesian optimization algorithm to optimize the following parameters: the number of hidden layer units, the number of residual blocks, the learning rate, and the L2 regularization factor.
[0176] The update module is used to periodically collect new data for distribution change detection, calculate the statistical feature changes of the newly collected data relative to the original data set, including the changes in mean, variance, and distribution diagram; and select the corresponding update strategy according to the detection results:
[0177] If it is detected that the data distribution change amplitude is less than the preset threshold, adopt the incremental learning method for model fine-tuning; if it is detected that the data distribution change amplitude is greater than or equal to the preset threshold, retrain, perform model update and deployment.
[0178] The deployment module is used to implement the deployment of the model in the actual scenario. This module provides two deployment methods: cloud platform deployment method and local deployment method.
[0179] In this embodiment, each module of the system works collaboratively to achieve a complete functional chain from data acquisition to prediction output, with strong engineering practicability. The modular design of the system enables each functional unit to be independently maintained and updated, improving the maintainability and scalability of the system.
[0180] The present invention provides a photovoltaic power prediction method and system based on adaptive machine learning; the method improves the traditional recurrent neural network and gated recurrent unit by introducing a dynamic time constant, constructs a dual-channel prediction model, and uses the Bayesian optimization algorithm to optimize the model structure and hyperparameters; the system adopts a modular design to achieve a complete functional chain from data acquisition to prediction output.
[0181] The present invention has the following beneficial effects:
[0182] Improved prediction accuracy: By introducing a dynamic time constant into the recurrent neural network and gated recurrent unit, the sensitivity of the model to sudden fluctuations in photovoltaic power is enhanced, enabling the model to better capture the temporal characteristics of the data.
[0183] Improved model adaptability: The Bayesian optimization is used to dynamically adjust the model structure, enabling the model depth to match the complexity of different scenarios. By optimizing parameters such as the number of hidden layer units and the number of residual blocks, the computational efficiency is improved while ensuring the prediction accuracy, saving computational resources.
[0184] Strong engineering practicability: A complete deployment plan and update strategy are provided, and the maintainability and scalability of the system are improved through modular design. The system supports two deployment methods, namely cloud platform and local, and a dynamic update mechanism based on data distribution detection is designed, enabling the model to continuously adapt to the changing characteristics of the data and having strong engineering value in practical applications.
[0185] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A photovoltaic power generation power prediction method, characterized in that: include: Data preprocessing: cleaning, enhancement and feature screening of photovoltaic power generation data; A dual-channel neural network model including a dynamic time constant is constructed, wherein the dual-channel neural network model includes: The first channel contains a recurrent neural network with a dynamic time constant, The second channel contains a gated recurrent unit with a dynamic time constant; Optimizing the hyperparameters of the dual-channel neural network model; Conduct model training; Use the trained model to make predictions and get the predicted values.
2. The prediction method according to claim 1, characterized in that: The data preprocessing step includes: Collect environmental factors and component data at preset time steps and combine them to form the original data set; Fill in missing values and process outliers in the original data set; Feature screening is performed based on mutual information analysis and correlation analysis.
3. The prediction method according to claim 1, characterized in that: The first channel is composed of a recurrent neural network with a time constant τ t , the unit state is represented by n t , through the weight matrix W τ Connect input x t With the previous state s t--1 And expressed as: τ t =τ(concat(x t ,s t-1 )W τ ); In the formula, concat represents the concatenation function, which converts the input vector x t and the previous state vector s t-1 Directly concatenate to form a new vector, whose length is the sum of the concatenated vectors; Current Status t The update is expressed as: In the formula, s t Indicates the current state, s t--1 Indicates the previous state, n t The unit status.
4. The prediction method according to claim 1, characterized in that: The second channel is composed of a gated recurrent unit with a time constant τ that changes with the unit state and input, retaining the reset gate r of the GRU. t and update gate z t Calculation of the candidate hidden state Calculation of Among them, the calculation of adding the dynamic time constant is expressed as: t t =τ(x t W τ +s t-1 U τ +b τ ); Where W τ and U τ They are τ t Connect to the current time step x t and the previous state S t-1 The weight matrix, b τ is the bias value for time constant update; At the same time, the time constant involved in the new state s t Update:
5. The prediction method according to claim 1, characterized in that: The process of optimizing the hyperparameters of the dual-channel neural network model includes: The Bayesian optimization algorithm is used to optimize the hyperparameters, including the number of hidden layer units, the number of residual blocks, the learning rate, and the L2 regularization factor of DTC-GRU and DTC-RNN.
6. The prediction method according to claim 1, characterized in that: The process of model training includes: Pre-encode the input data through DTC-RNN and DTC-GRU respectively; Perform feature extraction on the pre-encoded features; Concatenate the features extracted from the two channels; The concatenated features are passed through the fully connected layer and the fitting layer to obtain the predicted value; Update the weight parameters.
7. The prediction method according to claim 1, characterized in that: After making predictions, the model is deployed, either on a cloud platform or locally.
8. The prediction method according to claim 7, characterized in that: After deployment, the model update is also included: Regularly collect new data and perform distribution change detection, wherein the distribution change detection includes analyzing the mean, variance and distribution diagram of the data; If it is detected that the data distribution change is less than the preset threshold, the incremental learning method is used to fine-tune the model; If it is detected that the data distribution change is greater than or equal to the preset threshold, the new data is added to the original training set and Bayesian optimization and training are performed again; Use blue-green deployment or rolling update strategies to update models.
9. A photovoltaic power generation power prediction system, characterized in that: include: Data acquisition module, used to collect photovoltaic power generation related data; Preprocessing module for data cleaning, enhancement and feature screening; Prediction module, including a dual-channel neural network model; Model training module, used to implement model training; Optimization module, used to achieve hyperparameter optimization.
10. The prediction system according to claim 9, characterized in that The dual-channel neural network model includes: An input and precoding area, used to receive the preprocessed input data, and precode it through DTC-GRU and DTC-RNN respectively, to obtain a first initial feature vector and a second initial feature vector respectively; A DTC-GRU residual channel connected to the input and precoding area, configured to receive the first initial feature vector and output the first feature vector through N Residual-GRU residual blocks including DTC-GRU and batch normalization structures; A DTC-RNN residual channel connected to the input and precoding area, configured to receive the second initial feature vector, and output the second feature vector through N Residual-RNN residual blocks including DTC-RNN and batch normalization structures; The splicing and output area is connected to the DTC-GRU residual channel and the DTC-RNN residual channel respectively, and is used to splice the first feature vector and the second feature vector to obtain a fused feature vector, and process the fused feature vector through a fully connected layer and a fitting layer to obtain a photovoltaic power prediction value.
11. The prediction system according to claim 10, characterized in that: The DTC-GRU residual channel and the DTC-RNN residual channel both include a residual block structure, and the residual block includes: a batch normalization layer, a dynamic time constant neural network layer, and a skip connection structure; The input end of the jump connection structure is connected to the input end of the batch normalization layer and the output end of the dynamic time constant neural network layer respectively; The output end of the skip connection structure serves as the output end of the residual block and is connected to the input end of the next residual block, or is connected to the splicing and output area.
12. The prediction system according to claim 9, characterized in that It also includes an update module, wherein the update module is used to: Regularly collect new data for distribution change detection; Select incremental learning or retraining update method based on the test results; Perform model updates and deployments.
13. The prediction system according to claim 9, characterized in that It also includes a deployment module to implement the deployment of the model in actual scenarios.
Citation Information
Cited By
Power grid dispatching method, device and equipment based on BiTrackGCT model and medium
CN120710005A