Photovoltaic power prediction method and system based on clustering and hybrid neural network

By adopting clustering and hybrid neural network methods in photovoltaic power prediction, the problems of sample imbalance and outlier sensitivity in the prior art are solved, and higher prediction accuracy and reliability are achieved.

CN119134336BActive Publication Date: 2025-05-23SUZHOU INST OF TECH PHYSICS OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411604768.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-23
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing photovoltaic power prediction methods lack prediction accuracy and reliability when dealing with the problem of sample imbalance and outlier sensitivity in different weather types.

Method used

The clustering and hybrid neural network method is adopted to reduce the impact of weather conditions on photovoltaic power through clustering analysis, and dynamically adjust the model weights through adaptive weighting mechanisms to improve the training efficiency and performance of the model.

Benefits of technology

It improves the accuracy and reliability of photovoltaic power prediction, reduces model calculation complexity, and enhances the ability to adapt to unknown data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119134336B_ABST
    Figure CN119134336B_ABST
Patent Text Reader

Abstract

The present application discloses a photovoltaic power prediction method and system based on clustering and hybrid neural network, the method includes: preprocessing and sample division based on historical data of photovoltaic power station to be predicted, obtaining training samples and test samples of photovoltaic power station to be predicted; clustering the training samples and test samples of photovoltaic power station to be predicted, and determining clustering results; setting adaptive weights according to the distance between the feature centers of training samples and corresponding test samples in the same cluster cluster; constructing a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories according to the clustering results and adaptive weights; generating photovoltaic power prediction results of photovoltaic power station to be predicted based on the CNN-LSTM photovoltaic power prediction model. By providing a photovoltaic power prediction method based on clustering and hybrid neural network, and combining the advantages of clustering, convolutional neural network, and long short-term memory network, the aim is to improve the accuracy and reliability of photovoltaic power prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of photovoltaic technology, and in particular to a photovoltaic power prediction method and system based on clustering and hybrid neural network. Background Art

[0002] Photovoltaic power generation is one of the most accessible, low-cost and promising renewable energy sources. With the increasing energy demand in developing countries, photovoltaic power generation has increased year by year, alleviating the global energy and climate change crisis. However, due to the intermittent and random nature of photovoltaic power generation, the operation and planning of the existing power system pose challenges to the prediction accuracy of photovoltaic power. An effective photovoltaic power generation power prediction model can greatly improve the utilization rate of solar energy, and reliable prediction results can help improve system stability. Therefore, how to improve the prediction accuracy of photovoltaic power is a difficult research point and has become an urgent problem to be solved.

[0003] At present, traditional photovoltaic power prediction methods are mainly based on physical models and statistical models. Although the physical model method is simpler, its prediction accuracy is low and it is only suitable for medium- and long-term predictions. Due to the complexity and volatility of photovoltaic power generation, the statistical model method has complex and non-periodic statistical prediction models, and its performance on large-scale historical data is limited.

[0004] Compared with traditional methods, deep learning methods can better implement nonlinear mapping and feature extraction tasks, and achieve higher accuracy in photovoltaic power prediction. However, deep learning also has the problem of unbalanced samples of different weather types in photovoltaic power prediction, which makes the photovoltaic power prediction model sensitive to outliers such as noise samples, reducing the prediction ability of the model, thereby limiting the accuracy of photovoltaic power prediction. Summary of the invention

[0005] The purpose of this application is to provide a photovoltaic power prediction method and system based on clustering and hybrid neural networks to address the deficiencies in the prior art. It aims to improve the accuracy and reliability of photovoltaic power prediction by providing a photovoltaic power prediction method based on clustering and hybrid neural networks and combining the advantages of clustering, convolutional neural networks, and long short-term memory networks.

[0006] An embodiment of the present application provides a photovoltaic power prediction method based on clustering and hybrid neural network, the method comprising:

[0007] Preprocessing and sample division are performed based on historical data of the photovoltaic power station to be predicted, so as to obtain training samples and test samples of the photovoltaic power station to be predicted;

[0008] Perform clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determine the clustering results;

[0009] Set adaptive weights based on the distance between the feature centers of training samples and corresponding test samples in the same cluster;

[0010] According to the clustering results and adaptive weights, a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories is constructed;

[0011] The photovoltaic power prediction model based on CNN-LSTM generates the photovoltaic power prediction results of the photovoltaic power station to be predicted.

[0012] Optionally, the preprocessing of the historical data of the photovoltaic power station to be predicted includes:

[0013] Select historical data of the photovoltaic power station to be predicted within a preset time period;

[0014] Execute outlier removal, data missing value filling, data normalization and feature selection on the historical data of the photovoltaic power station to be predicted within the preset time period.

[0015] Optionally, clustering the training samples and test samples of the photovoltaic power station to be predicted to determine the clustering results includes:

[0016] The K-means clustering algorithm (K-means) based on the elbow method is used to cluster the training samples and test samples of the photovoltaic power station to be predicted, and the clustering results are obtained.

[0017] Optionally, after clustering the training samples and test samples of the photovoltaic power station to be predicted and determining the clustering results, the method further includes:

[0018] The weight is set based on the distance between the feature center of each training sample and the test sample in the same cluster calculated according to the clustering results, and the weight value of each sample is adaptively adjusted.

[0019] Optionally, the photovoltaic power prediction model based on CNN-LSTM includes:

[0020] A data input module, a data processing module and a data output module connected in communication, wherein:

[0021] The input data of the data input module include photovoltaic power, temperature, humidity, global horizontal radiation and global inclined radiation data of the photovoltaic power station to be predicted;

[0022] The data processing module includes a feature extraction unit, a time series learning unit and a fully connected layer, wherein the feature extraction unit includes three one-dimensional convolutional neural network (CNN) layers; the time series learning unit includes two long short-term memory neural network (LSTM) layers, wherein the return sequence of the first LSTM layer is set to true, and the return sequence of the second LSTM layer is set to false; and there are two fully connected layers after the time series learning unit;

[0023] The output data of the data output module is the predicted photovoltaic power, and the number of output neurons is determined according to the output characteristic number.

[0024] Optionally, the photovoltaic power prediction model based on CNN-LSTM adopts mean absolute error as the loss function.

[0025] Optionally, after the photovoltaic power prediction model based on CNN-LSTM generates a photovoltaic power prediction result of the photovoltaic power station to be predicted, the method further includes:

[0026] A preset method is used to perform a test and evaluation on a photovoltaic power prediction result of a photovoltaic power station to be predicted, wherein the preset method includes a root mean square error method, a mean absolute error method, and a correlation coefficient method.

[0027] Another embodiment of the present application provides a photovoltaic power prediction system based on clustering and hybrid neural network, the system comprising:

[0028] A data processing module is used to perform preprocessing and sample division based on historical data of the photovoltaic power station to be predicted, so as to obtain training samples and test samples of the photovoltaic power station to be predicted;

[0029] The sample clustering module is used to cluster the training samples and test samples of the photovoltaic power station to be predicted and determine the clustering results;

[0030] A weight acquisition module is used to set adaptive weights according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster;

[0031] Model training module, used to build a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories based on clustering results and adaptive weights;

[0032] The power prediction module is used to generate the photovoltaic power prediction results of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

[0033] Yet another embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the method described above.

[0034] Yet another embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the above-mentioned method when running.

[0035] Compared with the prior art, the present application first performs preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, and obtains the training samples and test samples of the photovoltaic power station to be predicted; then clusters the training samples and test samples of the photovoltaic power station to be predicted, and determines the clustering results; sets adaptive weights according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster cluster; constructs a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories according to the clustering results and the adaptive weights; finally, generates the photovoltaic power prediction results of the photovoltaic power station to be predicted based on the CNN-LSTM photovoltaic power prediction model. It provides a photovoltaic power prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A hardware structure block diagram of a computer terminal for a photovoltaic power prediction method based on clustering and hybrid neural network provided in an embodiment of the present application;

[0037] Figure 2 A schematic diagram of a process flow of a photovoltaic power prediction method based on clustering and hybrid neural network provided in an embodiment of the present application;

[0038] Figure 3 A schematic diagram of the Pearson correlation coefficient between the variables of the historical data provided in the embodiment of the present application;

[0039] Figure 4 It is a K-mean clustering flow chart based on the elbow method provided in an embodiment of the present application;

[0040] Figure 5 A K-means clustering curve diagram based on the elbow method provided in an embodiment of the present application;

[0041] Figure 6 A schematic diagram of a photovoltaic power prediction model architecture based on CNN-LSTM provided in an embodiment of the present application;

[0042] Figure 7 A schematic diagram of photovoltaic power prediction results provided in an embodiment of the present application;

[0043] Figure 8 A schematic diagram of a process flow of photovoltaic power prediction results of a photovoltaic power station to be predicted provided in an embodiment of the present application;

[0044] Fig. 9 A photovoltaic power prediction result curve graph using four deep learning methods provided in an embodiment of the present application;

[0045] Fig.10 A schematic diagram of the error comparison between the predicted power and the actual photovoltaic power using four deep learning methods provided in an embodiment of the present invention;

[0046] Fig.11 A schematic diagram of the structure of a photovoltaic power prediction system based on clustering and hybrid neural network provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as limiting the present application.

[0048] The embodiment of the present application provides a photovoltaic power prediction method based on clustering and hybrid neural network, which can be applied to electronic devices such as computer terminals, specifically ordinary computers, tablets, etc.

[0049] The following describes it in detail by taking running on a computer terminal as an example. Figure 1 The hardware structure block diagram of the computer terminal of the photovoltaic power prediction method based on clustering and hybrid neural network provided in the embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the computer terminal may also include a transmission device 106 for communication functions and an input and output device 108. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0050] The memory 104 can be used to store software programs and modules of application software, such as program instructions / modules corresponding to the photovoltaic power prediction method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0051] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0052] The existing deep learning methods mainly focus on spatial or temporal features, while ignoring the clustering strategy of different discriminant features and the impact of sample weights. This application establishes a CNN-LSTM network model with adaptive weights based on clustering, reduces the impact of weather conditions on photovoltaic power through clustering analysis, and uses adaptive weights to dynamically adjust the weights of the model according to real-time data during training, thereby improving the training efficiency and final performance of the model, thereby improving the accuracy and stability of photovoltaic power prediction.

[0053] See also Figure 2 , Figure 2 The schematic flow chart of the photovoltaic power prediction method based on clustering and hybrid neural network provided in the embodiment of the present application may include the following steps:

[0054] S201: Preprocessing and sample division are performed based on historical data of the photovoltaic power station to be predicted to obtain training samples and test samples of the photovoltaic power station to be predicted.

[0055] Specifically, the preprocessing of the historical data of the photovoltaic power station to be predicted may include:

[0056] Select historical data of the photovoltaic power station to be predicted within a preset time period; perform outlier removal, data missing value filling, data normalization and feature selection on the historical data of the photovoltaic power station to be predicted within the preset time period.

[0057] Data preprocessing is the initial step of data analysis, which aims to ensure the quality and consistency of data for effective model training. For the historical data of photovoltaic power plants to be predicted, preprocessing includes outlier removal, missing value filling with the average value of adjacent non-outlier data, data normalization and feature selection using the Pearson correlation coefficient.

[0058] Among them, check the missing values ​​in the data, and choose to delete records with missing values, fill missing values ​​(such as using mean, median, interpolation, etc.), or mark missing values ​​according to the specific situation. Check and delete duplicate data records to avoid the influence of duplicate information when training the model. Identify and process outliers in the data (such as power generation under extreme weather conditions), and choose to delete, replace or correct these values. You can also standardize (such as Z-score standardization) or normalize (such as Min-Max normalization) numerical features to ensure that the numerical ranges of different features have equivalent weights when training the model.

[0059] For example, publicly released data on a photovoltaic power station at location B of a solar energy center in country A can be selected to preprocess the photovoltaic power and related meteorological data, including time period selection, such as selecting data between preset time periods, removing outliers, filling missing values ​​with the average of adjacent non-outlier data, normalizing data, and performing feature selection using the Pearson correlation coefficient.

[0060] For example, the publicly released Alice Springs photovoltaic data and related weather data of the Australian Solar Energy Center are selected. The selected data are data from 25 stations from July 5, 2016 to October 6, 2023, and the data resolution is downsampled to 15 minutes. Since the quality of data seriously affects the performance of deep learning, in order to achieve accurate predictions, the data needs to be preprocessed. Preprocessing includes time period selection, outlier removal, missing value filling, data normalization and feature selection. This application selects data between 7:30-18:00, removes outliers from the selected data, and fills missing values ​​with the average of adjacent non-outlier data. After filling the missing values, data normalization and feature selection are required. This application focuses on data normalization and Pearson correlation coefficient.

[0061] Data normalization improves the overall performance by eliminating the scale differences between features, improving the model convergence speed, stability and prediction accuracy. However, the final result represents the predicted photovoltaic power generation value, which needs to be compared with the actual power generation to evaluate the prediction performance. Therefore, reverse normalization becomes particularly important. Normalization and reverse normalization can be performed using the following formula:

[0062]

[0063]

[0064] in, represents the sample value, represents the normalized sample value, , Indicates the minimum and maximum values ​​of the samples.

[0065] The Pearson correlation coefficient is used to measure the linear correlation between photovoltaic power and other variables. The value range is [-1, 1]. The stronger the correlation, the closer it is to 1 or -1, and the weaker the correlation, the closer it is to 0. A negative value indicates a negative correlation, and a positive value indicates a positive correlation. The formula is as follows:

[0066]

[0067] in, is the Pearson correlation coefficient, representing the variable and The linear correlation between is the number of samples of the variable, and The variables are and No. Sample values, yes The sample mean of yes The sample mean of . Figure 3 , Figure 3 A schematic diagram of the Pearson correlation coefficient between various variables of the historical data of the photovoltaic power station provided in the embodiment of the present application can select variables with larger absolute values ​​of the Pearson correlation coefficient in the diagram as subsequent model input variables, including five variables: photovoltaic power, temperature, humidity, global horizontal radiation and global tilted radiation. The model input includes five variables at each of the first three moments, for a total of 15 input variables.

[0068] Sample partitioning is to divide the preprocessed data set into training and test sets in order to train the model and evaluate its performance. For example, the data is stratified according to certain key features and then randomly divided within each layer. This method helps to maintain the distribution consistency of key features in the training and test sets.

[0069] S202: performing clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determining the clustering results.

[0070] Specifically, clustering the training samples and test samples of the photovoltaic power station to be predicted and determining the clustering results may include:

[0071] The K-means clustering algorithm based on the elbow method is used to cluster the training samples and test samples of the photovoltaic power station to be predicted, and the clustering results are obtained.

[0072] Specifically, the elbow method is used to determine the optimal cluster number K value, and the historical weather and photovoltaic power data are clustered by K-means, so that similar historical weather data are clustered into one category. For short-term photovoltaic power prediction, the weather characteristics and photovoltaic power at time T+1 and time T are closely related. In order to better predict the photovoltaic power at time T+1, the weather characteristics and photovoltaic power at time T are clustered, so that similar weather conditions are clustered into one cluster, which provides a basis for the subsequent clustering-based CNN-LSTM photovoltaic power prediction model.

[0073] This application uses the K-means method to achieve the goal of minimizing the distance between samples and cluster centers by optimizing the mean of data points in the iterative cluster. The clustering results of K-means clustering vary greatly when the number of clusters K is different. In order to overcome the sensitivity of K-means clustering to K, the elbow method is used to determine the optimal number of clusters. Figure 4 This is a K-mean clustering flow chart based on the elbow method provided in an embodiment of the present application. First, the optimal number of clusters, i.e., the K value, is obtained according to the elbow method; then, K data points are randomly selected as the initial cluster centers; the distance between each sample and the cluster center is calculated, and it is assigned to the cluster center with the nearest distance; finally, its cluster center is recalculated, and it is checked whether the algorithm has converged, that is, whether the change of the cluster center is less than a preset threshold, or whether the preset number of iterations has been reached. If any of the above conditions is met, the algorithm ends and the clustering result is output. If the conditions are not met, it returns to use the new cluster center to reallocate the samples until the clustering result is obtained.

[0074] Exemplarily, the present application uses the elbow method to determine the optimal number of clusters for obtaining the number of clusters of K-means clustering. By observing the elbow method curve, the number of clusters K is selected at the inflection point, that is, the "elbow" position. This position represents that after the number of clusters increases to a certain extent, the contribution of continuing to increase the number of clusters to the improvement of clustering performance is no longer significant, so the number of clusters corresponding to the inflection point is selected as the optimal number of clusters.

[0075] See also Figure 5 , Figure 5 A K-means clustering curve diagram based on the elbow method provided in the embodiment of the present application, in which the ordinate represents the clustering loss function, the abscissa represents the number of clusters, and the optimal number of clusters is found through the inflection point, such as Figure 5 As shown, when k=3, it is the elbow, so the number of clusters of the data is selected as 3, and K-means clustering is used to divide all samples into 3 clusters.

[0076] S203: Setting an adaptive weight according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster.

[0077] In an optional implementation, after clustering the training samples and test samples of the photovoltaic power station to be predicted and determining the clustering results, the method may further include:

[0078] According to the clustering results, the distance between each training sample and the feature center of the test sample is calculated, and the weight value of each sample is adaptively adjusted.

[0079] Specifically, the K-means algorithm can be used to cluster all data, and the cluster center and the cluster label to which each sample belongs can be recorded. For each training sample, calculate its distance to the cluster center of all test samples. For example, you can choose to use common distance measurement methods such as Euclidean distance and Manhattan distance. Its cluster center can be regarded as its feature center, or the training sample itself or some form of its average value, such as the mean vector, can be used as the feature center. Calculate a weight adjustment factor based on the distance, so that the closer the distance, the larger the factor, indicating that the weight increases; otherwise, the weight decreases. For example, use the reciprocal of the distance, the negative exponential function, etc. as the basis for weight adjustment. Multiply the original weight by the weight adjustment factor to obtain a new sample weight. If there is no original weight, you can set a basic weight value and then adjust it according to the adjustment factor.

[0080] Exemplarily, in order to obtain better prediction results, the distance between each training sample and the feature center of the test sample can be calculated according to the clustering results to set the weight, and the sample weight can be adaptively adjusted. In this embodiment, the average value of the feature vectors of all elements is used as the feature center of the cluster cluster. In order to obtain the best weight, the model trained in the same cluster category is biased towards the test sample of the same cluster category. The inverse of the distance between the feature center of the training sample and the test sample in the same cluster cluster is used as the weight coefficient, so that the training sample that is more similar to the test sample is given a large weight, and the training sample with little difference from the test sample is given a smaller weight.

[0081] Suppose the clustering category in the test sample is The number of samples is , This is the first test sample in this category. The characteristic vector of samples, then the cluster category is The feature center of the test sample As shown below:

[0082]

[0083] The clustering categories are No. The weight of a training sample is , the corresponding eigenvector is , The clustering category is The feature center of the test sample, the weight of the sample is as follows:

[0084]

[0085] Therefore, in the same cluster, by calculating the distance between the feature centers of the training samples and the test samples, the inverse of the distance is used as the adaptive weight, so that the training samples that are more similar to the test samples are given higher weights.

[0086] S204: According to the clustering results and the adaptive weights, a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories is constructed.

[0087] Specifically, independent CNN-LSTM networks are established according to the categories of clustering results, and the photovoltaic power prediction model of CNN-LSTM is trained and the hyperparameters are optimized using the data of the clustering categories.

[0088] According to the characteristics of CNN and LSTM, this application establishes a photovoltaic power prediction model based on clustering CNN-LSTM. Wherein, the photovoltaic power prediction model based on CNN-LSTM may include: a data input module, a data processing module and a data output module connected by communication, wherein the input data of the data input module includes the photovoltaic power, temperature, humidity, global horizontal radiation and global inclined radiation data of the photovoltaic power station to be predicted; the data processing module includes a feature extraction unit, a time series learning unit and a fully connected layer, wherein the feature extraction unit includes three one-dimensional CNN layers; the time series learning unit includes two LSTM layers, wherein the return sequence of the first LSTM layer is set to true, and the return sequence of the second LSTM layer is set to false; there are two fully connected layers after the time series learning unit; the output data of the data output module is the predicted photovoltaic power, and the number of output neurons is determined according to the output feature number.

[0089] It should be noted that the one-dimensional CNN layer extracts valuable features from the input data containing factor correlation information, and the one-dimensional convolution uses a convolution technique that improves data learning efficiency and representation modularity. The LSTM layer performs time series learning, extracting sequence pattern information as well as short-term and long-term dependencies. Therefore, the use of the CNN-LSTM hybrid deep learning model is expected to improve the accuracy of photovoltaic power prediction.

[0090] See also Figure 6 , Figure 6 A schematic diagram of the photovoltaic power prediction model architecture based on CNN-LSTM provided for an embodiment of the present application. The data input module is shown in the figure "Input", and the input data includes the photovoltaic power, temperature, humidity, global horizontal radiation and global tilt radiation data of the photovoltaic power station to be predicted. The data processing module includes a feature extraction unit, a time series learning unit and a fully connected layer, wherein the feature extraction unit is composed of three one-dimensional CNN layers, namely the Conv layer shown in the figure. Stacking multiple convolutional layers in the deep learning framework enables the initial layer to learn the low-level features of the input. Two LSTM layers are used in the time series learning unit. The return sequence of the previous LSTM layer is set to true so that the network will output the complete sequence of hidden states. The return sequence of the next LSTM layer is set to false so that the network outputs the hidden state at the last time step. After time series learning, there are two fully connected layers, namely the Dense layer shown in the figure, and the data output module is shown in the figure "Output". The number of neurons in the final output layer depends on the number of output features.

[0091] It should be noted that the activation function is used to enhance the model's ability to learn complex structures. The ReLU activation function used in this application can reduce the gradient vanishing problem and make the network more trainable. The photovoltaic power prediction model based on CNN-LSTM in this application uses the mean absolute error as the loss function and uses Adaptive moment estimation (Adam) as the optimizer. Adam is a commonly used stochastic gradient descent algorithm that combines the advantages of the adaptive gradient algorithm and the momentum method, and has a faster convergence speed and better generalization performance. The parameter settings of the CNN-LSTM network framework can be shown in Table 1 below:

[0092] Table 1: CNN-LSTM network parameter table

[0093] Network Layer Output size Number of parameters Conv (None, 13, 32) 128 Conv (None, 11, 64) 6280 Conv (None, 9, 128) 24704 Lstm (None, 9, 100) 91600 Lstm (None, 100) 80400 Dense (None, 64) 6464 Dense (None, 32) 2080 Output (None, 1) 33

[0094] It should be noted that this application tests different combinations of hyperparameters of the photovoltaic power prediction model based on CNN-LSTM to optimize their learning parameters. After multiple experiments, it was found that when the activation function is ReLu, the optimizer is Adam, the loss function is the root mean square error (RMSE), the number of iterations is 100, the batch size is 20, and the learning rate is 0.001, the photovoltaic power prediction model based on CNN-LSTM has the best performance. The first 80% of the data is used as training samples in chronological order, and the last 20% of the time samples are used as test samples to predict the photovoltaic power in the next 15 minutes. See Figure 7 , Figure 7 A schematic diagram of photovoltaic power prediction results provided in the embodiment of the present application shows the photovoltaic power prediction results for 10 consecutive days. For convenience, the photovoltaic power prediction method based on clustering and hybrid neural network proposed in the present application can be referred to as Cluster-CNN-LSTM method, where the actual power is Actual, and Figure 7 It can be seen that the power predicted by the method of the present application is very close to the actual power. And whether it is sunny or not, the Cluster-CNN-LSTM photovoltaic power prediction method has a relatively accurate prediction effect on photovoltaic power.

[0095] S205: Generate a photovoltaic power prediction result of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

[0096] For details, see Figure 8 , Figure 8A schematic flow chart of a photovoltaic power prediction result of a photovoltaic power station to be predicted is provided in an embodiment of the present application. Through the above steps S201-S204, the prediction results of each model are integrated to obtain a final prediction result.

[0097] After the photovoltaic power prediction model based on CNN-LSTM generates the photovoltaic power prediction result of the photovoltaic power station to be predicted, the method may further include:

[0098] The preset method is used to perform the inspection and evaluation of the photovoltaic power prediction results of the photovoltaic power station to be predicted, and RMSE, mean absolute error (MAE) and correlation coefficient (CC) are selected as photovoltaic power evaluation indicators.

[0099] Specifically, this application selects root mean square error, mean absolute error and correlation coefficient as photovoltaic power evaluation indicators, among which RMSE is a measure of the difference between the actual observed value and the model predicted value, which represents the average deviation between the model predicted value and the true value. MAE is the average value of the absolute error between the predicted value and the true value, which can better reflect the actual situation of the predicted value error. CC is used to measure the correlation between two variables. The closer it is to 1, the better the effect.

[0100] The root mean square error is the square root of the mean of the square of the difference between the predicted value and the true value, which reflects the degree of dispersion of the prediction error. The smaller the RMSE value, the higher the prediction accuracy of the model. The calculation formula is as follows:

[0101]

[0102] in, is the total number of samples, and is the i-th pair of actual PV power and predicted PV power values ​​being tested.

[0103] The mean absolute error is the mean of the absolute values ​​of the difference between the predicted value and the true value, which reflects the average level of prediction error. The smaller the MAE value, the smaller the prediction error of the model. The calculation formula is as follows:

[0104]

[0105] in, is the total number of samples, and is the i-th pair of actual PV power and predicted PV power values ​​being tested.

[0106] The correlation coefficient is used to measure the correlation between two variables. The calculation formula is as follows:

[0107]

[0108] in, is the total number of samples, and is the i-th pair of actual PV power and predicted PV power value being tested. and Respectively represent variables and variables The corresponding average value.

[0109] In summary, this application develops an emerging photovoltaic power intelligent prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction.

[0110] This application reduces the impact of weather conditions on photovoltaic power forecasting through clustering methods and reduces the computational complexity of the model. By clustering historical data, data with similar patterns are clustered together, and special model training and optimization are performed for each clustering category, so that the model can more accurately capture the characteristics of various types of data and achieve more accurate personalized predictions. Clustering divides the data set into multiple small-scale subsets, and the data within each subset is highly similar, so the model can be trained independently on these subsets, reducing the computational complexity of model training.

[0111] The adaptive weight mechanism proposed in this application is conducive to improving the generalization ability of the model. In the photovoltaic power prediction scenario, the variability of weather and environmental conditions may lead to large differences between the training set and the test set. The clustering-based adaptive weight mechanism can adjust the model's learning strategy according to the data distribution of different clustering categories, thereby improving the model's adaptability when facing unknown data. At the same time, the adaptive weight mechanism allows the model to dynamically adjust weights during training, avoiding complex optimization calculations on a global scale.

[0112] The core of photovoltaic power prediction is to accurately capture the complex relationship between weather factors and photovoltaic power generation. This application shows great potential in dealing with these complex relationships based on clustering and hybrid neural network prediction models. CNN has powerful feature extraction capabilities and can automatically extract representative features from high-dimensional data; LSTM is good at processing time series data and can remember long-term dependencies. Therefore, the combination of CNN and LSTM can better capture the spatiotemporal characteristics of photovoltaic power. The ultra-short-term photovoltaic power prediction method based on clustering and hybrid neural networks significantly improves the accuracy, generalization ability, computational efficiency and robustness of the prediction model by combining cluster analysis, CNN feature extraction and LSTM time series modeling, and has the potential to implement personalized predictions.

[0113] This application can help the power system to carry out more accurate energy dispatch, optimize the balance between power generation and power consumption, and improve energy efficiency. At the same time, accurate ultra-short-term photovoltaic power forecasting can optimize power trading strategies, thereby improving economic benefits, enhancing the flexibility and reliability of the power system, and providing important support for effective energy management and grid operation.

[0114] Compared with the prior art, the present application first performs preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, and obtains the training samples and test samples of the photovoltaic power station to be predicted; then clusters the training samples and test samples of the photovoltaic power station to be predicted, and determines the clustering results; according to the clustering results and adaptive weights, a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories is constructed; finally, based on the CNN-LSTM photovoltaic power prediction model, the photovoltaic power prediction results of the photovoltaic power station to be predicted are generated. It provides a photovoltaic power prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction.

[0115] It should be noted that this application discloses examples of comparative experiments of photovoltaic power prediction using four deep learning methods to prove the effectiveness of the method described in this application. For convenience, the method based on clustering and hybrid neural networks in this application is referred to as Cluster-CNN-LSTM method. This application selects three other classic deep learning methods for comparative experiments on the photovoltaic station data to be predicted, including LSTM, CNN-LSTM and GRU (gated recurrent unit, GRU) networks. Among them, GRU is a variant of LSTM. Table 2 is a comparison of photovoltaic power prediction accuracy. It can be seen from Table 2 that the photovoltaic power prediction accuracy of the method proposed in this invention is the highest, the corresponding MAE and RMSE are the smallest, the value of the correlation coefficient CC is the largest, the accuracy of the GRU method is the lowest, and relative to the CNN-LSTM method, the MAE of the method of this application is reduced by 0.071, the RMSE is reduced by 0.2, and the CC is increased by 0.023.

[0116] Table 2: Comparison of photovoltaic power prediction methods

[0117] method MAE RMSE CC LSTM 0.245 0.420 0.968 CNN-LSTM 0.216 0.413 0.969 GRU 0.255 0.429 0.967 Cluster-CNN-LSTM 0.145 0.213 0.992

[0118] See also Fig. 9 , Fig. 9 A photovoltaic power prediction result curve graph using four deep learning methods provided in the embodiment of the present application shows the comparison results of photovoltaic power prediction for 3 days. It can be seen from the figure that the curve fitting effect of the method proposed in the present application is the best for real power, followed by the CNN-LSTM method, and the prediction effect of the GRU method is the worst.

[0119] In order to more intuitively see the difference between the photovoltaic power prediction results of the four comparison methods and the actual photovoltaic power, see Fig.10 , Fig.10 A schematic diagram of the error comparison between the predicted power and the actual photovoltaic power using four deep learning methods provided in an embodiment of the present invention. As can be seen from the figure, the method of the present application is closer to the actual value and performs relatively stably, while the other three methods have significantly over- or under-predicted.

[0120] Therefore, this application proposes a photovoltaic power prediction model based on clustering CNN-LSTM. The specific mode and architectural layout of the combination of CNN and LSTM networks can better extract the spatial and temporal characteristics of the data. The data of the photovoltaic site to be predicted are used to prove the effectiveness of the method of this application. Four deep learning methods are used to predict photovoltaic power. The Cluster-CNN-LSTM method proposed in this application has the highest accuracy, followed by the CNN-LSTM method, and the GRU method is the worst. The deep learning method proposed in this application achieves high-precision prediction of photovoltaic power, which is of great significance for evaluating the performance and efficiency of photovoltaic systems.

[0121] Another embodiment of the present application provides a photovoltaic power prediction system based on clustering and hybrid neural network, such as Fig.11 The schematic diagram of the structure of the photovoltaic power prediction system based on clustering and hybrid neural network shown in FIG. 1 includes:

[0122] The data processing module 1101 is used to perform preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, so as to obtain training samples and test samples of the photovoltaic power station to be predicted;

[0123] The sample clustering module 1102 is used to perform clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determine the clustering results;

[0124] A weight acquisition module 1103 is used to set an adaptive weight according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster;

[0125] The model training module 1104 is used to construct a photovoltaic power prediction model of CNN-LSTM applicable to different clustering categories according to the clustering results and adaptive weights;

[0126] The power prediction module 1105 is used to generate a photovoltaic power prediction result of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

[0127] Compared with the prior art, the present application first performs preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, and obtains the training samples and test samples of the photovoltaic power station to be predicted; then clusters the training samples and test samples of the photovoltaic power station to be predicted, and determines the clustering results; according to the clustering results and adaptive weights, a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories is constructed; finally, based on the CNN-LSTM photovoltaic power prediction model, the photovoltaic power prediction results of the photovoltaic power station to be predicted are generated. It provides a photovoltaic power prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction.

[0128] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0129] Specifically, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0130] Specifically, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0131] S201: performing preprocessing and sample division based on historical data of the photovoltaic power station to be predicted, to obtain training samples and test samples of the photovoltaic power station to be predicted;

[0132] S202: performing clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determining the clustering results;

[0133] S203: setting an adaptive weight according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster;

[0134] S204: constructing a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories according to the clustering results and the adaptive weights;

[0135] S205: Generate a photovoltaic power prediction result of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

[0136] Compared with the prior art, the present application first performs preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, and obtains the training samples and test samples of the photovoltaic power station to be predicted; then clusters the training samples and test samples of the photovoltaic power station to be predicted, and determines the clustering results; according to the clustering results and adaptive weights, a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories is constructed; finally, based on the CNN-LSTM photovoltaic power prediction model, the photovoltaic power prediction results of the photovoltaic power station to be predicted are generated. It provides a photovoltaic power prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction.

[0137] An embodiment of the present application further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0138] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for performing the following steps:

[0139] S201: performing preprocessing and sample division based on historical data of the photovoltaic power station to be predicted, to obtain training samples and test samples of the photovoltaic power station to be predicted;

[0140] S202: performing clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determining the clustering results;

[0141] S203: setting an adaptive weight according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster;

[0142] S204: constructing a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories according to the clustering results and the adaptive weights;

[0143] S205: Generate a photovoltaic power prediction result of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

[0144] Specifically, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0145] Compared with the prior art, the present application first performs preprocessing and sample division based on the historical data of the photovoltaic power station to be predicted, and obtains the training samples and test samples of the photovoltaic power station to be predicted; then clusters the training samples and test samples of the photovoltaic power station to be predicted, and determines the clustering results; according to the clustering results and adaptive weights, a CNN-LSTM photovoltaic power prediction model suitable for different clustering categories is constructed; finally, based on the CNN-LSTM photovoltaic power prediction model, the photovoltaic power prediction results of the photovoltaic power station to be predicted are generated. It provides a photovoltaic power prediction method based on clustering and hybrid neural networks, and combines the advantages of clustering, convolutional neural networks, and long short-term memory networks, aiming to improve the accuracy and reliability of photovoltaic power prediction.

[0146] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0147] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0148] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the above-mentioned units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0149] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0151] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, read-only memories, random access memories, mobile hard disks, magnetic disks or optical disks.

[0152] The embodiments of the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A photovoltaic power prediction method based on clustering and hybrid neural network, characterized in that: The method comprises: Based on the historical data of the photovoltaic power station to be predicted, preprocessing and sample division are performed on the historical data of the photovoltaic power station to be predicted to obtain training samples and test samples of the photovoltaic power station to be predicted; Perform clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determine the clustering results; According to the clustering results, the distance between each training sample and the feature center of the test sample is calculated, and the weight value of each sample is adaptively adjusted. The inverse of the distance between the feature center of the training sample and the test sample is used as the weight coefficient; the weight of the sample is expressed as: in, Indicates that the clustering category is No. The weight of the training samples, represents the feature vector, Indicates that the clustering category is The characteristic center of the test sample satisfies , Indicates the first The feature vector of samples; Set adaptive weights based on the distance between the feature centers of training samples and corresponding test samples in the same cluster; Based on the clustering results and the adaptive weights, a photovoltaic power prediction model of CNN-LSTM suitable for different clustering categories is constructed; the photovoltaic power prediction model of CNN-LSTM includes: A data input module, a data processing module and a data output module connected in communication, wherein: The input data of the data input module include photovoltaic power, temperature, humidity, global horizontal radiation and global inclined radiation data of the photovoltaic power station to be predicted; The data processing module includes a feature extraction unit, a time series learning unit and a fully connected layer, wherein the feature extraction unit includes three one-dimensional CNN layers; the time series learning unit includes two LSTM layers, wherein the return sequence of the first LSTM layer is set to true, and the return sequence of the second LSTM layer is set to false; The number of neurons in the data output module is determined according to the output feature number; The photovoltaic power prediction model based on CNN-LSTM generates the photovoltaic power prediction results of the photovoltaic power station to be predicted.

2. The method according to claim 1, characterized in that The preprocessing of the historical data of the photovoltaic power station to be predicted includes: Select historical data of the photovoltaic power station to be predicted within a preset time period; Execute outlier removal, data missing value filling, data normalization and feature selection on the historical data of the photovoltaic power station to be predicted within the preset time period.

3. The method according to claim 2, characterized in that The performing of clustering processing on the training samples and the test samples of the photovoltaic power station to be predicted and determining the clustering results includes: The K-means clustering algorithm based on the elbow method is used to perform clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and the clustering results are obtained.

4. The method according to claim 3, characterized in that The photovoltaic power prediction model based on CNN-LSTM adopts mean absolute error as the loss function.

5. The method according to claim 4, characterized in that After the photovoltaic power prediction model based on CNN-LSTM generates the photovoltaic power prediction result of the photovoltaic power station to be predicted, the method further includes: A preset method is used to perform a test and evaluation on a photovoltaic power prediction result of a photovoltaic power station to be predicted, wherein the preset method includes a root mean square error method, a mean absolute error method, and a correlation coefficient method.

6. Photovoltaic power prediction system based on clustering and hybrid neural network, characterized in that: The system comprises: A partitioning module is used to perform preprocessing and sample partitioning on the historical data of the photovoltaic power station to be predicted based on the historical data of the photovoltaic power station to be predicted, so as to obtain training samples and test samples of the photovoltaic power station to be predicted; An execution module, used for performing clustering processing on the training samples and test samples of the photovoltaic power station to be predicted, and determining the clustering results; The adjustment module is used to calculate the distance between each training sample and the feature center of the test sample according to the clustering results, and adaptively adjust the weight value of each sample, and use the inverse of the distance between the feature center of the training sample and the test sample as the weight coefficient; wherein the weight of the sample is expressed as: in, Indicates that the clustering category is No. The weight of the training samples, represents the feature vector, Indicates that the clustering category is The characteristic center of the test sample satisfies , Indicates the first The feature vector of samples; A setting module, used for setting adaptive weights according to the distance between the feature centers of the training samples and the corresponding test samples in the same cluster; A construction module is used to construct a photovoltaic power prediction model of CNN-LSTM applicable to different clustering categories based on clustering results and adaptive weights; the photovoltaic power prediction model of CNN-LSTM includes: A data input module, a data processing module and a data output module connected in communication, wherein: The input data of the data input module include photovoltaic power, temperature, humidity, global horizontal radiation and global inclined radiation data of the photovoltaic power station to be predicted; The data processing module includes a feature extraction unit, a time series learning unit and a fully connected layer, wherein the feature extraction unit includes three one-dimensional CNN layers; the time series learning unit includes two LSTM layers, wherein the return sequence of the first LSTM layer is set to true, and the return sequence of the second LSTM layer is set to false; The number of neurons in the data output module is determined according to the output feature number; The generation module is used to generate the photovoltaic power prediction results of the photovoltaic power station to be predicted based on the photovoltaic power prediction model of CNN-LSTM.

7. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 5 when executed.