Photovoltaic power generation prediction system and method considering selection of similar days and weather types

By combining the evaluation criteria with similar numerical and similar shapes in the prediction of photovoltaic power generation, the problem of inaccurate selection of similar dailys in the existing technology is solved, and accurate prediction under different weather conditions is achieved, and the operation efficiency of photovoltaic power stations and the stability of the power system are improved.

CN118841945BActive Publication Date: 2025-07-11STATE GRID BEIJING ELECTRIC POWER CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410812363.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-07-11
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

The existing research on photovoltaic power generation prediction has similar daily selection criteria that are too one-sided. A single prediction model cannot effectively cope with various types of weather conditions, resulting in low prediction accuracy and inability to accurately capture the fluctuation and change trends of photovoltaic power generation, affecting the safe and stable operation and optimized scheduling of the power system.

Method used

The degree of joint similarity with similar numerical values and similar shapes is used as the judgment criteria, and the key influencing factors are determined through the Pearson coefficient method, combined with the European-style distance and twin neural network to form a comprehensive evaluation standard, and a fuzzy C-means clustering algorithm is used to generate similar day sets, and CNN-LSTM and DGM(2,1) models are constructed for predictions for ideal and non-ideal weather respectively.

Benefits of technology

It improves the accuracy and pertinence of photovoltaic power generation forecasts, enhances the safe and stable operation and scheduling management of the power system, and ensures the efficient operation of the photovoltaic power station.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118841945B_ABST
    Figure CN118841945B_ABST
Patent Text Reader

Abstract

The present invention discloses a photovoltaic power generation prediction system and method in the technical field of photovoltaic power generation power prediction, including: obtaining historical meteorological data and photovoltaic power generation data of a target power station; quantitatively evaluating the correlation between the historical meteorological data and the photovoltaic power generation data by using the Pearson coefficient method to determine the key influencing factors of the photovoltaic power generation; based on the key influencing factors of the photovoltaic power generation, using a combined method of Euclidean distance and siamese neural network to form a comprehensive evaluation criterion covering similar values and similar shapes. Compared with the traditional similar day selection method based on the clustering algorithm, the present invention uses the combined similarity degree of similar values and similar shapes as the judgment criterion, which helps to improve the efficiency and accuracy of clustering. In addition, for two weather types, a photovoltaic power generation prediction model of a pilot power station is constructed respectively, significantly improving the pertinence and advancement of the prediction model construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a photovoltaic power generation prediction system and method considering the selection of similar days and weather types, and belongs to the technical field of photovoltaic power generation power prediction. Background Technique

[0002] The excessive emission of greenhouse gases is the main cause of global warming, seriously hindering the stable and sustainable development of social economy. As a kind of renewable energy, solar energy has the advantages of sustainable utilization, rich resources, safety and stability, etc., and plays an important role in the process of energy transformation and structural adjustment. The photovoltaic industry in China is developing vigorously, and various indicators such as power generation and installed capacity are showing a rapid growth trend. However, the output power of the photovoltaic power generation system is easily affected by various factors such as local meteorological conditions, station equipment operation and maintenance, and human operation, and thus shows volatility and randomness, resulting in certain safety hazards in the process of photovoltaic power generation grid connection. Therefore, accurate prediction of photovoltaic power generation power is of great strategic significance for ensuring the safe and stable operation of the system after high-proportion photovoltaic access and guiding the power department to formulate scientific and perfect planning and dispatching strategies.

[0003] At present, scholars at home and abroad generally adopt a combined method of influencing factor identification, similar day selection, sequence decomposition and artificial intelligence prediction, and to a certain extent, achieve accurate prediction of photovoltaic power generation. However, the existing research on photovoltaic power generation prediction generally has problems such as too one-sided selection criteria for similar days, the inability of a single prediction model to effectively cope with various types of weather conditions, and the failure to form complementary advantages of various methods, resulting in low prediction accuracy, etc., and it is impossible to accurately capture the fluctuation change trend of photovoltaic power generation power, which is not conducive to giving full play to the supporting role of photovoltaic power generation prediction for the safe and stable operation and optimal dispatching of the power system. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art, and provide a photovoltaic power generation prediction system and method considering the selection of similar days and weather types. The combined proximity degree of numerical proximity and shape similarity is used as the judgment criterion, which helps to improve the efficiency and accuracy of clustering. In addition, for two weather types, a photovoltaic power generation prediction model of the pilot power station is constructed respectively, significantly improving the pertinence and advancement of the prediction model construction.

[0005] To achieve the above object, the present invention is implemented by adopting the following technical solution:

[0006] In the first aspect, the present invention provides a photovoltaic power generation prediction method considering the selection of similar days and weather types, including:

[0007] Obtain the historical meteorological data and photovoltaic power generation data of the target power station;

[0008] Use the Pearson coefficient method to quantitatively evaluate the correlation between historical meteorological data and photovoltaic power generation data, and determine the key influencing factors of photovoltaic power generation;

[0009] Based on the key influencing factors of photovoltaic power generation, use a combination method of Euclidean distance and siamese neural network to form a comprehensive evaluation criterion covering similar numerical values and shapes, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0010] Based on the key influencing factors of photovoltaic power generation and the selection results of historical similar days for the day to be predicted, divide the historical meteorological data into two categories: ideal weather and non-ideal weather, and construct two types of photovoltaic power generation prediction models for target power plants respectively;

[0011] According to the prediction results of the two types of photovoltaic power generation prediction models for target power plants, using the mean absolute error and root mean square error as evaluation indicators, compare with the baseline model to determine the best photovoltaic power generation prediction model for target power plants under two types of weather conditions;

[0012] Based on the best photovoltaic power generation prediction model, predict the photovoltaic power generation of the target power plant.

[0013] Furthermore, the historical meteorological data includes global irradiance, temperature, humidity, wind speed and wind direction. The ideal weather includes sunny days, and the non-ideal weather includes cloudy days, overcast days and rainy days.

[0014] Furthermore, the expression of the Pearson coefficient method is:

[0015]

[0016] In the formula: p is the correlation coefficient of the variable; M represents the candidate meteorological influencing factors of photovoltaic power generation; N represents the photovoltaic power; S is the number of samples.

[0017] Furthermore, based on the key influencing factors of photovoltaic power generation, use a combination method of Euclidean distance and siamese neural network to form a comprehensive evaluation criterion covering similar numerical values and shapes, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted, including:

[0018] S1: Use Euclidean distance to quantitatively evaluate the degree of numerical similarity between the photovoltaic power generation sequence and the meteorological element sequence. Among them, the Euclidean distance calculation formula is as follows:

[0019]

[0020] In the formula: D ik 、D(T,W) represents the Euclidean distance between the photovoltaic power generation sequence and the meteorological element sequence; T m 、W mThey respectively represent the m-th sample values of the photovoltaic power generation sequence and the meteorological element sequence; n is the total number of samples;

[0021] S2: Use a siamese neural network to measure the shape similarity between the photovoltaic power generation sequence and the meteorological element sequence, including: Input the sample data of the two sequences into two sub-networks with shared weights simultaneously, extract the morphological feature information of the two sequences, generate corresponding vectors, and determine their similarity by establishing a loss function. The calculation equation is as follows:

[0022] Z ik =Z(T snn ,W snn )=PG p (T snn )-G p (W snn )P

[0023] In the formula: Z ik represents the loss function, T snn and W snn respectively represent the data information of the two sequences input into the SNN; G P is the network model, whose function is to convert the input data into feature vectors; P represents the shared weight; Z(T snn ,W snn ) is the loss function value;

[0024] S3: Calculate the numerical proximity degree and the comprehensive distance of the numerical proximity degree between the photovoltaic power generation sequence and the meteorological element sequence, where: Based on the calculated Euclidean distance and loss function value, on the basis of normalizing them respectively, calculate the weighted sum value of the two as the comprehensive distance, comprehensively characterizing the proximity degree of the numerical values and shapes of the two data sequences. The specific form is as follows:

[0025]

[0026]

[0027]

[0028] In the formula: and are respectively the normalized Euclidean distance and loss function value, ρ is the comprehensive distance; σ and (1 - σ) are respectively the weight coefficients of the Euclidean distance and the loss function value;

[0029] S4: Use the comprehensive distance as the classification criterion and apply the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted, including:

[0030] S41: Divide the photovoltaic power generation volume sequence into c clusters (2 ≤ c ≤ n); then, select the comprehensive distance as the evaluation criterion, and calculate the membership matrix; finally, based on the membership matrix L = {l i1 , l i2 , …, l ik , …, l ic}, calculate the clustering center Z = {z1, z2, …, z k , …, z c}, and the expression is:

[0031]

[0032] In the formula: l ik is the membership degree of the i-th sample to the k-th cluster, ρ ik is the comprehensive distance of the i-th sample to the k-th clustering center, ρ ir is the comprehensive distance of the i-th sample to the r-th clustering center; s is the fuzzy weighting parameter, indicating the fuzziness of controlling the membership matrix L, and usually s takes 2; z k is the k-th clustering center, t i is the i-th sample;

[0033] S42: After the membership matrix L is initialized, continuously update and iterate to generate new Z and L. When the convergence condition ||Z t - z t+1 || ≤ ε or ||L t - L t+1 || ≤ ε is satisfied, where t represents the number of iterations, and the error ε takes 0.000001; or, the changes of L or Z are limited, the iterative calculation ends, and the clustering result is obtained according to the membership matrix. All historical days belonging to the same category as the day to be predicted are the historical similar day set of the day to be predicted.

[0034] Furthermore, for ideal weather, using the key influencing factors and photovoltaic power generation volume data as input variables, apply the CNN-LSTM model to construct a photovoltaic power generation volume prediction model for the target power station under ideal weather conditions. The calculation formula of the CNN convolution layer in the CNN-LSTM model is:

[0035]

[0036] Y r,s = δ(y r,s )

[0037] Y r,s+1 = pool(Y r,s )

[0038] In the formula: y r,sRepresents the output value of the s-th neuron on the r-th feature map in the convolutional layer; w r,i,j Represents the weight value at the i-th row and j-th column in the r-th convolutional kernel; x i,s+j-1 Represents the signal intensity sequence input at the i-th row and s + j - 1-th column; b r Represents the bias vector corresponding to the r-th convolutional kernel; M and N respectively correspond to the M-row and N-column vectors in the convolutional kernel; δ is the activation function; Y r,s Represents the feature map after the convolution operation; pool represents the pooling layer; Y r,s+1 Represents the feature map after the pooling operation; x and y respectively represent the input and output signal intensity sequences;

[0039] In the CNN-LSTM model, the LSTM layer includes a forget gate, an input gate, and an output gate, where:

[0040] The output expression of the forget gate is:

[0041]

[0042] In the formula: w f T Is the weight vector; f t Is the output value of the forget gate; x t Is the input of the t-th layer; h t-1 Is the output of the previous layer; b f Is the parameter value of the forget gate;

[0043] The specific expression of the input gate is:

[0044]

[0045] In the formula: C t Is the long-term state to be updated; w i T And w c T Are the weight vectors respectively; b i And b c Respectively represent the parameter values of the forget gate; c t Is the long-term state; c t-1 Is the long-term state of the previous layer, i t Is the data situation of the updated t-th layer;

[0046] The output gate expression is:

[0047]

[0048] In the formula: o t Represents the output value of the output gate; w o T Is the weight vector; xt is the input of the t-th layer; h t-1 is the output of the previous layer; b o is the output gate parameter; c t is the long-term state; h t represents the information for determining the output at the current moment.

[0049] Furthermore, for non-ideal weather, using the historical photovoltaic power of adjacent days of the day to be predicted corrected by the deviation ratio based on the ideal day as input data, applying the DGM(2,1) model to construct a prediction model for the photovoltaic power generation of the target power station under non-ideal weather conditions, the sample correction based on adjacent days and ideal days includes:

[0050] Calculating the deviation ratio of non-ideal weather using the daily total power value of the ideal day: First, select historical day samples that can truly reflect the power generation power change trend of the day to be predicted from the three sub-categories of non-ideal weather, and define them as ideal days; then, divide the daily total power value of the ideal day by the total power of the historical day of the same weather type as the ideal day to obtain the deviation ratio β i ; Finally, calculate the average value of the deviation ratio values under the same weather type to obtain the deviation ratio value under this weather, and so on, calculate the deviation ratio values corresponding to the three sub-categories of non-ideal weather respectively. The specific calculation formula is as follows:

[0051]

[0052] In the formula: W t is the daily total power value of the ideal day of the sub-category weather under non-ideal weather conditions; W r is the actual daily total power value of the sub-category weather under non-ideal weather conditions; a is the total number of days of the sub-category weather under non-ideal weather conditions; β 非理想天气 is the deviation ratio under non-ideal weather conditions;

[0053] Using the deviation ratio to correct the power of the historical adjacent days of the day to be predicted: Based on the deviation ratios of the three sub-categories of weather calculated above, correct the adjacent day samples of the day to be predicted, so as to obtain input samples with smaller fluctuations and gentler change trends. The specific calculation formula is as follows:

[0054]

[0055] In the formula: y (0) (m) is the daily total power value corresponding to the m-th sample in the adjacent day sample sequence; is the daily total power value corresponding to the i-th sample in the corrected adjacent day sample sequence, and β is the deviation ratio;

[0056] The specific formula of the DGM(2,1) prediction model is as follows:

[0057] The non - negative initial data sequence of the system variable is Y (0) , and a first - order cumulative subtraction sequence is generated, as shown below:

[0058] Y (0) =(y (0) (1), y (0) (2),..., y (0) (m))

[0059] α (1) Y (0) =(α (1) y (0) (2), α (1) y (0) (3), ΛΛ, α (1) y (0) (m))

[0060] In the formula: Y (0) is the non - negative initial data sequence of the system variable; y (0) (m) is the initial data sequence; m is the number of initial data; α (1) is the first - order cumulative sequence;

[0061] Based on the obtained first - order cumulative subtraction sequence for optimizing the non - linear equations, while fixing other variables, the initial sequence can be continuously optimized to obtain the formula:

[0062] α (1) y (0) (k)=y (0) (k)-y (0) (k - 1), k = 2, 3,..., m

[0063] Define as the whiting equation of the DGM(2,1) model a (1) y (0) (k)+ay (0) (k)=b. On the basis that B and X satisfy the numerical format stability, convergence, and as small as possible numerical dissipation, the least - squares estimation of parameters a and b satisfies The parameters a and b are obtained, as shown below:

[0064]

[0065] In the formula: B represents the time - series increment information and a constant sequence; X represents the information of the cumulative generation sequence; y (0) (m) is the initial data sequence; m is the number of initial data;

[0066] The solution is:

[0067]

[0068] Where: n is the number of generated sequences; y (0) (k) is the initial data sequence; k is the number of initial data; α (1) is the first-order accumulated sequence;

[0069] Generate a prediction sequence, specifically as follows:

[0070]

[0071] Where: represents the prediction sequence; a and b are parameters; y (0) (1) is the data sequence; k is the number of initial data; e -ak is the exponential decay factor in the model.

[0072] Furthermore, the calculation formulas for the mean absolute error and the root mean square error are as follows:

[0073]

[0074] Where: MAE is the mean absolute error; RMSE is the root mean square error; N represents the number of data to be predicted; Z(v) and Y(v) are the true value and the predicted value of the photovoltaic power generation system respectively; v represents the data to be predicted at a certain moment.

[0075] In a second aspect, the present invention provides a photovoltaic power generation prediction system considering similar day selection and weather type, including:

[0076] Data acquisition module: used to acquire the historical meteorological data and photovoltaic power generation data of the target power station;

[0077] Quantitative evaluation module: used to quantitatively evaluate the correlation between the historical meteorological data and the photovoltaic power generation data by using the Pearson coefficient method, and determine the key influencing factors of the photovoltaic power generation;

[0078] Comprehensive evaluation module: used to form a comprehensive evaluation criterion covering numerical similarity and shape similarity based on the key influencing factors of the photovoltaic power generation by using a combination method of Euclidean distance and siamese neural network, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0079] Model construction module: used to divide the historical meteorological data into two categories of ideal weather and non-ideal weather based on the key influencing factors of the photovoltaic power generation and the selection result of the historical similar days for the day to be predicted, and construct two types of photovoltaic power generation prediction models for the target power station respectively;

[0080] Model determination module: It is used to compare with the baseline model based on the prediction results of the photovoltaic power generation prediction models of two types of target power stations, with the mean absolute error and the root mean square error as evaluation indicators, and determine the best photovoltaic power generation prediction model for the target power station under two types of weather conditions;

[0081] Photovoltaic prediction module: It is used to predict the photovoltaic power generation of the target power station based on the best photovoltaic power generation prediction model.

[0082] Thirdly, the present invention provides a photovoltaic power generation prediction system device considering the selection of similar days and weather types, including a processor and a storage medium;

[0083] The storage medium is used to store instructions;

[0084] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the above.

[0085] Fourthly, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the above are implemented.

[0086] Compared with the prior art, the beneficial effects achieved by the present invention:

[0087] First, by collecting and sorting out the historical power generation data and historical meteorological factor data of photovoltaic power stations, and on the basis of selecting the key meteorological elements of photovoltaic power generation by using the Pearson correlation coefficient method, the present invention innovatively uses the Euclidean distance and the numerically similar and shape-similar indexes generated by the siamese neural network as the classification criteria of the fuzzy C-means clustering algorithm to select historical similar days, effectively solving the problem of ignoring the shape similarity of multi-dimensional data in traditional sample division, and accurately measuring the shape and numerical similarity degree between the sample data and the clustering center;

[0088] II. Based on the similar day selection results, in view of the significant differences in photovoltaic power generation under different weather conditions, a photovoltaic power generation prediction model under different weather types is established, significantly improving the pertinence and effectiveness of the model. For ideal weather (i.e., sunny days), a prediction model based on CNN-LSTM is constructed; for non-ideal weather (i.e., cloudy, overcast, and rainy days), by innovatively introducing the concepts of adjacent days and ideal days, the problem of similar day failure caused by the long time interval between historical similar days and the day to be predicted is solved, ensuring that the input samples have small fluctuations and gentle trends, and the DGM(2,1) model with a simple calculation process and low sample quality requirements is used to predict the photovoltaic power generation, avoiding the errors generated during the conversion process of the photovoltaic sequence from discrete to continuous, significantly improving the generalization and diversity of the photovoltaic power generation prediction model. Under non-ideal weather conditions, by integrating the concepts of adjacent days and ideal days, the influence of random disturbance factors on the DGM(2,1) model is weakened, realizing the accurate prediction of photovoltaic power generation under different weather types, improving the operation efficiency of photovoltaic power stations while strengthening power dispatching management to ensure the safe, stable, and economic operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments and descriptions thereof of the invention are used to explain the invention and do not constitute an improper limitation of the invention. In the drawings:

[0090] Figure 1 It is a technical roadmap for the model construction of the photovoltaic power generation prediction method considering similar day selection and weather types provided in Embodiment 1 of the present invention;

[0091] Figure 2 It is the photovoltaic power generation power situation of typical months in four seasons of 2019 in the pilot power station provided in Embodiment 1 of the present invention;

[0092] Figure 3 It is a schematic diagram of the error comparison of the photovoltaic power generation prediction model provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0093] The present invention will be described in detail below with reference to the drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0094] The following detailed descriptions are all exemplary descriptions, aiming to provide a further detailed description of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs. The terms used in the present invention are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention.

[0095] Example 1:

[0096] Please refer to Figures 1-3 , a photovoltaic power generation prediction method considering the selection of similar days and weather types. Based on the results of similar day selection, this example proposes corresponding photovoltaic power generation prediction methods for pilot power stations for different weather types. The method includes the following steps:

[0097] Step A: By consulting historical data, collect and organize the historical meteorological data and photovoltaic power generation data of the target power station, including photovoltaic power generation, global radiation irradiance, temperature, humidity, wind speed, wind direction, etc. In this example, through on-site investigation of the target power station and consulting relevant websites of the China Meteorological Administration, collect and organize the meteorological factor data including global radiation irradiance, temperature, humidity, wind speed, wind direction, etc. and photovoltaic power data of the photovoltaic power station in 2019. Taking January 1, 2019 as an example, Table 1 lists the main meteorological parameters and photovoltaic power of the pilot power station. From 8:00 in the morning to 18:00 in the evening, the time interval is 15 minutes, with a total of 40 data; Figure 2 Shows the photovoltaic power generation situation of typical months in four seasons of the pilot power station in 2019. The typical months in the four seasons are April, July, October, and January of the next year respectively.

[0098]

[0099]

[0100]

[0101] Table 1 Summary Table of Meteorological Elements and Photovoltaic Power Data of the Pilot Power Station

[0102] Step B: Based on the historical photovoltaic power generation and meteorological parameters collected in Step A, use the Pearson correlation coefficient method (PCC) to quantitatively evaluate the correlation between photovoltaic power generation and meteorological factors such as global radiation irradiance, temperature, humidity, wind speed, and wind direction, and determine the key influencing factors of photovoltaic power generation. Among them, the PCC method is used to quantitatively evaluate the correlation between the historical photovoltaic power generation data of the pilot power station and meteorological factors including global radiation irradiance, temperature, humidity, wind speed, and wind direction, and determine the key influencing factors of photovoltaic power generation. The specific PCC method involved is as follows:

[0103]

[0104] Where: p is the correlation coefficient of the variable; M represents the candidate meteorological influencing factors of photovoltaic power generation; N represents the photovoltaic power generation; and S is the number of samples. The p value calculated using PCC ranges from -1 to 1. A positive value represents a positive correlation, and a negative value represents a negative correlation. The closer the |p| value is to 1, the stronger the correlation between the two sequences.

[0105] Table 2 shows the correlation calculation results and rankings of meteorological elements and photovoltaic power generation. Among them, the top three influencing factors in the ranking are regarded as key influencing factors, and subsequent ones are used as input variables for the photovoltaic power generation prediction model.

[0106]

[0107]

[0108] Table 2 Screening Results Table of Key Meteorological Factors for Photovoltaic Power Generation

[0109] Step C: According to the key meteorological factors determined by the aforementioned PCC method and the power generation data of the photovoltaic power station, use a combination method of Euclidean Distance (ED) and Siamese neural network (SNN) to form a comprehensive evaluation criterion covering similar numerical values and shapes, and use this as the classification criterion for the Fuzzy C–Means algorithm (FCM) to generate a set of historical similar days for the day to be predicted, specifically including:

[0110] Step C1: Use Euclidean distance to quantitatively evaluate the degree of numerical similarity between two sequences.

[0111] The Euclidean distance method is used to measure the degree of numerical similarity between two sequences. Assume that the photovoltaic power generation sequence and the meteorological element sequence are T = [T1, T2, …, T n and W = [W1, W2, ……, W n , and the Euclidean distance D ik between them is calculated as follows:

[0112]

[0113] Where: D(T, W) is the Euclidean distance between the photovoltaic power generation sequence and the meteorological element sequence; T m , W m represent the m-th sample values of the photovoltaic power generation sequence and the meteorological element sequence respectively; and n is the total number of samples.

[0114] Step C2: Use a Siamese neural network to measure the degree of shape similarity between two sequences.

[0115] The SNN consists of two neural networks (such as CNN, FCN, RNN, etc.) with the same parameters and arranged side by side. Its principle is to judge the similarity between sequences by calculating the loss of input features on the basis of feature extraction of the two sequences.

[0116] Step C201: Input the sample data of the two sequences into two sub-networks with shared weights at the same time, extract the morphological feature information of the two sequences, and generate corresponding vectors;

[0117] Step C202: Determine the similarity degree between the two by establishing a loss function, and its calculation equation is as follows:

[0118] Z ik =Z(T snn ,W snn )=PG p (T snn )-G p (W snn )P(2b)

[0119] In the formula: Z ik represents the loss function, T snn and W snn respectively represent the data information of the two sequences input into the SNN; G P is the network model, and its function is to convert the input data into feature vectors; P represents the shared weight; Z(T snn ,W snn ) is the loss function value. The smaller the function value, the higher the similarity degree of the shapes of the two sequence samples; on the contrary, the similarity degree is relatively low.

[0120] Step C3: Calculate the numerical proximity degree between the two sequences and the comprehensive distance of the numerical proximity degree.

[0121] Based on the Euclidean distance and the loss function value respectively calculated in the previous two steps, on the basis of normalizing them by using equations (2c) and (2d) respectively, use equation (2e) to calculate the weighted sum value of the two as the comprehensive distance, comprehensively representing the numerical and shape proximity degrees of the two data sequences. The specific form is as follows:

[0122]

[0123] In the formula: and are the Euclidean distance and the loss function value after normalization respectively, ρ is the comprehensive distance; σ and (1 - σ) are the weight coefficients of the Euclidean distance and the loss function value respectively, and are set to 0.5.

[0124] Step C4: Using the comprehensive distance calculated in Step C3 as the classification criterion, apply the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted.

[0125] Step C401: Calculate the membership matrix. The FCM algorithm aims to minimize the objective function. Based on the membership function that reflects the fuzzy membership relationship between historical data and cluster centers, it iteratively and optimizes to determine that the sample data belongs to a certain cluster center to achieve sample classification. As described below, first, divide T into c clusters (2 ≤ c ≤ n); then, select the comprehensive distance calculated by the above formula (2e) as the evaluation criterion, and use formula (2f) to calculate the membership matrix; finally, based on the membership matrix L = {l i1 , l i2 , …, l ik , …, l ic}, use formula (2g) to calculate the cluster center Z = {z1, z2, …, z k , …, z c} of each cluster.

[0126]

[0127] In the formula: l ik is the membership degree of the i-th sample to the k-th cluster, ρ ik is the comprehensive distance of the i-th sample to the k-th cluster center, ρ ir is the comprehensive distance of the i-th sample to the r-th cluster center; s is the fuzzy weighting parameter, indicating the fuzziness of controlling the membership matrix L, usually s takes 2; z k is the k-th cluster center, t i is the i-th sample.

[0128] Step C402: Generate the clustering result.

[0129] After the matrix L is initialized in Step C401, continuously update and iterate to generate new Z and L. When the convergence condition is met, that is, ||Z t - z t+1 || ≤ ε or ||L t - L t+1 || ≤ ε, where t represents the number of iterations, and the error ε takes 0.000001; or, the changes in L or Z are limited, the iterative calculation ends, and according to the membership matrix, the clustering result is obtained. All historical days that belong to the same category as the day to be predicted are the set of historical similar days for the day to be predicted.

[0130] Table 3 shows the weather type of the day to be predicted and the number of historical similar days belonging to the same category as it.

[0131]

[0132]

[0133] Table 3 Summary of Weather Types for the Day to be Predicted

[0134] Step D: Based on the key influencing factors determined in Steps B and C and the selection results of the historical similar days of the day to be predicted, the historical data is divided into two major categories: ideal weather (i.e., sunny days) and non-ideal weather (i.e., cloudy days, overcast days, and rainy days). Table 4 shows the number of days corresponding to the two weather types.

[0135]

[0136] Table 4 Specific Number Distribution Table of Two Weather Types

[0137] Step E: For the characteristics of ideal weather (i.e., sunny days) and non-ideal weather (i.e., cloudy days, overcast days, and rainy days), respectively construct a target power generation prediction model for the power station with strong adaptability, specifically including:

[0138] Step E1: For ideal weather, using the key meteorological influencing factor data and historical power generation data as input variables, apply the CNN-LSTM model to construct a target power generation prediction model for the power station under ideal weather conditions.

[0139] Step E01: Construct a CNN model.

[0140] The Convolutional Neural Networks-Long Short-Term Memory (CNN-LSTM) model is a hybrid neural network architecture, usually including an input layer, a CNN convolutional layer, a pooling layer, an LSTM layer, and an output layer. The organic combination of CNN and LSTM can achieve the joint modeling of time series and spatial features, thereby improving the prediction accuracy. Among them, the CNN part is mainly used to extract the main features of the input data. Its core is the convolutional layer with the characteristics of local connection and weight sharing, which can effectively reduce the number of parameters and improve the applicable range. As for the introduction of the activation function, it can effectively approximate nonlinear functions of various complexities, improve the expression ability of the model, and avoid problems such as gradient disappearance. The specific calculation formula is as follows:

[0141]

[0142] Y r,s =δ(y r,s )(3b)

[0143] Y r,s+1 =pool(Y r,s )(3c)

[0144] Where: yr,s represents the output value of the s-th neuron on the r-th feature map in the convolutional layer; w r,i,j represents the weight value at the i-th row and j-th column in the r-th convolutional kernel; x i,s+j-1 represents the signal intensity sequence input at the i-th row and the (s + j - 1)-th column; b r represents the bias vector corresponding to the r-th convolutional kernel; M and N respectively correspond to the M-row and N-column vectors in the convolutional kernel; δ is the activation function; Y r,s represents the feature map after the convolutional operation; pool represents the pooling layer; Y r,s+1 represents the feature map after the pooling operation; x and y respectively represent the input and output signal intensity sequences.

[0145] Step E02: Construct an LSTM model.

[0146] The LSTM model is an improved deep learning model based on the Recursive Neural Network (RNN). On the basis of retaining the memory caching function of the RNN, it introduces a "memory block" solution, increasing the long-term memory function of the model while solving the inherent "information loss and gradient explosion" problems in the RNN link structure. Compared with traditional models, the LSTM network introduces a "gate" structure. Among them, the forget gate is mainly responsible for avoiding the transmission of invalid information when storing key information; the input gate and output gate are responsible for identifying, processing, and transmitting data to the next level.

[0147] (1) The forget gate uses the activation function sigm to process the input x of the t-th layer t and the output h of the previous layer t-1 to process the data. On the basis of retaining the data information of the long-term cell state c t , it transmits the processing result to the long-term state c of the previous level t-1 . The output f of the forget gate t is expressed as:

[0148]

[0149] In the formula: w f T is the weight vector; f t is the output value of the forget gate; x t is the input of the t-th layer; h t-1 is the output of the previous layer; b f is the parameter value of the forget gate.

[0150] (2) The input gate is designed to control the data information retained after being processed by the input gate. For example, using the sigm activation function for x t and h t-1Perform data processing to determine the data condition i of the t-th layer after update t or the long-term state C to be updated obtained after being processed by the tanh function t The specific expression is as follows:

[0151]

[0152] In the formula: C t is the long-term state to be updated; w i T and w c T are weight vectors respectively; b i and b c represent the parameter values of the forget gate respectively; c t is the long-term state; c t-1 is the long-term state of the previous layer, and i t is the data condition of the t-th layer after update.

[0153] (3) The output gate is mainly used to control the data processed by the output gate and output the information determined to be output at the current moment. The specific expression is as follows:

[0154]

[0155] In the formula: o t represents the output value of the output gate; w o T is the weight vector; x t is the input of the t-th layer; h t-1 is the output of the previous layer; b o is the output gate parameter; c t is the long-term state; h t represents the information determined to be output at the current moment.

[0156] Step E2: For non-ideal weather, using the historical photovoltaic power of the adjacent days of the day to be predicted corrected by the deviation ratio based on the ideal day as the input data, apply the DGM(2,1) model to construct a prediction model for the photovoltaic power generation of the target power station under non-ideal weather conditions;

[0157] Step E201: Sample correction based on adjacent days and ideal days.

[0158] Since the number of historical samples of non-ideal weather is limited, resulting in a large time and space interval of the photovoltaic sequence of its historical similar day set and a phenomenon of violent fluctuation changes, making the historical similar days ineffective; if not processed and directly using artificial intelligence algorithms to construct a prediction model, it is impossible to achieve accurate prediction of photovoltaic power generation. Therefore, ideal days and adjacent days are introduced to correct the original sequence, and the specific process is as follows:

[0159] (1) Calculate the deviation ratio of non-ideal weather using the daily total power value of the ideal day. First, select historical day samples from the three sub-categories of non-ideal weather that can truly reflect the changing trend of the power generation on the day to be predicted, and define them as ideal days; then, divide the daily total power value of the ideal day by the total power of the historical day of the same weather type as the ideal day to obtain the deviation ratio β i ; finally, calculate the average value of the β i values under the same weather type to obtain the β value for this weather. By analogy, calculate the β values corresponding to the three sub-categories of non-ideal weather respectively. The specific calculation formula is as follows:

[0160]

[0161] In the formula: W t is the daily total power value of the ideal day of the sub-category weather under non-ideal weather conditions; W r is the actual daily total power value of the sub-category weather under non-ideal weather conditions; a is the total number of days of the sub-category weather under non-ideal weather conditions; β 非理想天气 is the deviation ratio under non-ideal weather conditions.

[0162] (2) Use the deviation ratio to correct the power of the historical adjacent days of the day to be predicted. Based on the deviation ratios of the three sub-categories of weather calculated above, correct the adjacent day samples of the day to be predicted, so as to obtain input samples with smaller fluctuations and gentler changing trends. The specific calculation formula is as follows:

[0163]

[0164] In the formula: y (0) (m) is the daily total power value corresponding to the m-th sample in the adjacent day sample sequence; is the daily total power value corresponding to the i-th sample in the corrected adjacent day sample sequence, and β is the deviation ratio.

[0165] Step E202: Construct a DGM(2,1) prediction model.

[0166] Grey system theory (Grey model, GM) takes the uncertain system of "small sample" and "poor information" as the research object, and can effectively process non-linear systems with a large amount of unknown data information. In order to avoid the error caused by the transition from discrete to continuous, based on the DGM(2,1) model of the grey theory model, on the basis of representing the historical sequence as a sequence in exponential distribution form, the differential equation is used to approximate and fit the changing trend of the sequence, so that it has high sensitivity and fitting ability, and can accurately predict the historical sequence with significant fluctuating changing trends. The specific formula is as follows:

[0167] (1) Assume that the non-negative initial data sequence of the system variable is Y(0) , generate a first-order cumulative subtraction sequence, as shown below:

[0168] Y (0) = (y (0) (1), y (0) (2),..., y (0) (m))(4d)

[0169] α (1) Y (0) = (α (1) y (0) (2), α (1) y (0) (3), ΛΛ, α (1) y (0) (m))(4e)

[0170] Where: Y (0) is a non-negative initial data sequence of system variables; y (0) (m) is the initial data sequence; m is the number of initial data; α (1) is the first-order cumulative sequence;

[0171] The first-order cumulative subtraction sequence used to optimize the non-linear equations obtained based on formula (4e) can optimize the initial sequence continuously while fixing other variables, thus obtaining formula (4f):

[0172] α (1) y (0) (k) = y (0) (k) - y (0) (k - 1), k = 2, 3,..., m (4f)

[0173] (2) Definition is the whiting equation of the DGM(2,1) model a (1) y (0) (k) + ay (0) (k) = b. On the basis that B and X satisfy the numerical format stability, convergence, and as little numerical dissipation as possible, the least squares estimation of parameters a and b satisfies Parameters a and b are obtained according to formulas (4h) and (4i), as shown below:

[0174]

[0175] Where: B represents the time series increment information and a constant column; X represents the information of the cumulative generation sequence; y (0) (m) is the initial data sequence; m is the number of initial data.

[0176] Solved:

[0177]

[0178] Where: n is the number of generated sequences; y (0) (k) is the initial data sequence; k is the number of initial data; α (1) is the first-order accumulated sequence;

[0179] (3) Generate the prediction sequence, specifically as follows:

[0180]

[0181] Where: represents the prediction sequence; a and b are parameters; y (0) (1) is the data sequence; k is the number of initial data; e -ak is the exponential decay factor in the model.

[0182] Under ideal weather conditions, select the key meteorological impact factor data on sunny days and the historical photovoltaic power generation data as input variables, and use the CNN-LSTM model to construct the photovoltaic power generation prediction model of the pilot power station under ideal weather conditions; use MAE and RMSE as evaluation indicators to examine the prediction effect of the CNN-LSTM model. The specific situation is shown in Table 5.

[0183]

[0184] Table 5 Summary of the prediction effect of the CNN-LSTM model under ideal weather conditions

[0185] Under non-ideal weather conditions, use the historical photovoltaic power of the adjacent days of the day to be predicted corrected by the deviation ratio based on the ideal day as input data, and use the DGM(2,1) model to construct the photovoltaic power generation prediction model of the pilot power station under non-ideal weather conditions. According to the daily total power data collected above, calculate the deviation ratios of the three types of non-ideal weather. The specific situation is shown in Table 6.

[0186]

[0187] Table 6 Summary of the deviation ratios of the three types of non-ideal weather

[0188] Specifically, take February 28th, the day to be predicted, as an example. Its weather type is rainy. The meteorological conditions and power generation power of its adjacent days are shown in Table 7. Use the deviation ratio of the rainy weather type calculated above to correct the total output power of the adjacent days. The correction results are shown in Table 7.

[0189]

[0190] Table 7 Meteorological data of the adjacent days of the typical day to be predicted and the total output power before and after correction

[0191] Taking the total daily power of the corrected adjacent days as the input variable, the DGM(2,1) model is used to predict the total daily power of the typical day to be predicted (i.e., February 28), and the prediction accuracy rate is 90.42%. And so on, Table 8 shows the prediction of the DGM(2,1) model for 3 days to be predicted under non-ideal weather conditions.

[0192]

[0193]

[0194] Table 8 Prediction of the DGM(2,1) model for 3 days to be predicted under non-ideal weather conditions

[0195] Step F, select BP and LSTM as the benchmark models, and respectively construct the power generation prediction models of the pilot photovoltaic power station under ideal and non-ideal weather conditions. Taking the Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) as the evaluation indicators, compare and evaluate the prediction effects of the CNN-LSTM model, DGM(2,1) model and the two benchmark models, and determine the most suitable power generation prediction model of the target power station under different weather conditions to achieve accurate power prediction of the photovoltaic power station. The calculation formulas of the two indicators are as follows:

[0196]

[0197] In the formula: MAE is the Mean Absolute Error; RMSE is the Root Mean Squared Error; N represents the number of data to be predicted; Z(v) and Y(v) are the true value and predicted value of the photovoltaic power generation system respectively; v represents. For specific situations, please refer to Table 9 and Table 10.

[0198]

[0199]

[0200] Table 9 Comparison of prediction results of various models under ideal weather conditions

[0201]

[0202] Table 10 Comparison of prediction results of various models under non-ideal weather conditions

[0203] Step G, according to the introduction of the photovoltaic power generation prediction method of the pilot power station in the previous steps A - F, successively construct a data collection module, a key influencing factor identification module, a similar day selection module, a photovoltaic power generation system prediction module corresponding to ideal and non-ideal weather conditions, and a model evaluation module, and complete the deployment of the virtual device for predicting the photovoltaic power generation of the pilot power station, specifically including:

[0204] Step G1: Construction of data collection module: Store the historical photovoltaic power generation data and candidate meteorological factor data of the collected typical pilot power stations into the data collection module;

[0205] Step G2: Construction of key influencing factor identification module: Call the historical photovoltaic power generation data and meteorological factor data of the pilot power stations in the data collection module, and use the Pearson coefficient method to quantitatively evaluate the correlation between the historical photovoltaic power generation data and meteorological factors, and determine the key influencing factors;

[0206] Step G3: Similar day selection module: Based on the previously selected key meteorological factors, use a combination method of Euclidean distance and twin neural network to form a comprehensive evaluation criterion covering similar values and shapes, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0207] Step G4: Construction of photovoltaic power prediction module under ideal weather conditions: Based on the clustering results of historical days and combined with the actual weather conditions in the local area, divide the set of historical similar days for the day to be predicted into two major categories and four sub-categories of ideal weather (sunny days) and non-ideal weather (cloudy days, overcast days, rainy days). Using the key meteorological factor data and historical photovoltaic power data as input variables, use the CNN-LSTM model to construct a prediction model for the photovoltaic power generation of the target power station under ideal weather conditions;

[0208] Step G5: Construction of photovoltaic power prediction module under non-ideal weather conditions: For non-ideal weather, use the historical photovoltaic power of the adjacent days of the day to be predicted corrected by the deviation ratio based on the ideal day as the input data, and use the DGM(2,1) model to construct a prediction model for the photovoltaic power generation of the target power station under non-ideal weather conditions;

[0209] Step G6: Construction of model evaluation module: Call the prediction results of the above two photovoltaic power prediction modules, use the mean absolute error and root mean square error as evaluation indicators, compare with the baseline model, and select the best photovoltaic power generation prediction model for the power station under different weather conditions.

[0210] A method for predicting photovoltaic power generation considering the selection of similar days and the classification of weather types in a pilot power station. First, conduct on-site research on the target power station, consult relevant websites of the China Meteorological Administration, and collect meteorological factor data including annual global irradiance, temperature, humidity, wind speed, wind direction, etc., and photovoltaic power generation data of the pilot photovoltaic power station. Secondly, use the Pearson correlation coefficient method to quantitatively evaluate the correlation between the photovoltaic power of the pilot power station and the above meteorological elements, and determine the key meteorological factors affecting the photovoltaic power generation of the pilot power station. Thirdly, use a combination method of Euclidean distance and Siamese neural network to form a comprehensive evaluation criterion covering similar values and similar shapes, and use this as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted. Next, based on the above-selected and determined main meteorological factors and the set of historical similar days for the day to be predicted, divide the historical data into two categories: ideal weather (i.e., sunny days) and non-ideal weather (i.e., cloudy, overcast, and rainy days). For the significant differences in photovoltaic power generation under different weather conditions, establish an artificial intelligence prediction model for effectively processing stationary sequences under ideal weather conditions, and a DGM(2,1) model incorporating the concepts of adjacent days and ideal days under non-ideal weather conditions. Then, use the mean absolute error and root mean square error as evaluation indicators to quantitatively evaluate the prediction performance of the above two types of models and the benchmark model, select the most suitable power generation prediction model corresponding to different weather conditions for the target power station, and achieve accurate prediction of the power of the photovoltaic power station. Finally, construct a data collection module, a key influencing factor identification module, a similar day selection module, a photovoltaic power generation prediction module corresponding to ideal weather and non-ideal weather conditions, and a model evaluation module in sequence to complete the deployment of the virtual device for predicting photovoltaic power generation in the pilot power station, achieve accurate prediction of the photovoltaic power generation system in the pilot power station, improve the operation efficiency of the photovoltaic power station, and strengthen power dispatching management to ensure the safe, stable and economic operation of the power system.

[0211] Embodiment 2:

[0212] A photovoltaic power generation prediction system considering the selection of similar days and weather types can implement the photovoltaic power generation prediction method considering the selection of similar days and weather types described in Embodiment 1, including:

[0213] A data acquisition module: used to acquire historical meteorological data and photovoltaic power generation data of the target power station;

[0214] A quantitative evaluation module: used to quantitatively evaluate the correlation between historical meteorological data and photovoltaic power generation data by using the Pearson coefficient method, and determine the key influencing factors of photovoltaic power generation;

[0215] Comprehensive evaluation module: It is used to form a comprehensive evaluation criterion covering similar values and shapes based on the key influencing factors of photovoltaic power generation, using a combination method of Euclidean distance and siamese neural network, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0216] Model construction module: It is used to divide historical meteorological data into two categories: ideal weather and non-ideal weather based on the key influencing factors of photovoltaic power generation and the selection result of historical similar days for the day to be predicted, and construct two types of photovoltaic power generation prediction models for target power stations respectively;

[0217] Model determination module: It is used to compare with the baseline model according to the prediction results of the two types of photovoltaic power generation prediction models for target power stations, with the mean absolute error and root mean square error as evaluation indicators, and determine the best photovoltaic power generation prediction model for the target power station under two types of weather conditions;

[0218] Photovoltaic prediction module: It is used to predict the photovoltaic power generation of the target power station based on the best photovoltaic power generation prediction model.

[0219] Embodiment 3:

[0220] The embodiment of the present invention also provides a vehicle fault indicator signal processing device, which can implement the photovoltaic power generation prediction method considering the selection of similar days and weather types described in Embodiment 1, including a processor and a storage medium;

[0221] The storage medium is used to store instructions;

[0222] The processor is used to operate according to the instructions to execute the steps of the following method:

[0223] Obtain the historical meteorological data and photovoltaic power generation data of the target power station;

[0224] Use the Pearson coefficient method to quantitatively evaluate the correlation between historical meteorological data and photovoltaic power generation data, and determine the key influencing factors of photovoltaic power generation;

[0225] Based on the key influencing factors of photovoltaic power generation, use a combination method of Euclidean distance and siamese neural network to form a comprehensive evaluation criterion covering similar values and shapes, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0226] Based on the key influencing factors of photovoltaic power generation and the selection result of historical similar days for the day to be predicted, divide historical meteorological data into two categories: ideal weather and non-ideal weather, and construct two types of photovoltaic power generation prediction models for target power stations respectively;

[0227] According to the prediction results of the photovoltaic power generation prediction models for two types of target power stations, with the mean absolute error and root mean square error as evaluation indicators, compare with the baseline model to determine the best photovoltaic power generation prediction model for the target power station under two types of weather conditions;

[0228] Based on the best photovoltaic power generation prediction model, predict the photovoltaic power generation of the target power station.

[0229] Embodiment 4:

[0230] The embodiment of the present invention also provides a computer-readable storage medium, which can implement the photovoltaic power generation prediction method considering the selection of similar days and weather types described in Embodiment 1. There is a computer program stored thereon. When the program is executed by a processor, the steps of the following method are implemented:

[0231] Obtain the historical meteorological data and photovoltaic power generation data of the target power station;

[0232] Use the Pearson coefficient method to quantitatively evaluate the correlation between the historical meteorological data and the photovoltaic power generation data, and determine the key influencing factors of the photovoltaic power generation;

[0233] Based on the key influencing factors of the photovoltaic power generation, use a combined method of Euclidean distance and siamese neural network to form a comprehensive evaluation criterion covering similar values and similar shapes, and use it as the classification criterion for the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted;

[0234] Based on the key influencing factors of the photovoltaic power generation and the selection results of the historical similar days for the day to be predicted, divide the historical meteorological data into two categories: ideal weather and non-ideal weather, and respectively construct two types of photovoltaic power generation prediction models for target power stations;

[0235] According to the prediction results of the two types of photovoltaic power generation prediction models for target power stations, with the mean absolute error and root mean square error as evaluation indicators, compare with the baseline model to determine the best photovoltaic power generation prediction model for the target power station under two types of weather conditions;

[0236] Based on the best photovoltaic power generation prediction model, predict the photovoltaic power generation of the target power station.

Claims

1. A photovoltaic power generation prediction method considering similar day selection and weather types, characterized in that, Including: Obtain the historical meteorological data and photovoltaic power generation data of the target power station; Use the Pearson coefficient method to quantitatively evaluate the correlation between the historical meteorological data and the photovoltaic power generation data, and determine the key influencing factors of the photovoltaic power generation; Based on the key influencing factors of the photovoltaic power generation, use the Euclidean distance to quantitatively evaluate the degree of numerical similarity, use the Siamese neural network to measure the degree of shape similarity, calculate the comprehensive distance of the degree of numerical similarity and the degree of numerical similarity, and use it as the classification criterion of the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted; Based on the key influencing factors of the photovoltaic power generation and the selection results of the historical similar days of the day to be predicted, divide the historical meteorological data into two categories: ideal weather and non-ideal weather, and construct two types of photovoltaic power generation prediction models for the target power station respectively; According to the prediction results of the two types of photovoltaic power generation prediction models for the target power station, use the mean absolute error and the root mean square error as evaluation indicators, compare with the baseline model, and determine the best photovoltaic power generation prediction model for the target power station under the two types of weather conditions; Based on the best photovoltaic power generation prediction model, predict the photovoltaic power generation of the target power station.

2. The photovoltaic power generation prediction method considering similar day selection and weather type according to claim 1, characterized in that The historical meteorological data includes global radiation irradiance, temperature, humidity, wind speed and wind direction. The ideal weather includes sunny days, and the non-ideal weather includes cloudy days, overcast days and rainy days.

3. The photovoltaic power generation prediction method considering the selection of similar days and weather types according to claim 1, characterized in that, The expression of the Pearson coefficient method is: Where: p is the correlation coefficient of the variable; M represents the candidate meteorological influencing factors of the photovoltaic power generation; N represents the photovoltaic power; S is the number of samples.

4. The photovoltaic power generation prediction method considering the selection of similar days and weather types according to claim 1, characterized in that Based on the key influencing factors of the photovoltaic power generation, use the Euclidean distance to quantitatively evaluate the degree of numerical similarity, use the Siamese neural network to measure the degree of shape similarity, calculate the comprehensive distance of the degree of numerical similarity and the degree of numerical similarity, and use it as the classification criterion of the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted, including: S1: Use the Euclidean distance to quantitatively evaluate the degree of numerical similarity between the photovoltaic power generation sequence and the meteorological element sequence. The Euclidean distance calculation formula is as follows: Where: D ik , D(T, W) represents the Euclidean distance between the photovoltaic power generation sequence and the meteorological element sequence; T m , W m respectively represent the m-th sample values of the photovoltaic power generation sequence and the meteorological element sequence; n is the total number of samples; S2: Use the Siamese neural network to measure the degree of shape similarity between the photovoltaic power generation sequence and the meteorological element sequence, including: input the sample data of the two sequences into two sub-networks with shared weights at the same time, extract the morphological feature information of the two sequences, generate corresponding vectors, and determine the similarity between the two by establishing a loss function. The calculation equation is as follows: Z ik = Z(T snn , W snn ) = PG p (T snn ) - G p (W snn )P Where: Z ik represents the loss function, T snn and W snn respectively represent the data information of two sequences input into the siamese neural network; G P is the network model, whose role is to convert the input data into feature vectors; P represents the shared weight; Z(T snn ,W snn ) is the loss function value; S3: Calculate the comprehensive distance of the degree of numerical similarity and the degree of numerical similarity between the photovoltaic power generation sequence and the meteorological element sequence, where: based on the calculated Euclidean distance and loss function value, on the basis of normalizing them respectively, calculate the weighted sum value of the two as the comprehensive distance, comprehensively representing the degree of numerical and shape similarity of the two data sequences. The specific form is as follows: Where: D ik and Z ik are the Euclidean distance and the loss function value after normalization, respectively, ρ is the comprehensive distance; σ and (1 - σ) are the weight coefficients of the Euclidean distance and the loss function value, respectively; S4: Use the comprehensive distance as the classification criterion, and use the fuzzy C-means clustering algorithm to generate a set of historical similar days for the day to be predicted, including: S41: Divide the photovoltaic power generation amount sequence into c clusters, where 2 ≤ c ≤ n; then, select the comprehensive distance as the evaluation criterion and calculate the membership degree matrix; finally, based on the membership degree matrix L = {l i1 , l i2 , …, l ik , …, l ic}, calculate the clustering center Z = {z1, z2, …, z k , …, z c}, and the expression is: where: l ik is the membership degree of the i-th sample to the k-th cluster set, ρ ik is the comprehensive distance of the i-th sample to the k-th clustering center, ρ ir is the comprehensive distance of the i-th sample to the r-th clustering center; s is the fuzzy weighting parameter, representing the fuzziness of controlling the membership degree matrix L, and s takes 2; z k is the k-th clustering center, t i is the i-th sample; S42: After the membership matrix L has been initialized, continuously update and iterate to generate new Z and L. When the convergence condition ||Z t - z t+1 || ≤ ε or ||L t - L t+1 || ≤ ε is satisfied, where t represents the number of iterations and the error ε is taken as 0.000001; alternatively, the changes in L or Z are limited, the iterative calculation ends, and the clustering result is obtained based on the membership matrix. All historical days that belong to the same category as the day to be predicted are the set of historical similar days of the day to be predicted.

5. The photovoltaic power generation prediction method considering the selection of similar days and weather types according to claim 1, characterized in that, For ideal weather, using the key influencing factors and photovoltaic power generation data as input variables, a photovoltaic power generation prediction model for the target power station under ideal weather conditions is constructed by applying the CNN-LSTM model. The calculation formula of the CNN convolutional layer in the CNN-LSTM model is as follows: where: y r,s represents the output value of the s-th neuron in the r-th feature map in the convolutional layer; w r,i,j represents the weight value at the i-th row and j-th column in the r-th convolutional kernel; x i,s+j-1 represents the signal intensity sequence input at the i-th row and the (s + j - 1)-th column; b r represents the bias vector corresponding to the r-th convolutional kernel; M and N respectively correspond to the M-row and N-column vectors in the convolutional kernel; δ is the activation function; Y r,s represents the feature map after the convolution operation; pool represents the pooling layer; Y r,s+1 represents the feature map after the pooling operation; x and y respectively represent the input and output signal intensity sequences; In the CNN-LSTM model, the LSTM layer includes a forget gate, an input gate, and an output gate, where: The output expression of the forget gate is: Wherein: is the weight vector; f t is the output value of the forgetting gate; x t is the input of the t-th layer; h t-1 is the output of the previous layer; b f is the parameter value of the forgetting gate; The specific expression of the input gate is: Where: C t is the long-term state to be updated; w i T and w c T are weight vectors respectively; b i and b c represent the parameter values of the forget gate respectively; c t is the long-term state; c t-1 is the long-term state of the previous layer, and i t is the data condition of the t-th layer after update; The output gate expression is: Where: o t represents the output value of the output gate; w o T is the weight vector; x t is the input of the t-th layer; h t-1 is the output of the previous layer; b o is the output gate parameter; c t is the long-term state; h t represents the information for determining the output at the current moment.

6. The photovoltaic power generation prediction method considering similar day selection and weather type according to claim 1, characterized in that For non-ideal weather, using the historical photovoltaic power of adjacent days of the day to be predicted corrected by the deviation ratio based on the ideal day as input data, a photovoltaic power generation prediction model for the target power station under non-ideal weather conditions is constructed. The sample correction based on adjacent days and ideal days includes: The deviation ratio of non-ideal weather is calculated using the total daily power value of the ideal day: First, historical day samples that can truly reflect the changing trend of the power generation of the day to be predicted are selected from the three sub-categories of non-ideal weather and defined as ideal days; then, the total daily power value of the ideal day is divided by the total power of the historical day of the same weather type as the ideal day to obtain the deviation ratio β i ; finally, the average value of the deviation ratio values under the same weather type is calculated to obtain the deviation ratio value for this weather, and so on. The deviation ratio values corresponding to the three sub-categories of non-ideal weather are calculated respectively. The specific calculation formula is as follows: Where: W t is the daily total power value of the ideal day of the sub - type weather under non - ideal weather conditions; W r is the actual daily total power value of the sub - type weather under non - ideal weather conditions; a is the total number of days of the sub - type weather under non - ideal weather conditions; β 非理性天气 is the deviation ratio under non - ideal weather conditions; Using the deviation ratio to correct the power of the historical adjacent days of the day to be predicted: Based on the deviation ratios of the three small categories of weather calculated above, correct the adjacent day samples of the day to be predicted, so as to obtain input samples with smaller fluctuations and gentler change trends. The specific calculation formula is as follows: where: y (0) (m) is the daily total power value corresponding to the m-th sample in the adjacent-day sample sequence; is the daily total power value corresponding to the i-th sample in the corrected adjacent-day sample sequence, and β is the deviation ratio; The specific formula of the DGM(2,1) prediction model is as follows: The non - negative initial data sequence of the system variable is Y (0) , generate a sequence of successive subtractions, which is specifically as follows: Y (0) = (y (0) (1), y (0) (2),..., y (0) (m)) α (1) Y (0) = (α (1) y (0) (2), α (1) y (0) (3), ΛΛ, α (1) y (0) (m)) where: Y (0) is a non - negative initial data sequence of system variables; y (0) (m) is the initial data sequence; m is the number of initial data; α(1) is the first - order accumulated sequence; Based on the obtained first-order accumulated difference sequence for optimizing the non-linear equations, while fixing other variables, continuously optimize the initial sequence, so as to obtain the formula: α (1) y (0) y(k) = y (0) y(k) - y (0) y(k - 1), k = 2, 3, ..., m Definition is the DGM(2,1) model a (1) y (0) (k)+ay (0) (k) = b's grey equation. On the basis that B and X satisfy the numerical format stability, convergence and as little numerical dissipation as possible, the least squares estimation of parameters a and b satisfies Obtain parameters a and b, which are specifically as follows: Where: B represents the incremental information of the time series and a constant sequence; X represents the information of the accumulated generating sequence; y (0) (m) is the initial data sequence; m is the number of initial data; Solve to get: Where: n is the number of generated sequences; y (0) (k) is the initial data sequence; k is the number of initial data; α (1) is the first-order accumulated sequence; Generate a prediction sequence, specifically as follows: In the formula: represents the prediction sequence; a and b are parameters; y (0) (1) is the data sequence; k is the number of initial data; e -ak is the exponential decay factor in the model.

7. The photovoltaic power generation prediction method considering the selection of similar days and weather types according to claim 1, wherein The calculation formulas of the mean absolute error and root mean square error are as follows: In the formula: MAE is the mean absolute error; RMSE is the root mean square error; N represents the number of data to be predicted; Z(v) and Y(v) are the true value and predicted value of the photovoltaic power generation system respectively; v represents the data to be predicted at a certain moment.

8. A photovoltaic power generation prediction system considering the selection of similar days and weather types, characterized in that, Including: Data acquisition module: Used to acquire the historical meteorological data and photovoltaic power generation data of the target power station; Quantitative evaluation module: Used to quantitatively evaluate the correlation between historical meteorological data and photovoltaic power generation data by using the Pearson coefficient method, and determine the key influencing factors of photovoltaic power generation; Comprehensive evaluation module: Used to quantitatively evaluate the degree of numerical similarity by using the Euclidean distance based on the key influencing factors of photovoltaic power generation, measure the degree of shape similarity by using a siamese neural network, calculate the comprehensive distance of the degree of numerical similarity and the degree of numerical similarity, and use it as the classification standard of the fuzzy C-means clustering algorithm to generate a set of historical similar days of the day to be predicted; Model construction module: Used to divide the historical meteorological data into two categories, ideal weather and non-ideal weather, based on the key influencing factors of photovoltaic power generation and the selection result of the historical similar days of the day to be predicted, and construct two types of photovoltaic power generation prediction models for the target power station respectively; Model determination module: Used to compare the baseline model with the prediction results of the two types of photovoltaic power generation prediction models for the target power station, and use the mean absolute error and root mean square error as evaluation indicators to determine the best photovoltaic power generation prediction model for the target power station under the two types of weather conditions; Photovoltaic prediction module: Used to predict the photovoltaic power generation of the target power station based on the best photovoltaic power generation prediction model.

9. Photovoltaic power generation prediction system device considering selection of similar days and weather types, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Improved-fuzzy-clustering-algorithm-based photovoltaic power prediction method

    CN106251001A

  • Short-term power forecasting method of photovoltaic power station based on a Kmeans-GRA-Elman model

    CN109002915A