A Photovoltaic Power Generation Prediction Method for Cold Snap Weather under Semi-Supervised Learning

The outliers and missing values ​​in photovoltaic power generation data are processed through semi-supervised learning methods, and data augmentation and screening are used for deep learning models, which solves the problems of degraded data quality and insufficient generalization capabilities in the existing technology, and achieves higher prediction accuracy and more complex timing relationship capture.

CN119853021BActive Publication Date: 2025-06-13NANJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510318637.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The prior art ignores outliers and missing values ​​in the data in the prediction of photovoltaic power generation under major weather conditions, resulting in a decline in data quality and affecting prediction accuracy. At the same time, it is difficult for shallow models to capture complex timing relationships.

Method used

The semi-supervised learning method was used to process outliers and missing values ​​through Grubbs test and segmented cube Hermite interpolation method, and data augmentation and screening were used using SMOTE algorithm and K-medoids clustering, and photovoltaic power generation prediction was performed by combining WOA-optimized CNN-LSTM model.

Benefits of technology

It improves data quality and generalization capabilities of models, can effectively deal with short-term power fluctuations in major weather conditions, capture more complex timing relationships, and improve prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119853021B_ABST
    Figure CN119853021B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting photovoltaic power generation under cold wave weather with semi-supervised learning, which comprises the following steps: (1) obtaining the historical power generation data and the synchronous meteorological data of a photovoltaic power station, and defining the cold wave discrimination criteria based on the characteristics of sudden temperature drop; (2) detecting and removing outliers by using the Grubbs test, complementing the missing data through the piecewise cubic Hermite interpolation method, and constructing a cold wave small sample data set; (3) using the SMOTE algorithm to expand the data of the cold wave small samples, screening high-confidence samples by combining K-medoids clustering, and training and initializing the CNN-LSTM model with the labeled data; (4) optimizing the parameters of the CNN-LSTM model by using the whale optimization algorithm WOA, finally outputting the photovoltaic power prediction result and conducting error verification; the present invention has important significance for the technical research in the fields of new energy power prediction, power grid dispatching, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power prediction, and particularly relates to a photovoltaic power generation prediction method under semi-supervised learning for cold wave weather. Background Art

[0002] In recent years, affected by multiple factors such as the intensification of the greenhouse effect, major weather events such as extreme precipitation and extreme weather have occurred frequently across the country, and heavy rain, strong wind, cold wave, etc. affect the prediction of new energy power generation output. With the further sharp increase in the new energy installed capacity, especially the proposal of the goal of building a new power system, the impact of major weather processes on the balance of the power system has become increasingly prominent. Therefore, how to achieve the prediction of the output of photovoltaic power stations under major weather has become a technical problem to be solved urgently.

[0003] In order to achieve the prediction of photovoltaic power generation power under major weather, the existing technology performs grey relational analysis on the environmental factors of photovoltaic power stations to extract typical meteorological factors; then extracts the differential load characteristics of power generation according to the improved GMM clustering algorithm, and uses VMD variational mode decomposition to smooth the photovoltaic power station power prediction data; finally, according to the KELM kernel limit conversion operation problem, models and predicts the modal functions of each scale, but there are the following technical problems:

[0004] In the existing technical solution, grey relational analysis is used to extract key meteorological factors for meteorological data, and the GMM clustering algorithm is used to extract the differential load characteristics of power generation, ignoring the existence of outliers and missing values in the data. Direct use will lead to a decline in data quality and affect the prediction accuracy.

[0005] The existing technical solution ignores the problem of the scarcity of relevant photovoltaic data and meteorological data under major weather conditions, and has insufficient adaptability to rare weather events, affecting the prediction generalization ability; the VMD variational mode decomposition used requires manual setting of the decomposition layer number, and it is difficult to effectively model the short-term power fluctuations under major weather conditions; at the same time, the KELM adopted is a shallow model, and it is difficult to capture complex time series relationships. Summary of the Invention

[0006] To solve the above problems in the existing technology, the present invention proposes a photovoltaic power generation prediction method under semi-supervised learning for cold wave weather, which solves the problems of missing and abnormal data samples; solves the problem of scarce sample data under major weather conditions and improves the generalization ability of the model; at the same time, adopts a deep learning model, which can even cope with the short-term power fluctuations under major weather conditions and capture more complex time series relationships.

[0007] To achieve the purpose of the present invention, the present invention adopts the following technical solutions: A photovoltaic power generation prediction method under semi-supervised learning for cold wave weather described in the present invention includes the following steps:

[0008] (1) Obtain the historical power generation data and synchronous meteorological data of the photovoltaic power station, define the cold wave discrimination criteria based on the characteristics of sudden temperature drop, and divide the historical data into conventional weather and cold wave weather data sets;

[0009] (2) Conduct data preprocessing on the cold wave samples, use Grubbs test to detect and remove outliers, and complete the missing data through piecewise cubic Hermite interpolation method to construct a small cold wave sample data set;

[0010] (3) Use the SMOTE algorithm to expand the data of the small cold wave samples, combine K-medoids clustering to screen high-confidence samples, and use the labeled data to train the initialized CNN-LSTM model;

[0011] (4) Generate pseudo-labels for the unlabeled data through the initialized model, fuse the labeled data and high-confidence pseudo-label data, use the whale optimization algorithm WOA to optimize the parameters of the CNN-LSTM model, and finally output the photovoltaic power prediction results and conduct error verification.

[0012] Further, in step (2), the calculation process of the Grubbs test satisfies the following formula:

[0013] ;

[0014] ;

[0015] where: G is the Grubbs statistic of the data points in the sample under cold wave weather; is the i-th data point of the sample under cold wave weather; N is the total number of data points in the sample under cold wave weather; is the Grubbs test threshold under cold wave weather; t is the critical value of the distribution under the sample degrees of freedom N-2 and significance level.

[0016] Further, in step (2), the implementation process of the piecewise cubic Hermite interpolation method includes:

[0017] Construct the time series and the photovoltaic feature quantity series ,

[0018] Calculate the interpolation function:

[0019] ;

[0020] ;

[0021] ;

[0022] where:

[0023] T is the time series in the sample under cold snap weather after removing outliers;

[0024] P is a certain photovoltaic feature quantity series in the sample under cold snap weather after removing outliers;

[0025] is the nth time point in the sample under cold snap weather after removing outliers;

[0026] is a certain photovoltaic feature quantity corresponding to the nth time point in the sample under cold snap weather after removing outliers;

[0027] n is a certain position in the sample under cold snap weather after removing outliers;

[0028] i is a certain position less than n in the sample under cold snap weather after removing outliers;

[0029] is a certain time point less than in the sample under cold snap weather after removing outliers;

[0030] is the derivative of the function curve at the time point in the sample under cold snap weather after removing outliers;

[0031] is the correction term at the time point in the sample under cold snap weather after removing outliers;

[0032] is the basic interpolation at the time point in the sample under cold snap weather after removing outliers;

[0033] is the fitting interpolation at the time point in the sample under cold snap weather after removing outliers.

[0034] Furthermore, in step (3), the SMOTE data augmentation process specifically includes: calculating the Euclidean distance between the sample and its k-nearest neighbor samples in the feature space:

[0035] ;

[0036] ;

[0037] Generating new samples:

[0038] ;

[0039] where:

[0040] is a sample point in the small sample dataset under cold snap weather;

[0041] is the time corresponding to the sample point in the small sample dataset under cold snap weather;

[0042] is the sample point in the small sample dataset under cold snap weather corresponding wind direction angle;

[0043] is the sample point in the small sample dataset under cold snap weather corresponding wind speed;

[0044] is the sample point in the small sample dataset under cold snap weather corresponding rainfall;

[0045] is the sample point in the small sample dataset under cold snap weather corresponding temperature;

[0046] is the sample point in the small sample dataset under cold snap weather corresponding humidity;

[0047] is the sample point in the small sample dataset under cold snap weather corresponding air pressure;

[0048] is the sample point in the small sample dataset under cold snap weather corresponding light intensity;

[0049] is the Euclidean distance between two sample points in the small sample dataset under cold snap weather;

[0050] is a random number for controlling the distance between the new sample and the original sample, and the value range is .

[0051] Furthermore, in step (3), the silhouette coefficient is used for the K-medoids clustering quality evaluation:

[0052] ;

[0053] ;

[0054] ;

[0055] where:

[0056] i is a sample in the augmented sample dataset under cold wave weather;

[0057] j is a sample in the augmented sample dataset under cold wave weather;

[0058] is the cluster where sample i is located in the augmented sample dataset under cold wave weather;

[0059] is the total number of samples of sample i in the cluster where the sample is located in the augmented sample dataset under cold wave weather;

[0060] is the average distance from sample i to other samples in the same cluster in the augmented sample dataset under cold wave weather;

[0061] is a certain feature quantity in sample i in the augmented sample dataset under cold wave weather;

[0062] is the average distance of the nearest cluster of sample i in the augmented sample dataset under cold wave weather;

[0063] is the silhouette coefficient value.

[0064] Further, in step (4), pseudo-labels of unlabeled data are generated through the initialized model, and the labeled data and high-confidence pseudo-label data are fused, specifically as follows: The initialized trained CNN-LSTM model is used to predict the unlabeled cold wave weather sample data to obtain the corresponding prediction output; The formula is as follows:

[0065] ;

[0066] ;

[0067] ;

[0068] Where:

[0069] is the x-th feature map of the v-th convolutional layer in the model;

[0070] is the activation function in the model;

[0071] is the feature map set of the input layer in the model

[0072] is the weight of the x-th feature map and the -th feature map of the v-th convolutional layer in the model;

[0073] is the additive bias of the th feature map of the v-th convolutional layer in the model;

[0074] is the th feature map of the v-th pooling layer in the model;

[0075] is the multiplicative bias of the th feature map of the v-th pooling layer in the model;

[0076] is the pooling function used in the model;

[0077] is the th neuron in the v-th layer of the LSTM hidden layer in the model.

[0078] Furthermore, in step (4), the confidence of all prediction samples is calculated, and the data with a confidence higher than the set threshold is screened out and assigned its prediction label as the pseudo-label; the formula is as follows:

[0079] ;

[0080] where:

[0081] is the variance of the prediction sample;

[0082] P is the number of samples of the prediction sample;

[0083] is a certain prediction value of the prediction sample;

[0084] k is the number of forward propagation times of the input value of the unlabeled cold wave weather sample data;

[0085] is the uncertainty threshold;

[0086] True indicates that the sample is a pseudo-label;

[0087] False indicates that the sample cannot be used as a pseudo-label.

[0088] Furthermore, in step (4), the WOA optimization process is as follows:

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] Wherein:

[0095] is the solution of the th whale in the model;

[0096] D is the dimension of the problem in the model;

[0097] is the position of the

[0098] th whale in the model on the D-th dimension;

[0099] is the current overall optimal solution of the model;

[0100] is the solution in the model corresponding to the objective function value;

[0101] is a certain whale in the model;

[0102] is the th whale's new position in the model;

[0103] is the th whale's current optimal solution in the model;

[0104] A is the coefficient in the model that controls the exploration behavior;

[0105] is the distance between the current whale and the optimal solution in the model;

[0106] C is the coefficient in the model that controls the exploitation behavior;

[0107] r is a random function in the model, with a range of ;

[0108] u is the coefficient in the model that balances the exploration behavior and the exploitation behavior, gradually decreasing from the maximum value to 0.

[0109] Furthermore, in step (4), the error verification metrics include:

[0110] ;

[0111] ;

[0112] ;

[0113] Wherein:

[0114] is the number of sample points predicted using the model;

[0115] is the predicted value of predicting the th sample point using the model;

[0116] is the th actual value of the sample point;

[0117] is the average value of the actual values.

[0118] Furthermore, the semi-supervised learning process specifically includes: training a benchmark model using 30% labeled data in the initialization stage, retaining high-confidence samples with a predicted variance σ² < 0.05 in the pseudo-label generation stage, and fusing 50% of the original labeled data and 50% of the pseudo-label data for model optimization in the final training stage.

[0119] Beneficial effects: Compared with the prior art, the Grubbs test adopted by the present invention can effectively detect and eliminate outliers, while the piecewise cubic Hermite interpolation method can accurately complete missing data to ensure data quality; in the process of small sample data augmentation, the SMOTE method combined with K-medoids clustering adopted can optimize the data distribution and improve data reliability; in the process of photovoltaic power prediction, the CNN-LSTM model optimized by WOA adopted can effectively extract time series features and improve prediction accuracy. The cold wave weather photovoltaic power prediction method proposed by the present invention has the advantages of accurate outlier detection, effective small sample augmentation, and high prediction accuracy, and is of great significance for technical research in fields such as new energy power prediction and power grid dispatching. Brief Description of the Drawings

[0120] Figure 1 is the flowchart of the present invention;

[0121] Figure 2 is the line chart of the daily maximum temperature and minimum temperature in November of the present invention. Detailed Embodiments

[0122] The present invention will be further explained and illustrated below in conjunction with the drawings and specific embodiments. It should be understood that the embodiments are only used to illustrate and explain the present invention, and do not impose any limitation on the scope of implementation of the present invention.

[0123] The embodiment of the present invention provides a cold wave weather photovoltaic power prediction method under semi-supervised learning, including the following steps:

[0124] (1)Obtain the historical power generation data of the PV power station from September to November in recent years and the meteorological data during the same period. The meteorological data includes but is not limited to core meteorological parameters such as temperature and humidity. Then, according to the characteristic trends and mutation situations of the meteorological data, define the discrimination criteria for cold wave weather events. After the discrimination criteria are clarified, divide the historical data of the PV power station into normal weather and cold wave weather;

[0125] As Figure 2 shown, the discrimination criteria for cold wave weather are as follows: When it is monitored that the temperature drops by ≥8°C within 24 hours and the minimum temperature ≤5°C, it is determined as the start of the cold wave; when the temperature remains >5°C for more than 12 hours, it is determined as the end of the cold wave.

[0126] (2)Preprocess the selected cold wave sample data. Check the data for outliers through Grubbs test, and then supplement the missing values generated by the above method through piecewise cubic Hermite interpolation method to complete the establishment of a small sample data set under cold wave weather;

[0127] (2.1)Use Grubbs test to screen the samples under the cold wave weather. Assume that the data is normally distributed, and judge whether it is an outlier by comparing the deviation between the data point and the data mean value to obtain the samples under the cold wave weather after removing the outliers. The formula for Grubbs test is as follows:

[0128] ;

[0129] ;

[0130] where: G is the Grubbs statistic of the data point in the sample under the cold wave weather; is the i-th data point of the sample under the cold wave weather; N is the total number of data points in the sample under the cold wave weather; is the Grubbs test threshold under the cold wave weather; t is the critical value distributed according to the sample degree of freedom N - 2 and the significance level.

[0131] (2.2)Use the piecewise cubic Hermite interpolation method (PCHIP) to supplement the missing values for the samples under the cold wave weather after removing the outliers to obtain a small sample data set under the cold wave weather. Among them, the formula for the Hermite interpolation method is as follows:

[0132] Construct the time series and the PV characteristic quantity series ,

[0133] Calculate the interpolation function:

[0134] ;

[0135] ;

[0136] ;

[0137] Wherein:

[0138] T is the time series in the sample under cold wave weather after removing outliers;

[0139] P is a certain photovoltaic characteristic quantity series in the sample under cold wave weather after removing outliers;

[0140] is the nth time point in the sample under cold wave weather after removing outliers;

[0141] is a certain photovoltaic characteristic quantity corresponding to the nth time point in the sample under cold wave weather after removing outliers;

[0142] n is a certain position in the sample under cold wave weather after removing outliers;

[0143] i is a certain position less than n in the sample under cold wave weather after removing outliers;

[0144] is a certain time point less than in the sample under cold wave weather after removing outliers;

[0145] is the derivative of the function curve at the time point in the sample under cold wave weather after removing outliers;

[0146] is the correction term at the time point in the sample under cold wave weather after removing outliers;

[0147] is the basic interpolation at the time point in the sample under cold wave weather after removing outliers;

[0148] is the fitting interpolation at the time point in the sample under cold wave weather after removing outliers.

[0149] (3) Expand the small sample data based on the SMOTE model, merge the expanded sample with the original small sample data, use K-medoids clustering to detect the data quality, initially screen out the data samples with higher confidence, and use part of them as labeled data to train the initialized CNN-LSTM model;

[0150] (3.1)For the small sample dataset under the cold snap weather, by analyzing the characteristics of the original data, the SMOTE algorithm is used to generate new synthetic sample data based on the K-nearest neighbor strategy in the feature space, expand the sample quantity under the cold snap weather, and obtain the expanded sample dataset under the cold snap weather. Among them, the sample expansion formula based on SMOTE is as follows:

[0151] ;

[0152] ;

[0153] Generate new samples:

[0154] ;

[0155] Among them:

[0156] is a sample point in the small sample dataset under the cold snap weather;

[0157] is the time corresponding to the sample point in the small sample dataset under the cold snap weather;

[0158] is the wind direction angle corresponding to the sample point in the small sample dataset under the cold snap weather;

[0159] is the wind speed corresponding to the sample point in the small sample dataset under the cold snap weather;

[0160] is the rainfall corresponding to the sample point in the small sample dataset under the cold snap weather;

[0161] is the temperature corresponding to the sample point in the small sample dataset under the cold snap weather;

[0162] is the humidity corresponding to the sample point in the small sample dataset under the cold snap weather;

[0163] is the air pressure corresponding to the sample point in the small sample dataset under the cold snap weather;

[0164] is the light intensity corresponding to the sample point in the small sample dataset under the cold snap weather;

[0165] is the Euclidean distance between two sample points in the small sample dataset under cold snap weather;

[0166] is a random number for controlling the distance between the new sample and the original sample, and its value range is .

[0167] (3.2) Apply the K-medoids clustering algorithm to the expanded sample dataset under cold snap weather for data quality detection. According to the distribution of samples in the feature space, divide the data into several clusters, select their central points as typical samples, preliminarily screen out sample data with higher confidence, and label some of them to obtain the sample dataset under cold snap weather with labeled data. Among them, the formula for the silhouette coefficient value evaluating the quality of the K-medoids clustering algorithm is as follows:

[0168] ;

[0169] ;

[0170] ;

[0171] Among them:

[0172] i is a certain sample in the expanded sample dataset under cold snap weather;

[0173] j is a certain sample in the expanded sample dataset under cold snap weather;

[0174] is the cluster where sample i is located in the expanded sample dataset under cold snap weather;

[0175] is the total number of samples of sample i in the cluster where the sample is located in the expanded sample dataset under cold snap weather;

[0176] is the average distance from sample i to other samples in the same cluster in the expanded sample dataset under cold snap weather;

[0177] is a certain feature quantity in sample i in the expanded sample dataset under cold snap weather;

[0178] is the average distance of the nearest cluster to sample i in the expanded sample dataset under cold snap weather;

[0179] is the silhouette coefficient value.

[0180] (4) Use the preliminarily trained model to predict the unlabeled data, then screen out the samples with high confidence as pseudo-labels, combine the labeled data and the pseudo-label data, retrain the CNN-LSTM model optimized by WOA to predict the photovoltaic output, and verify the model accuracy through data metrics.

[0181] (4.1.1) Use the initialized and trained CNN-LSTM model to predict the unlabeled cold wave weather sample data and obtain the corresponding prediction output. Among them, the relevant formulas of the CNN-LSTM model are as follows:

[0182] ;

[0183] ;

[0184] ;

[0185] Among them:

[0186] is the x-th feature map of the v-th convolutional layer in the model;

[0187] is the activation function in the model; the Relu function is selected in the present invention;

[0188] is the feature map set of the input layer in the model;

[0189] is the weight between the x-th feature map of the v-th convolutional layer in the model and the -th feature map

[0190] is the additive bias of the -th feature map of the v-th convolutional layer in the model;

[0191] is the -th feature map of the v-th pooling layer in the model;

[0192] is the multiplicative bias of the -th feature map of the v-th pooling layer in the model;

[0193] is the pooling function used in the model;

[0194] is the -th neuron of the v-th layer of the LSTM hidden layer in the model.

[0195] After obtaining the corresponding model output, based on the confidence evaluation mechanism, calculate the confidence of all prediction samples, and screen out the data with a confidence higher than the set threshold, and assign its prediction label as a pseudo-label. Among them, the uncertainty estimation formula based on the prediction result is as follows:

[0196] ;

[0197] Among them:

[0198] is the variance of the prediction sample;

[0199] P is the number of samples of the prediction sample;

[0200] is a certain predicted value of the prediction sample;

[0201] k is the number of forward propagation times of the input value of the unlabeled cold wave weather sample data;

[0202] is the uncertainty threshold;

[0203] True indicates that the sample is a pseudo-label;

[0204] False indicates that the sample cannot be used as a pseudo-label.

[0205] (4.2.1)Merge the screened high-confidence pseudo-label samples with the original labeled data, and use the CNN-LSTM model optimized by WOA to find the optimal network parameters through iterative update to predict the photovoltaic output. Among them, the formula of the CNN-LSTM model optimized by WOA is as follows:

[0206] ;

[0207] ;

[0208] ;

[0209] ;

[0210] ;

[0211] Among them:

[0212] is the th whale solution in the model;

[0213] D is the dimension of the problem in the model;

[0214] is the position of the th whale in the Dth dimension of the model;

[0215] M is the number of whales in the model;

[0216] is the current total optimal solution of the model;

[0217] is the solution in the model corresponding objective function value;

[0218] is a certain whale in the model;

[0219] is the new position of the whale in the model;

[0220] is the current optimal solution of the whale in the model;

[0221] A is the coefficient controlling the exploration behavior in the model;

[0222] is the distance between the current whale and the optimal solution in the model;

[0223] C is the coefficient controlling the exploitation behavior in the model;

[0224] r is a random function in the model, with a range of ;

[0225] u is the coefficient controlling the balance between exploration behavior and exploitation behavior in the model, gradually decreasing from the maximum value to 0.

[0226] (4.2.2) Perform error verification on the photovoltaic output prediction results of the CNN-LSTM model optimized by the WOA. The present invention mainly uses indicators such as mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination for evaluation. The relevant calculation formulas are as follows:

[0227] ;

[0228] ;

[0229] ;

[0230] Among them:

[0231] is the number of sample points predicted by the model;

[0232] is the The predicted value of a sample point;

[0233] is the actual value of the sample point;

[0234] is the average value of the actual values.

Claims

1. A method for predicting photovoltaic power generation in cold weather under semi-supervised learning, characterized in that: The following steps are involved: (1) Obtain historical power generation data of photovoltaic power stations and meteorological data of the same period, define cold wave discrimination criteria based on the characteristics of sudden temperature drop, and divide historical data into regular weather and cold wave weather data sets; (2) Data preprocessing was performed on the cold wave samples. Outliers were detected and removed using the Grubbs test. Missing data were filled in using the piecewise cubic Hermite interpolation method to construct a small sample dataset of cold waves. (3) The SMOTE algorithm was used to expand the data of the small sample of cold waves, and K-medoids clustering was used to screen high-confidence samples. The labeled data was used to train and initialize the CNN-LSTM model. (4) Generate pseudo labels for unlabeled data by initializing the model, fuse the labeled data with high-confidence pseudo-label data, and use the whale optimization algorithm (WOA) to optimize the CNN-LSTM model parameters. Finally, output the PV output prediction results and perform error verification.

2. According to the method for predicting photovoltaic power generation in cold weather under semi-supervised learning in claim 1, it is characterized in that: In step (2), the calculation process of the Grubbs test satisfies the following formula: ; ; Where: G is the Grubbs statistic of the data points in the sample under cold wave weather; is the i-th data point of the sample under cold wave weather; N is the total number of data points in the sample under cold wave weather; is the Grubbs test threshold under cold wave weather; t is the critical value distributed according to the sample degrees of freedom N-2 and the significance level.

3. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1 is characterized in that: In step (2), the implementation process of the piecewise cubic Hermite interpolation method includes: Constructing a time series and photovoltaic characteristic series , Compute the interpolation function: ; ; ; in: T is the time series in the sample under cold wave weather after removing outliers; P is a photovoltaic characteristic sequence in the sample under cold wave weather after removing outliers; is the nth time point in the sample under cold wave weather after removing outliers; is a photovoltaic characteristic quantity corresponding to the nth time point in the sample under cold wave weather after removing outliers; n is a certain location in the sample under cold wave weather after removing outliers; i is a position less than n in the sample under cold wave weather after removing outliers; is the number of samples less than 100% in cold wave weather after removing outliers. a certain point in time; In the sample under cold wave weather after removing outliers The derivative of the function curve at a time point; In the sample under cold wave weather after removing outliers Correction items at time points; In the sample under cold wave weather after removing outliers Basic interpolation of time points; In the sample under cold wave weather after removing outliers Fitted interpolation of time points.

4. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1 is characterized in that: In step (3), the SMOTE data expansion process specifically includes: calculating samples in the feature space With k nearest neighbor samples The Euclidean distance of: ; ; Generate new samples: ; in: is a sample point in a small sample data set under cold wave weather; The time corresponding to the sample point in the small sample data set under cold wave weather; The sample points of the small sample data set under cold wave weather The corresponding wind direction angle; The sample points of the small sample data set under cold wave weather The corresponding wind speed; The sample points of the small sample data set under cold wave weather The corresponding rainfall amount; The sample points of the small sample data set under cold wave weather The corresponding temperature; The sample points of the small sample data set under cold wave weather The corresponding humidity; The sample points of the small sample data set under cold wave weather The corresponding air pressure; The sample points of the small sample data set under cold wave weather The corresponding illumination amplitude; A random number that controls the distance between the new sample and the original sample, with a value range of .

5. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: In step (3), the K-medoids clustering quality is evaluated using the silhouette coefficient: ; ; ; in: i is a sample in the expanded sample data set under cold wave weather; j is a sample in the expanded sample data set under cold wave weather; is the cluster where sample i is located in the expanded sample data set under cold wave weather; is the total number of samples i in the cluster where the samples are located in the sample data set under the expanded cold wave weather; is the average distance from sample i to other samples in the same cluster in the expanded sample data set under cold wave weather; is a characteristic quantity of sample i in the expanded sample data set under cold wave weather; is the average distance to the nearest cluster of sample i in the expanded sample data set under cold wave weather; is the silhouette coefficient value.

6. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: In step (4), the pseudo labels of the unlabeled data are generated by initializing the model, and the labeled data and the high-confidence pseudo-label data are fused. Specifically, the unlabeled cold wave weather sample data is predicted using the initialized trained CNN-LSTM model to obtain the corresponding prediction output; the formula is as follows: ; ; ; in: is the xth feature map of the vth convolutional layer in the model; is the activation function in the model; It is the feature atlas of the input layer in the model; is the xth feature map of the vth convolutional layer in the model and the The weight of the feature map; is the vth convolutional layer in the model. Additive bias of feature maps; is the vth pooling layer in the model feature maps; is the vth pooling layer in the model The multiplicative bias of the feature maps; is the pooling function used in the model; is the vth layer of the LSTM hidden layer in the model. A neuron.

7. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: In step (4), the confidence of all predicted samples is calculated, and the data with confidence higher than the set threshold is screened out, and its predicted label is assigned as a pseudo label; the formula is as follows: ; in: is the variance of the prediction sample; P is the number of samples of the prediction sample; is a predicted value of the predicted sample; k is the number of forward propagation of the unlabeled cold wave weather sample data input value; is the uncertainty threshold; True means the sample is a pseudo label; False means that the sample cannot be used as a pseudo label.

8. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: In step (4), the WOA optimization process is as follows: ; ; ; ; ; in: For the model Head whale solution; D is the dimension of the problem in the model; is the position of the first whale in the model in the Dth dimension; M is the number of whales in the model; is the current optimal solution of the model; Solution for the model The corresponding objective function value; For a certain whale in the model; For the model the new position of the first whale; For the model The current optimal solution for the first whale; A is the coefficient of control exploration behavior in the model; is the distance between the current whale and the optimal solution in the model; C is the coefficient of control development behavior in the model; r is the random function in the model, ranging from ; u is the coefficient that controls the balance between exploration behavior and development behavior in the model, which gradually decreases from the maximum value to 0.

9. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: In step (4), the error verification indicators include: ; ; ; in: The number of sample points predicted by the model; To use the model to predict The predicted value of sample points; For the The actual value of the sample points; is the average of the actual values.

10. The method for predicting photovoltaic power generation in cold weather under semi-supervised learning according to claim 1, characterized in that: The semi-supervised learning process specifically includes: using 30% labeled data to train the benchmark model in the initialization phase, retaining high-confidence samples with prediction variance σ²<0.05 in the pseudo-label generation phase, and integrating 50% of the original labeled data and 50% of the pseudo-label data for model optimization in the final training phase.

Citation Information

Patent Citations

  • Cold and strong wind weather sample expansion method and system, storage medium and equipment

    CN117151488A

  • Pollution source over-limit emission studying and judging method based on pseudo-label semi-supervised learning

    CN118627920A