Distributed photovoltaic random output scene generation method and system

Through the combination of graph convolution generation adversarial networks and clustering algorithms, the problem of difficulty in capturing the spatial and temporal correlation of photovoltaic unit output is solved in the existing technology, and more efficient and accurate photovoltaic output scenario generation is achieved, and the generated scenarios are more representative and diverse.

CN119988857AActive Publication Date: 2025-05-13SICHUAN UNIV

Patent Information

Application Number
CN202411866902.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-13
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The existing photovoltaic output scenario generation method is difficult to effectively capture the spatial and temporal correlation in photovoltaic unit output, resulting in a large deviation from the generated scene and the actual scene, and cannot meet the requirements of scene diversity.

Method used

Graph convolution generation adversarial network (GCGAN) combined with ST-DBSCAN and K-medoids clustering algorithms is used to mine high-dimensional nonlinear features of historical data through implicit modeling methods, learn the time-space correlation in photovoltaic unit output, and generate representative photovoltaic output scenarios.

Benefits of technology

The accuracy of the generation of uncertain scenarios for photovoltaic output is improved, and the generated scenarios are more complete and comprehensive, which can effectively preserve the spatial and temporal correlations between photovoltaic station outputs and meet the requirements of scene diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988857A_ABST
    Figure CN119988857A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed photovoltaic random output scene generation method. The method comprises the following steps: S1, collecting historical output data of each photovoltaic station; s2, inputting the processed historical output data into a graph convolutional generative adversarial network, and training the graph convolutional generative adversarial network; s3, inputting the collected output data of each photovoltaic station into the trained graph convolution generative adversarial network, generating a photovoltaic output scene, and completing the data enhancement of the output scene; s4, clustering the power data of the photovoltaic power station by using an ST-DBSCAN clustering algorithm, selecting typical photovoltaic stations, and splicing the output of each typical photovoltaic station according to the time dimension; and S5, selecting a typical output date by using a K-medoids clustering method, and selecting and generating an output scene to obtain a typical output scene with a plurality of photovoltaic sites. According to the invention, the distributed photovoltaic output scene is more comprehensive and diversified, and the accuracy of the distributed photovoltaic output scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of photovoltaic scene generation, and in particular to a distributed photovoltaic random output scene generation method and system. Background Art

[0002] With the increasing penetration of renewable energy, especially large-scale integration, the uncertainty of its power generation has brought great challenges to the operation and planning of power systems. The scenario generation method is an effective method to deal with the uncertainty of renewable energy generation, and is widely used in stochastic optimization problems such as power system economic dispatch, energy storage configuration, and market bidding. The basic idea of ​​the scenario generation method is to convert various uncertain factors into a combination of multiple deterministic factors, each combination represents a scenario. Then optimize each deterministic scenario to reduce the difficulty of modeling and solving. The key to the scenario generation method is how to select multiple representative scenarios to ensure the accuracy of the model.

[0003] Scenario generation methods include scenario generation and scenario reduction. The academic community has proposed many methods for the study of scenario generation. The most commonly used method is the probability model-based method, which establishes a probability distribution that conforms to the characteristics of random power generation output by analyzing historical data, and generates scenarios by sampling. Pinson P uses Gaussian Copula to fit the interdependent structure of wind power prediction error and prediction distribution based on the probability prediction of wind power. In the scenario analysis research of multiple power generation locations, Li Jinghua considered the spatial correlation of power generation of multiple wind farms, avoided the direct construction of the joint distribution function of multiple wind farms, and proposed a joint scenario simulation method based on Copula to achieve effective characterization of power generation scenarios in multiple geographical locations; Jin Tianran established a wind power and photovoltaic time correlation model based on multivariate normal, and at the same time established a Copula model of vine structure, and used multidimensional variable vine structure to sample and generate scenarios. The existing multi-site joint output based on model probability often generates the output of a single photovoltaic site first, and then uses the copula function to construct the joint probability function between multiple sites. However, the scenario analysis method based on probability model often needs to make assumptions about the regular characteristics of unknown distribution data based on statistical experience, and requires manual selection of model functions. The rationality of the probability assumption is related to the quality and reliability of the scenario analysis. The scenarios generated often deviate greatly from the actual scenarios and cannot meet the requirements of scenario diversity.

[0004] Another popular method is the time series method. Zou Bin uses the Auto Regressive Moving Average (ARMA) model to characterize wind speed and simulate random production based on the effective capacity distribution of generator sets; Lei Yu uses the ARMA model to model the dynamic power output of wind power generation to generate an initial scenario set. In the study of scenario simulation at multiple power generation locations, Duong D.Le uses the time series method to generate scenarios of multiple spatially related locations, so that the generated scenarios can maintain the temporal correlation of the time series and characterize the spatial correlation of different geographical locations. However, the autoregressive model and state space representation are prone to problems of model overfitting and pattern recognition errors, and it is difficult to use time series models to characterize the diversity of renewable energy generation dynamics.

[0005] In recent years, machine learning methods have also begun to be used in scenario generation. Vagropoulos SI, proposed the use of artificial neural networks (ANN) to generate scenarios and used them for scenario generation of power load, solar energy and wind power; Cui Mingjian used neural networks to generate scenarios and used the generated scenarios to predict the ramp events of wind power. With the development of deep learning, more and more generative models based on deep learning are used in scenario generation; Yize Chen proposed a renewable energy scenario generation method based on generative adversarial networks, which can correctly capture the temporal and spatial correlation of solar and wind power generation, and can efficiently generate data that conforms to the observation characteristics by adding conditional label information; Dong Xiaochong proposed a day-ahead scenario generation method based on conditional generative adversarial networks to characterize the uncertainty of the day-ahead output of renewable energy by the relationship between the noise distribution that meets the prediction conditions and the prediction scenario. Generative models provide a new solution for capturing the true distribution of data. There is no need to analyze the distribution characteristics of data and no probability assumptions are required. The unknown distribution of historical data can be learned in an unsupervised manner, and data samples consistent with the distribution characteristics of training data can be generated.

[0006] In the research on the reduction of photovoltaic power stations and wind farms, Wang Qun, Yang Xian and others used the improved K-medoids clustering and improved K-means clustering algorithms to reduce the scenes, improving the speed and accuracy of scene reduction; J. Dupaová proposed the predecessor and descendant algorithm based on probability distance, and Heitsch H improved the algorithm on the basis of J. Dupaová and proposed a fast predecessor elimination algorithm for scene reduction to improve the calculation speed and accuracy; Heitsch H also proposed a synchronous back generation reduction method, which is now widely used in scene reduction. However, the existing photovoltaic output scene reduction often only reduces the time series dimension or the space dimension separately, ignoring the temporal and spatial correlation between each photovoltaic unit. Summary of the invention

[0007] The purpose of the present invention is to provide a distributed photovoltaic random output scenario generation method, which is aimed at studying the power system problems under the background of a high proportion of renewable energy. The large number of original photovoltaic output scenarios generated are reduced to a few representative scenarios to achieve the purpose of streamlining data and reducing the subsequent calculation amount, and the uncertainty problem is converted into an effective deterministic problem to be solved, providing a data basis for subsequent power system planning, operation, scheduling and other problems.

[0008] To achieve the above object, the present invention provides the following technical solution: a method for generating a distributed photovoltaic random output scenario, comprising the following steps:

[0009] S1. Collect historical output data of each photovoltaic site and normalize and standardize the historical output data;

[0010] S2. Use the noise as the input of the generator in the graph convolutional generative adversarial network, use the generator to generate a large amount of output data, use the output data generated by the generator and the processed historical output data as the input of the discriminator in the graph convolutional generative adversarial network, train the graph convolutional generative adversarial network until the discriminator cannot determine whether the data is generated by the generator or the real historical output data, and the generative adversarial network reaches Nash equilibrium, and it is believed that the samples generated by the generator are basically consistent with the real data in statistical distribution;

[0011] The noise is data sampled from a simple distribution, where the simple distribution includes a Gaussian distribution or a uniform distribution;

[0012] S3, input the noise into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios and complete data enhancement of the output scenarios;

[0013] S4. Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish the densely concentrated data from the scattered data, select typical photovoltaic sites, and splice the output of each typical photovoltaic site according to the time dimension;

[0014] S5. Use the K-medoids clustering method to select typical output dates, select output scenarios from the completely generated photovoltaic output scenarios according to the dates, and obtain typical photovoltaic output scenarios.

[0015] In some embodiments, in S1, the historical output data of the photovoltaic sites are divided into rainy season and dry season, the data format is set to a three-dimensional matrix, and the photovoltaic power information of all sites on each day is recorded in the form of a two-dimensional matrix {x n,t}.

[0016] In some embodiments, the layers in the generator and discriminator of the graph convolutional generative adversarial network both use graph convolutional networks, and both the generator and the discriminator include graph filters in the spatial dimension and feature filters in the temporal dimension. Graph filters are used in the spatial dimension to capture the spatial correlation between photovoltaic sites; in the temporal dimension, one-dimensional convolution filters are used to process temporal features to complete feature extraction.

[0017] In some embodiments, the steps include:

[0018] Use the input data matrix as parameters for feature filters:

[0019]

[0020] The last layer uses the sigmoid function to generate output values ​​in the range of (0,1) as photovoltaic output data;

[0021] Through the feature aggregation of adjacent nodes of the graph convolution, the correlation coefficient after exponential transformation is used as the weight of the graph filter. The calculation expression is:

[0022]

[0023] The matrix multiplication term in formula (6) Considered as the output of the feature filter, using a length of And a one-dimensional convolution filter with trainable weight coefficients, the elements of the j-th column are:

[0024]

[0025] In the formula, For each layer The input data matrix is ​​a matrix of features, where L is the total number of hidden layers in the generator; σ(·) is the nonlinear activation function; A is the graph filter, is the weight matrix; C i,j is the correlation coefficient between node i and node j; m is the position offset of the sliding window in the one-dimensional convolution; is half the width of the convolution window; is the position feature matrix for m offsets; is the feature of the j-th column of the input matrix at the m-offset position.

[0026] In some embodiments, the loss function of the generator is:

[0027]

[0028] The loss function of the discriminator is:

[0029]

[0030] Where z is the noise variable; P z is a simple distribution; G(z) is the noise variable mapped to a generated data space; P r is the probability distribution of real samples; x is the generated sample; G is the generator; D is the discriminator; P G is the probability distribution of generated samples; E is the expectation.

[0031] In some embodiments, in S4, selecting a typical photovoltaic site includes the following steps:

[0032] S41, setting the input field radius Eps, the minimum number minpts and the time neighborhood Δt;

[0033] S42, starting from any unlabeled point, a certain point is selected as the core point p of the scan, and the neighborhood is scanned within the time neighborhood Δt of the core point p, and the number of neighborhood points that meet the distance threshold in the spherical area Eps(s) with the core point p as the center and the radius Eps from the neighborhood is scanned within the time neighborhood Δt;

[0034] S43, if the number of neighboring points of the core point p is less than the set threshold MinPts, mark the point as a noise point;

[0035] S44, if the number of neighborhood points is greater than the set threshold MinPts, the point is marked as a core point p, a cluster number is generated as cluster C1, and all points of the core point p in the neighborhood Eps(s) are added to the cluster C1 to be scanned;

[0036] S45, for the unmarked points in cluster C1, repeat steps S42 to S44 until all the points in cluster C1 are marked;

[0037] S46. Continue to select marked points from the data set, repeat steps S42 to S45, and cluster into another class until all points are marked.

[0038] In some embodiments, the specific method of splicing the output of each typical photovoltaic site according to the time dimension is: splicing the typical photovoltaic sites according to the corresponding dates to form a one-dimensional sequence:

[0039] V d =[P1(1,d),P1(2,d),...,P1(t,d),...,P i (1,D),P i (2,D),...,P i (t,D)] T (9);

[0040]

[0041] Where P i (t,D) is the output of the i-th typical station at the t-th hour on the D-th day; t=1,2,...,24, d=1,2,...,D, D is the total number of days, D=365; i=1,2,...,M, where M is the number of typical stations; It is a 24×M matrix.

[0042] In some embodiments, S5 includes the following steps:

[0043] S51, randomly select r sequences from all sequences as initial clustering centers;

[0044] S52, according to the principle of being closest to the cluster center, assign the remaining objects to each class, and calculate the distance between the two sequences that are closest to each other;

[0045] S53, minimizing the sum of weighted distances from all scenes to the cluster center;

[0046] S54, judging whether it converges, if not, re-performing S52, if it converges, then the r cluster centers obtained by clustering are the typical days of photovoltaic output;

[0047] S55. Calculate the probability of occurrence of each sequence after clustering.

[0048] In some embodiments, the calculation expression for calculating the distance between the two sequences with the closest distance is:

[0049]

[0050] In the formula, d(u i ,u j ) is the distance between two sequences, for each sequence u i , select the cluster center u with the smallest distance j To calculate the distance; T is the number of time points in the sequence; For scene u i The value at the tth time point; the distance d(u i ,u j ) is the sum of the absolute values ​​of the differences between the two series at each time point.

[0051] A distributed photovoltaic random output scene generation system, using the distributed photovoltaic random output scene generation method, comprises:

[0052] Data acquisition module: used to obtain the historical output data of each photovoltaic site and normalize and standardize the historical output data;

[0053] Graph convolutional generative adversarial network processing module: Use the noise as the input of the generator, and use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the graph convolutional generative adversarial network processing module to train the graph convolutional generative adversarial network;

[0054] The noise is input into the generator of the trained graph convolutional generative adversarial network processing module to generate photovoltaic output scenarios and complete data enhancement of the output scenarios;

[0055] Spatial clustering processing module: Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish densely concentrated data from scattered data, select typical photovoltaic sites, and splice the output of each typical photovoltaic site according to the time dimension;

[0056] Time clustering processing module: Use the K-medoids clustering method to select typical output dates, select output scenarios from the completely generated photovoltaic output scenarios according to the dates, and obtain typical photovoltaic output scenarios.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. The present invention mines the high-dimensional nonlinear features of historical data through implicit modeling based on the graph convolution generative adversarial network, uses graph filters in the spatial dimension, and uses one-dimensional convolution filters in the temporal dimension to complete feature extraction, learns the time-space correlation and meteorological correlation contained in the output of multiple photovoltaic units, and improves the accuracy of generating uncertain photovoltaic output scenarios. At the same time, the graph convolution generative adversarial network learns the existing photovoltaic output data distribution, so that the generator learns to simulate various potential output situations, including marginal and rare events; the discriminator promotes the generator to improve the quality, so that the generated results are closer and closer to the real output distribution, thereby generating a more complete and comprehensive photovoltaic output scenario and completing the data enhancement of the output scenario.

[0059] 2. The present invention proposes a method for reducing the output scenarios of multiple photovoltaic sites. The method uses spatial clustering to obtain typical sites, reconstructs the output data of typical sites, and then clusters the time to obtain the typical output date. The complete output of each site corresponding to the typical output date is the typical output scenario. The method reduces the scenario from the two dimensions of space and time, retains the spatial correlation and time correlation of the output of each photovoltaic site, and obtains a typical output scenario with the output data of multiple photovoltaic sites. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0061] Figure 2A schematic diagram of the data format for input into a generative adversarial network in Embodiment 1 of the present invention;

[0062] Figure 3 This is a schematic diagram of the graph convolution generative adversarial network structure in Embodiment 1 of the present invention;

[0063] Figure 4 A schematic diagram of a data reconstruction form according to an embodiment of the present invention;

[0064] Figure 5 This is a schematic diagram of the system structure in Embodiment 2 of the present invention. DETAILED DESCRIPTION

[0065] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0066] Embodiment 1

[0067] See also Figure 1-Figure 4 , a distributed photovoltaic random output scenario generation method, comprising the following steps:

[0068] 1. Photovoltaic output scenario generation based on graph convolutional generative adversarial network.

[0069] (I) Data processing:

[0070] In this embodiment, the historical output data of each photovoltaic site is firstly collected, and the historical photovoltaic data is divided into the rainy season and the dry season.

[0071] The input data for generating the photovoltaic output curve scenario is the photovoltaic output power information of N sites in one year. There are 365 days in a year. The data format is converted into a three-dimensional matrix. The data size is expressed as N×T×D, where N is the number of total photovoltaic sites; T is 24, indicating 24 hours of data; and D is 365, indicating that there are 365 days in a year. The photovoltaic power information of all sites on each day can be recorded in the form of a two-dimensional matrix {x n,t}, where n is the number of sites, t represents the time, and x n,t Represents the photovoltaic output power information of the nth site at time t, and the input matrix form is shown as follows:

[0072]

[0073] Before inputting data into the model, it is also necessary to scale the data features, map the original data to a specific range, remove the unit and unify the dimensions, and make it a pure data sample, that is, normalize and standardize the data. Data normalization is to limit the data to a certain range, usually mapping the data to the range of -1 to 1 or 0 to 1. The former is called Min-Max normalization. The commonly used Min-Max normalization calculation expression is:

[0074]

[0075] The data obtained through the above operation is in the form of Figure 1 shown.

[0076] (2) Use graph convolutional generative adversarial network (GCGAN) to generate photovoltaic output scenarios.

[0077] Generative adversarial network is a deep learning framework, which includes a generator network and a discriminator network.

[0078] The goal of the Generative Adversarial Network is to generate realistic samples that are difficult for the discriminator network to distinguish between true and false. To achieve this goal, it is necessary to define the loss function for the generator network training and the loss function for the discriminator network training.

[0079] When the discriminator network is fixed and the generator network is trained, the probability of the generated sample G(z) should be larger, and the loss function L G The smaller.

[0080] The training goal of the discriminator network is to distinguish P as much as possible. r and P G When the parameters θ of the generator network are fixed G , update the discriminator network parameters θ D When , the greater the difference in discrimination, the stronger the judgment ability of the discriminator network is, that is, maximizing the difference between E[D(·)] and E[D(G(·))]. According to the basic principle of generative adversarial networks, when the discriminator network can distinguish input samples well, its loss function is smaller.

[0081] The loss function of the generator is:

[0082]

[0083] The loss function of the discriminator is:

[0084]

[0085] Where z is the noise variable; P zis a simple distribution; G(z) is the noise variable mapped to a generated data space; P r is the probability distribution of real samples; x is the generated sample; G is the generator; D is the discriminator; P G is the probability distribution of generated samples; E is the expectation; E[D(·)] is the expectation of the discriminator's real data; E[D(G(·))] is the expectation of the discriminator accepting the generated data.

[0086] The zero-sum game between the two networks is denoted as V(D,G). For the optimization of the loss function, both networks need to update the parameters of their own network according to the update of the other network. The two networks are continuously trained alternately and their parameters are continuously updated until the adversarial network reaches Nash equilibrium. When the training converges, a pair of optimal parameters can be obtained. Because there are multiple extreme points in the high-dimensional data space, the parameters at this time and are the two extreme points of V(D,G). The principle of generating adversarial networks is a minimax optimization problem, which is defined as follows:

[0087]

[0088] In the formula, V(D,G) is a binary cross entropy function, the ultimate goal of which is to minimize the probability distribution P of the generated samples. G And the true sample probability distribution P r The JS distance between them.

[0089] like Figure 2 As shown, the layers in the generator and discriminator of the graph convolutional generative adversarial network use graph convolutional networks (GCNs), which can effectively predict node signals with graph dependencies. Each layer of the graph convolutional network uses a graph filter that embeds potential dependencies into a set of nodes that can aggregate input signals from other connected nodes. In this way, the output data of the graph convolutional network will show strong correlations between interconnected nodes.

[0090] Design Filters To match the actual spatial correlation between N PV plants. Let L be the total number of hidden layers in the generator, For each layer The input data matrix is ​​a matrix of features, where

[0091] The input of the first layer is a random noise matrix, namely X (1) = Z and K1 = K. In order to gradually increase the time dimension, set For each layer Using the Matrix As a trainable feature filter parameter, it can be expressed as:

[0092]

[0093] In the formula, σ(·) is a nonlinear activation function; this embodiment selects ReLU as the activation function of most layers, and the last layer uses the sigmoid function to generate output values ​​in the range of [-1,1] as photovoltaic output data. The final output layer can show the correlation between the features of each node.

[0094] The layers of the discriminator are similar to those of the generator. The first layer uses sample data or real data generated by the generator. As input. The difference between the graph filter A is that the number of features along with The activation functions of the remaining layers all use Leaky ReLU, and the last layer uses the sigmoid function to generate values ​​in the range of (0,1) to complete the binary classification.

[0095] Both the generator and the discriminator contain graph filters in the spatial dimension and feature filters in the temporal dimension. In the spatial dimension, graph filters are used to capture the spatial correlation between photovoltaic sites. By aggregating the features of adjacent nodes through graph convolution, the correlation coefficient after exponential transformation is used as the weight of the graph filter, as shown in formula (7). This exponential transformation can improve the significance of the filter weights between related positions.

[0096]

[0097] In the time dimension, one-dimensional convolution is used to process the time series characteristics to reduce the parameter scale and maintain the time series characteristics of photovoltaic output. Considered as the output of the feature filter. Using a length of And a one-dimensional convolution filter with trainable weight coefficients, the elements of the j-th column are shown as follows:

[0098]

[0099] In the formula, Each layer has K l The input data matrix of features is , where l∈{1,...,L}, L is the total number of hidden layers in the generator; σ(·) is the nonlinear activation function; A is the graph filter, is the weight matrix; C i,j is the correlation coefficient between node i and node j; m is the position offset of the sliding window in the one-dimensional convolution; Ml is half the width of the convolution window; is the position feature matrix for m offsets; is the feature of the j-th column of the input matrix at the m-offset position.

[0100] The above settings can match the time correlation of photovoltaic output scenarios and reduce the complexity of the training process.

[0101] After the above data preprocessing, the randomly sampled noise is used as the input of the graph convolutional generative adversarial network generator, and the historical output data of each photovoltaic site is used as the input of the discriminator. Then, the graph filter of the spatial dimension and the one-dimensional filter of the time dimension are calculated using equations (7) and (8) respectively. Finally, the graph convolutional generative adversarial network is trained with the training data until the Nash equilibrium is reached. When the generative adversarial network training is completed, the photovoltaic output scene generated by the generator that is similar to the real data is obtained.

[0102] 2. Photovoltaic output reduction scenarios based on spatiotemporal clustering.

[0103] (III) Select typical photovoltaic output sites using ST-DBSCAN spatial clustering;

[0104] This embodiment uses the ST-DBSCAN clustering algorithm to cluster the photovoltaic power station power data, distinguish the densely concentrated data from the scattered data, and select typical photovoltaic sites. The following steps are included:

[0105] Step 1, set the input field radius Eps, the minimum number minpts and the time neighborhood Δt;

[0106] Step 2: Starting from any unlabeled point, select a point as the core point p for scanning, scan the neighborhood within the time neighborhood Δt of the core point p, and scan the number of neighborhood points that meet the distance threshold in the spherical area Eps(s) with the core point p as the center and the radius Eps from the neighborhood within the time neighborhood Δt;

[0107] Step 3: If the number of neighboring points of the core point p is less than the set threshold MinPts, the point is marked as a noise point;

[0108] Step 4: If the number of neighborhood points is greater than the set threshold MinPts, the point is marked as a core point p, a cluster number is generated as cluster C1, and all points of the core point p in the neighborhood Eps(s) are added to the cluster C1 to be scanned;

[0109] Step 5: For the unmarked points in cluster C1, repeat steps S42 to S44 until all the points in cluster C1 are marked;

[0110] Step 6: Continue to select labeled points from the data set, repeat steps 2 to 5, and cluster them into another class until all points are labeled.

[0111] (iv) Data reconstruction;

[0112] like Figure 3 As shown in the figure, the typical PV site outputs obtained through spatial clustering are spliced ​​into a one-dimensional sequence according to the corresponding dates.

[0113] The output of typical station i obtained by spatial clustering is spliced ​​into a matrix according to the time dimension. Then the 24-hour output data of all stations on the dth day are spliced ​​into a vector:

[0114] V d =[P1(1,d),P1(2,d),...,P1(t,d),...,P i (1,D),P i (2,D),...,P i (t,D)] T (9);

[0115]

[0116] Where P i (t,D) is the output of the i-th typical station at the t-th hour on the D-th day; t=1,2,...,24, d=1,2,...,D, D is the total number of days, D=365; i=1,2,...,M, where M is the number of typical stations; It is a 24×M matrix.

[0117] (V) Select typical output dates using time clustering;

[0118] The K-medoids clustering method is used to select typical output dates. The specific steps are as follows:

[0119] Step 1: Randomly select r sequences from all sequences as the initial clustering centers, using express.

[0120] Step 2: According to the principle of being closest to the cluster center, the remaining objects are assigned to each class, and the distance between the two closest sequences is calculated;

[0121]

[0122] In the formula, d(u i ,u j ) is the distance between two sequences, for each sequence u i , select the cluster center u with the smallest distance jTo calculate the distance; T is the number of time points in the sequence; For scene u i The value at the tth time point; the distance d(u i ,u j ) is the sum of the absolute values ​​of the differences between the two series at each time point.

[0123] Step 3: The goal is to minimize the sum of weighted distances from all scenes to the cluster center. According to the principle of minimizing the total loss of formula (11), find a new cluster center to replace the original cluster center.

[0124]

[0125] In the formula, S is the set of all sequences; J is the set of cluster centers, the purpose is to find an optimal subset J of S to replace S, so that J contains as much information as possible from S; p i For the sequence u i The probability of occurrence.

[0126] Step 4: Determine whether it converges. If not, repeat step 2. If it converges, the cluster centers obtained by clustering are It is a typical day for photovoltaic output;

[0127] Step 5: Calculate the probability p of each sequence appearing after clustering i That is, p i is the ratio of the number of scenes in the first category to the total number of scenes.

[0128] The present invention considers that the complete output of each site corresponding to a typical date is a typical scenario of photovoltaic output. After obtaining the typical output date, the generated output data is returned, and the complete output slice corresponding to the output date is selected as the typical scenario of photovoltaic output.

[0129] Embodiment 2

[0130] like Figure 5 As shown, the present invention also provides a distributed photovoltaic random output scene generation system, comprising:

[0131] Data acquisition module: used to obtain the historical output data of each photovoltaic site and normalize and standardize the historical output data.

[0132] Graph convolutional generative adversarial network processing module: Use noise as the input of the generator, use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the graph convolutional generative adversarial network processing module, and train the graph convolutional generative adversarial network.

[0133] The noise is input into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios and complete data enhancement of the output scenarios.

[0134] Spatial clustering processing module: Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish densely concentrated data from scattered data, select typical photovoltaic sites, and splice the output of each typical photovoltaic site according to the time dimension.

[0135] Time clustering processing module: Use the K-medoids clustering method to select typical output dates, select output scenarios from the completely generated photovoltaic output scenarios according to the dates, and obtain typical photovoltaic output scenarios.

[0136] The distributed photovoltaic random output scene generation system of the present invention can be installed in a computer device. The computer device includes a processor, a memory, and a computer program stored in the memory and run on the processor, such as a distributed photovoltaic random output scene generation program. Among them, the memory includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example: SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. The processor is the control core of the electronic device, and uses various interfaces and lines to connect the various components of the entire computer device, and executes various functions of the computer device and processes data by running or executing programs or modules stored in the memory, and calling data stored in the memory.

[0137] The module described in the present invention refers to a series of computer program segments that can be executed by a processor of a computer device and can complete fixed functions, and is stored in a memory of the computer device.

[0138] The present invention uses graph convolution to generate a network to track and simulate the photovoltaic output characteristics, focusing on the internal laws of distributed photovoltaic output, exploring the spatiotemporal correlation of the outputs of multiple distributed photovoltaic units, and generating a joint output scenario of multiple units. On the basis of taking into account the randomness and correlation of the output of distributed photovoltaic units, artificial data with the same dimension and similar probability distribution as the original photovoltaic output data is generated, making the distributed photovoltaic output scenario more comprehensive and diverse, while improving the accuracy of the distributed photovoltaic output scenario. Using a scene reduction method that clusters from two dimensions, space and time, the spatial correlation and temporal correlation between the outputs of each photovoltaic site are retained, typical photovoltaic sites and typical photovoltaic output days are extracted, and a typical output scenario with multiple photovoltaic sites is obtained.

Claims

1. A method for generating a distributed photovoltaic random output scenario, characterized in that: The following steps are involved: S1. Collect historical output data of each photovoltaic site and normalize and standardize the historical output data; S2. Use the noise as the input of the generator in the graph convolutional generative adversarial network, use the generator to generate output data, use the output data generated by the generator and the processed historical output data together as the input of the discriminator in the graph convolutional generative adversarial network, and train the graph convolutional generative adversarial network until it reaches Nash equilibrium; S3, input the noise into the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios and complete data enhancement of the output scenarios; S4. Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish the densely concentrated data from the scattered data, select typical photovoltaic sites, and splice the output of each typical photovoltaic site according to the time dimension; S5. Use the K-medoids clustering method to select typical output dates, select output scenarios from the completely generated photovoltaic output scenarios according to the dates, and obtain typical photovoltaic output scenarios.

2. A distributed photovoltaic random output scenario generation method according to claim 1, characterized in that: In S1, the historical output data of photovoltaic sites are divided into rainy season and dry season, and the data format is set as a three-dimensional matrix. The photovoltaic power information of all sites on each day is recorded in the form of a two-dimensional matrix {x n,t }.

3. A distributed photovoltaic random output scenario generation method according to claim 1, characterized in that: The layers in the generator and discriminator of the graph convolutional generative adversarial network both use graph convolutional networks. Both the generator and the discriminator contain graph filters in the spatial dimension and feature filters in the temporal dimension. Graph filters are used in the spatial dimension to capture the spatial correlation between photovoltaic sites. In the temporal dimension, one-dimensional convolutional filters are used to process temporal features to complete feature extraction.

4. A distributed photovoltaic random output scenario generation method according to claim 3, characterized in that: The following steps are involved: Use the input data matrix as parameters for feature filters: X (l+1) =σ(AX (l) W (l) ),l=1,...,L (6); The last layer uses the sigmoid function to generate output values ​​in the range of (0,1) as photovoltaic output data; Through the feature aggregation of adjacent nodes of the graph convolution, the correlation coefficient after exponential transformation is used as the weight of the graph filter. The calculation expression is: The matrix multiplication term in formula (6) Considered as the output of the feature filter, the length is (2M l +1) and a one-dimensional convolution filter with trainable weight coefficients, the elements of the j-th column are: In the formula, Each layer has K l The input data matrix of features is , where l∈{1,...,L}, L is the total number of hidden layers in the generator; σ(·) is the nonlinear activation function; A is the graph filter, is an N×N matrix; is the weight matrix, K l ×K l+1 The matrix of i,j is the correlation coefficient between node i and node j; m is the position offset of the sliding window in the one-dimensional convolution; M l is half the width of the convolution window; is the weight of the convolution kernel; is the feature of the j-th column of the input matrix at the m-offset position.

5. A distributed photovoltaic random output scenario generation method according to claim 3, characterized in that: The loss function of the generator is: The loss function of the discriminator is: Where z is the noise variable; P z is a simple distribution; G(z) is the noise variable mapped to a generated data space; P r is the probability distribution of real samples; x is the generated sample; G is the generator; D is the discriminator; P G is the probability distribution of generated samples; E is the expectation.

6. A distributed photovoltaic random output scenario generation method according to claim 1, characterized in that: In S4, selecting a typical photovoltaic site includes the following steps: S41, setting the input field radius Eps, the minimum number minpts and the time neighborhood Δt; S42, starting from any unlabeled point, a certain point is selected as the core point p of the scan, and the neighborhood is scanned within the time neighborhood Δt of the core point p, and the number of neighborhood points that meet the distance threshold in the spherical area Eps(s) with the core point p as the center and the radius Eps from the neighborhood is scanned within the time neighborhood Δt; S43, if the number of neighboring points of the core point p is less than the set threshold MinPts, mark the point as a noise point; S44, if the number of neighborhood points is greater than the set threshold MinPts, the point is marked as a core point p, a cluster number is generated as cluster C1, and all points of the core point p in the neighborhood Eps(s) are added to the cluster C1 to be scanned; S45, for the unmarked points in cluster C1, repeat steps S42 to S44 until all the points in cluster C1 are marked; S46. Continue to select marked points from the data set, repeat steps S42 to S45, and cluster into another class until all points are marked.

7. A distributed photovoltaic random output scenario generation method according to claim 6, characterized in that: The specific method of splicing the output of each typical photovoltaic site according to the time dimension is: splicing the typical photovoltaic sites according to the corresponding dates to form a one-dimensional sequence: V d =[P1(1,d),P1(2,d),...,P1(t,d),...,P i (1,D),P i (2,D),...,P i (t,D)] T (9); Where P i (t,D) is the output of the i-th typical station at the t-th hour on the D-th day; t=1,2,...,24, d=1,2,...,D, D is the total number of days, D=365; i=1,2,...,M, where M is the number of typical stations; It is a 24×M matrix.

8. The method for generating a distributed photovoltaic random output scenario according to claim 1, characterized in that: S5 includes the following steps: S51, randomly select r sequences from all sequences as initial clustering centers; S52, according to the principle of being closest to the cluster center, assign the remaining objects to each class, and calculate the distance between the two sequences that are closest to each other; S53, minimizing the sum of weighted distances from all scenes to the cluster center; S54, judging whether it converges, if not, re-performing S52, if it converges, then the r cluster centers obtained by clustering are the typical days of photovoltaic output; S55. Calculate the probability of occurrence of each sequence after clustering.

9. A distributed photovoltaic random output scenario generation method according to claim 8, characterized in that: The calculation expression for calculating the distance between the two sequences with the closest distance is: In the formula, d(u i ,u j ) is the distance between two sequences, for each sequence u i , select the cluster center u with the smallest distance j To calculate the distance; T is the number of time points in the sequence; For scene u i The value at the tth time point; the distance d(u i ,u j ) is the sum of the absolute values ​​of the differences between the two series at each time point.

10. A distributed photovoltaic random output scene generation system, using the distributed photovoltaic random output scene generation method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module: used to obtain the historical output data of each photovoltaic site and normalize and standardize the historical output data; Graph convolutional generative adversarial network processing module: Use the noise as the input of the generator, and use the data generated by the generator and the historical output data processed by the data acquisition module as the input of the discriminator in the graph convolutional generative adversarial network processing module to train the graph convolutional generative adversarial network; The noise is input into the generator of the trained graph convolutional generative adversarial network to generate photovoltaic output scenarios and complete data enhancement of the output scenarios; Spatial clustering processing module: Use the ST-DBSCAN clustering algorithm to cluster the power data of photovoltaic power stations, distinguish densely concentrated data from scattered data, select typical photovoltaic sites, and splice the output of each typical photovoltaic site according to the time dimension; Time clustering processing module: Use the K-medoids clustering method to select typical output dates, select output scenarios from the completely generated photovoltaic output scenarios according to the dates, and obtain typical photovoltaic output scenarios.

Citation Information

Patent Citations

  • Photovoltaic power station typical scene generation method based on multi-scene model

    CN112541546A

  • Wind-solar active power output scene generation method and device, electronic equipment and storage medium

    CN114066236A

  • Photovoltaic monthly scene generation method and system

    CN115841280A

  • Power distribution network typical operation scene generation method based on Gaussian mixture model

    CN116226689A

  • Photovoltaic output typical scene acquisition method and device, equipment and storage medium

    CN116933100A

Cited By

  • Evaluation method and device for voltage sag of power distribution network, computer equipment and program product

    CN120784881A

  • Photovoltaic power generation prediction method and system based on uncertainty graph convolution

    CN121052461A