A Clustering Method for Typical Photovoltaic Output Scenarios Considering Comprehensive Similarity Measure
Through the comprehensive similarity measurement and two-stage center of mass optimization method, combined with the comprehensive evaluation of the entropy weight Topsis method, the problem of curve similarity measurement, center of mass extraction and evaluation indicators in photovoltaic output scenario clustering is solved, and a more accurate and comprehensive clustering of typical photovoltaic output scenarios is achieved.
Patent Information
- Application Number
- CN202310727085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The existing photovoltaic output scene clustering methods have problems in curve similarity measurement, typical center of mass extraction and scene set quality evaluation, such as insufficient consideration of fluctuation time shift differences, susceptibility to outliers and noise, and insufficient comprehensive evaluation indicators.
A method of clustering of typical photovoltaic output scenes that consider comprehensive similarity metrics is proposed. The distance of the photovoltaic curve is comprehensively measured by the distance of the photovoltaic curve, the morphological trend and the fluctuation position. The two-stage center of mass optimization method is used to extract typical scenes, and the scene set results are comprehensively evaluated by the entropy weight Topsis method.
Effectively identify different photovoltaic output curve types, taking into account the morphological representativeness and numerical magnitude representativeness of the center of mass curve, providing a more comprehensive scenario set evaluation, and solving the result evaluation error caused by different algorithm curve distance measurement methods.
Smart Images

Figure CN116992319B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photovoltaic output scenario reduction, and in particular relates to a photovoltaic typical output scenario clustering method considering comprehensive similarity measurement. Background Art
[0002] By the end of 2022, my country's cumulative installed photovoltaic capacity will reach approximately 392 million kilowatts, ranking first in the world. As the scale of photovoltaic power generation increases rapidly year by year, photovoltaic power generation data presents the characteristics of massive, high-dimensional and highly random. In order to study the optimization problems of power system planning, operation, and scheduling under the background of penetration of renewable energy such as photovoltaic power generation, it is necessary to effectively reduce the photovoltaic output scene and reflect the most realistic photovoltaic power generation characteristics with the least amount of data to achieve the purpose of reducing the amount of calculation. Therefore, how to accurately and efficiently construct a typical output scene from large-scale data is a theoretical and practical problem that needs to be solved urgently.
[0003] Clustering, as an unsupervised machine learning method, is widely used in aspects such as power system load classification and extraction of characteristic scenarios of new energy sources such as wind and solar. Traditional clustering algorithms divide scenarios into different categories through appropriate similarity measurement methods and then extract typical scenarios in the same category. However, there are still several problems in application: ① The problem of curve similarity measurement. Similarity measurement is the basis of unsupervised clustering of time series and other time series analyses. Currently, mainly single metric distance or a combination of multiple metric distances is used to judge the similarity of curves, such as the classical Euclidean distance, Wasserstein probability distance, dynamic time warping distance, correlation coefficient distance, Euclidean distance + morphological distance, etc. Although these methods can measure the similarity between curves from the overall numerical value or overall morphology, they rarely consider the fluctuation time shift difference between photovoltaic power generation curves caused by weather state changes. ② The problem of extracting typical centroids. As the representative curve in the clustering cluster, the centroid needs to reflect the overall characteristics of the cluster and the differences between clusters. Among them, methods such as the arithmetic average method, median method, fuzzy centroid method, and probability density method are widely used. However, these methods are easily affected by outliers and noise, which may reduce the accuracy of the clustering results. Some scholars solve the centroid extraction as an optimization problem, which improves the quality of the centroid to a certain extent, but most of them optimize with a single target and rarely consider the multi-dimensional representativeness of the centroid, such as numerical value, morphology, etc. ③ The problem of quality evaluation of the typical scenario set. Most existing studies use clustering validity indicators such as the SSE index, DBI index, and CHI index to evaluate the quality of the clustering results from the inter-class dispersion and intra-class compactness. Most of these indicators measure the similarity between samples with the Euclidean distance. Some scholars also improve the validity indicators using different distance functions to evaluate the clustering results. Although it can effectively solve the result evaluation error caused by different measurement methods, it may still not be able to evaluate different algorithms uniformly, and currently, few people establish different comprehensive index systems according to the characteristics of the research object to evaluate the scenario sets generated by different algorithms at the same time. Summary of the Invention
[0004] The purpose of the present invention is to provide a clustering method for typical photovoltaic output scenarios considering comprehensive similarity measurement on the basis of existing clustering methods, so as to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A clustering method for typical photovoltaic output scenarios considering comprehensive similarity measurement, comprising the following steps:
[0006] Step 1, Comprehensive similarity measurement of photovoltaic curves:
[0007] Considering the characteristics of photovoltaic power generation, comprehensively measure the distance of the curve with the similarity of the electricity quantity size, morphological trend similarity, and fluctuation position similarity of the photovoltaic output curve;
[0008] Step 2, construct a photovoltaic scenario aggregation type reduction model:
[0009] Select the initial center points using the dissimilarity matrix, use the comprehensive metric distance as the basis for sample division, extract typical scenarios using the two-stage centroid optimization method, determine the optimal number of clusters using the effectiveness evaluation index, and improve the traditional K-means algorithm respectively;
[0010] Step 3, verify the typical scenario set:
[0011] Select six indicators including the average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility, and electricity quantity, and comprehensively evaluate the results of the typical scenario set through the entropy weight Topsis method.
[0012] The detailed steps of the comprehensive similarity metric of the photovoltaic curve in Step 1 are as follows:
[0013] Assume that two photovoltaic output curves are x = (x1, x2,..., x m ) and y = (y1, y2,..., y m ), where m represents the total number of curve sampling points; x1, x2,..., x m represent the output of curve x at times 1, 2,..., m, with the unit: MW; y1, y2,..., y m represent the output of curve y at times 1, 2,..., m, with the unit: MW;
[0014] Step 1.1, use the difference in the area enclosed by the output curves x and y and the coordinate axes as the similarity of the electricity quantity between the two, and the distance formula for the similarity of the electricity quantity is expressed as follows:
[0015]
[0016] In the formula, ΔE(x, y) represents the distance of the electricity quantity similarity between curves x and y; t represents the t-th sampling point; m represents the total number of curve sampling points; x t and y t represent the output of curves x and y at the t-th sampling point respectively, with the unit: MW; represents the absolute value of the difference in electricity quantity between curves x and y; represents the larger value of the electricity quantities of curves x and y, and the smaller the value of ΔE(x, y), the more similar the electricity quantities of the two;
[0017] Step 1.2, to ensure the invariance of curve scaling, select the Z-Score standardization method to standardize the two curves to reduce the impact of the change in output amplitude on the correlation analysis;
[0018] Step 1.3, when calculating the cross - correlation of two curves, fix one curve and slide y along the horizontal axis from the initial position with a step size of 1 until it aligns with x. Replace the data at the original position of curve y with 0 at the positions where curve y does not align with curve x. Thus, the expression of the cross - correlation sequence function of curves x and y can be obtained as follows:
[0019]
[0020]
[0021] In the formula, represents the cross - correlation sequence of curves x and y, where, represents taking the cross - correlation value of curves x and y when represents the length of the sequence when the moving step size is s, s represents the moving step size of curve y along the horizontal axis, s ∈ [-m, m]; m represents the total number of sampling points of the curve; R s (x, y) represents the sum of the dot products of the two curves each time curve y moves along the horizontal axis; h represents the subscript of the position where curve y aligns with x. The position corresponding to the maximum cross - correlation value can be obtained through the cross - correlation sequence Then, according to the displacement of curve y at this time and the optimal displacement curve of curve y can be obtained;
[0022] Step 1.4, to ensure the invariance of curve displacement, the coefficient normalization method is used to process the cross - correlation sequence so that the zero - lag autocorrelation coefficient is equal to 1. The coefficient normalization formula is as follows:
[0023]
[0024] In the formula, NCC c (x, y) represents the cross - correlation sequence of curves x and y after coefficient normalization. The data in the NCC c (x, y) sequence ranges from [-1, 1]; represents the cross - correlation sequence of curves x and y; represents the length of the sequence when the moving step size is s; s represents the moving step size of curve y along the horizontal axis, s ∈ [-m, m]; R0(x, x) represents the autocorrelation value of curve x; R0(y, y) represents the autocorrelation value of curve y; when the data in the NCC c (x, y) sequence is 1 or -1, it means that the two curves are positively and negatively correlated in shape at this time. For the convenience of calculation, the shape trend similarity distance between curves x and y is defined as follows:
[0025] SBD(x,y) = 1 - max(NCC c (x,y)) (5)
[0026] Wherein, SBD(x,y) represents the morphological trend similarity distance between curves x and y, and its range is [0,2]; max(NCC c (x,y)) represents the maximum value in the cross-correlation sequence after normalizing the coefficients of curves x and y. The smaller the value of SBD(x,y), the more similar the morphological trends of curves x and y are;
[0027] Step 1.5: Take the deviation of the occurrence positions of the fluctuations of the two curves as the fluctuation position similarity of the two curves. The formula for the fluctuation position similarity distance is expressed as follows:
[0028]
[0029] Wherein, δ(x,y) represents the fluctuation position similarity distance between curves x and y; represents the absolute value of the displacement when curve y slides to the most similar morphological trend to x; m represents the total number of curve sampling points. The smaller the value of δ(x,y), the more similar the fluctuation positions of curves x and y are;
[0030] Step 1.6: To reflect the comprehensive similarity between power generation curves, considering the morphological trend similarity and the fluctuation position similarity on the basis of the similarity of power amounts, a comprehensive similarity metric distance applicable to photovoltaic power generation curves is obtained. The formula is expressed as follows:
[0031] ESD(x,y) = ΔE(x,y)·SBD(x,y)·δ(x,y) (7)
[0032] Wherein, ESD(x,y) represents the comprehensive similarity distance between curves x and y; ΔE(x,y) represents the similarity distance of the power amounts of curves x and y; SBD(x,y) represents the morphological trend similarity distance between curves x and y; δ(x,y) represents the fluctuation position similarity distance between curves x and y. The smaller the comprehensive similarity metric distance ESD(x,y), the more similar the two curves are.
[0033] The detailed steps for constructing the photovoltaic scenario clustering reduction model in the said Step 2 are as follows:
[0034] Step 2.1: Data preprocessing: Use the Lagrange interpolation method to complete the filling of missing photovoltaic output data, remove data with large deviations, and complete the cleaning of the data;
[0035] Step 2.2: Initialize the clustering parameters: Set the maximum number of clusters, the number of iterations, and the initial number of clusters;
[0036] Step 2.3, Determine the initial clustering centers: Based on the constructed dissimilarity matrix, two measurement criteria, namely mean dissimilarity and overall dissimilarity, are defined. The initial clustering centers are determined according to these criteria. The dissimilarity matrix is an object structure that stores the dissimilarities formed between several curves. The dissimilarity between two curves is defined by the comprehensive similarity distance. Suppose several photovoltaic output curves are represented by the matrix X = [x 1 , x 2 ,..., x i ,..., x n T , where x 1 , x 2 ,..., x i ,..., x n represent the 1st, 2nd,..., i-th,..., n-th curves, represents the output of the i-th curve at the 1st, 2nd,..., m-th moments, with the unit of MW. i represents the curve number, n represents the total number of curves, and m represents the total number of curve sampling points. Thus, the dissimilarity matrix between curves is as follows:
[0037]
[0038] In the formula, disM represents the dissimilarity matrix of n curves; n represents the total number of curves; dis(x i , x j ) represents the dissimilarity between curves x i and x j ; x i represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers;
[0039] The mean dissimilarity is defined as follows:
[0040]
[0041] In the formula, Adis(x i ) represents the mean dissimilarity of curve x i ; dis(x i , x j ) represents the dissimilarity between curves x i and x j ; x i represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers; n represents the total number of curves; Adis(x i ) reflects the position of curve x i , and the larger its value, the more it indicates x i The sparser the distribution of the surrounding curves and the higher the degree of discreteness from other curves, and vice versa, the denser.
[0042] The overall dissimilarity is defined as follows:
[0043]
[0044] In the formula, Tdis represents the overall dissimilarity of n curves; dis(x i , x j ) represents the dissimilarity between curve x i and x j ; n represents the total number of curves; x i represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers; Tdis is related to the distribution of all curves and reflects the sparsity of the curve distribution;
[0045] Step 2.4, similarity measure division: First, calculate the three metric distances from each curve in the sub-interval to the initial clustering center according to Step 1.1, Step 1.4, and Step 1.5 respectively. Then, combine the three metric methods and calculate the comprehensive similarity distance according to Step 1.6. Then, divide the curves into the category that is most similar to them according to the comprehensive similarity between each curve in the interval and the clustering centroid;
[0046] Step 2.5, update the clustering center: First, because cross-correlation is used to calculate the similarity between curve shapes, the optimization goal in the first stage is set to maximize the sum of the squares of the comprehensive similarity distances between the shape centroid and other curves in the cluster. The formula for the shape centroid is expressed as follows:
[0047]
[0048] In the formula, represents the shape centroid of the k-th class; represents the original curve set where the centroid to be solved is located; represents the 1st, 2nd,..., i-th,..., n k th curves; n k represents the number of curves in the k-th class; μ k represents the actual centroid of the k-th class; k represents the class label; represents the sum of the squares of the cross-correlations of all curves in the cluster;
[0049] Secondly, adopt the idea of multiplying by the same ratio to magnify the shape centroid as the second stage of centroid solution. The shape centroid is obtained by optimizing after standardizing the curve set. Set the multiplying ratio base as the ratio of the mean of the original sample curves in the cluster and the mean of the shape centroid. The formula for the actual centroid is expressed as follows:
[0050]
[0051] In the formula, μ k represents the actual centroid of the k-th class; n k represents the number of curves in the k-th class; represents the output of the i-th curve in the k-th class at the t-th moment, with the unit of MW; k represents the class label; i represents the curve number; t represents the t-th sampling point; m represents the total number of curve sampling points; represents the value of the morphological centroid of the k-th class at the t-th moment; represents the morphological centroid of the k-th class;
[0052] Step 2.6, clustering validity evaluation: Based on the principle of minimizing the sum of the comprehensive similarity distances from the samples within the class to the centroid, a clustering validity index W is proposed to determine the optimal number of clusters. The formula for W is expressed as follows:
[0053]
[0054] In the formula, W represents the sum of the comprehensive similarity distances from the centroid curve to the curves within the class; K represents the number of clusters; n k represents the number of curves in the k-th class; k represents the class label; i represents the curve number; represents the i-th curve in the k-th class; represents the curve and μ k The comprehensive similarity distance; W will gradually decrease as the number of clusters increases. The inflection point of the W curve indicates that after this point, when the number of clusters is further increased, the decrease in the sum of the comprehensive similarity distances from the centroid to the curves within the class is very small. That is, the K value corresponding to the inflection point of the W curve is the optimal number of clusters;
[0055] After the clustering for the current K value is completed, calculate the clustering validity index W according to Step 2.6. Then, let K = K + 1, and return to Step 2.3. When K reaches the maximum, end the loop, and take the K value corresponding to the inflection point of the W curve as the optimal number of clusters.
[0056] The detailed steps for verifying the typical scenario set in Step 3 are as follows:
[0057] Step 3.1, time interval division: The twenty-four solar terms, as a very important time division method in traditional Chinese culture, reflect the position of the sun on the earth and the change of light intensity. The power generation of the photovoltaic system is directly related to the solar light intensity. At the same time, taking the twenty-four solar terms as the interval can also consider the influence of seasonal and climatic factors. For example, the temperature gradually rises in spring, reaches a peak in summer, gradually drops in autumn, and reaches a low in winter. These factors will also affect the output of the photovoltaic power generation system. Therefore, using the twenty-four solar terms as the clustering interval can better reflect the output characteristics and change laws of the photovoltaic system in different time periods:
[0058] Step 3.2, Generation of Scenario Set Replacement: According to the principle that the clustering labels correspond to the original date labels, several typical output scenarios that match the optimal number of clusters within the solar term interval are used to replace the typical output curves obtained by clustering with the original position output curves, so as to obtain new typical solar term output scenarios;
[0059] Step 3.3, Selection of Scenario Set Evaluation Indexes: Based on the characteristics of photovoltaic power generation, an index system including curve fluctuation characteristics and electricity quantity characteristics is established to verify the quality of the typical scenario set. Specifically, it includes 6 indexes: average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility rate, and electricity quantity. The formulas for the average number of fluctuations and extreme volatility rate are as follows:
[0060] Average number of fluctuations, this index reflects the average number of output reversals during the research period:
[0061]
[0062] In the formula, K fluc represents the average number of fluctuations, and the unit is: times; G represents the total number of days in the research period; g represents the gth day in the research period; t represents the tth sampling point; m represents the total number of curve sampling points; P t-1 , P t , P t+1 represent the outputs at the (t - 1)th, tth, and (t + 1)th moments respectively, and the unit is: MW. Among them, if (P t - P t-1 )(P t - P t+1 ) > 0 indicates 1 output reversal, and if (P t - P t-1 )(P t - P t+1 ) ≤ 0 indicates that there is no output reversal;
[0063] Extreme volatility rate, this index reflects the proportion of the number of days when the photovoltaic power generation exceeds or is lower than the specified extreme fluctuation threshold to the total number of days in the research period:
[0064]
[0065] In the formula, PFR represents the extreme volatility rate, and the unit is: %; ε% represents the extreme fluctuation threshold; G represents the total number of days in the research period; G + (ε%) represents the number of days with a fluctuation amplitude greater than the maximum threshold; G - (ε%) represents the number of days with a fluctuation amplitude less than the minimum threshold;
[0066] Step 3.4, Index Comprehensive Evaluation: The entropy weight TOPSIS method is used to comprehensively analyze the six indexes of the average fluctuation times, average fluctuation amplitude, output distribution skewness, output distribution kurtosis, extreme volatility, and electricity quantity. First, the entropy weight method is used to calculate the weights of each index in the scenario set, and then the TOPSIS analysis method is used for comprehensive scoring. Since the quality of the typical scenario set needs to be judged based on the indexes of the original scenario set, the six index types of the average fluctuation times, average fluctuation amplitude, output distribution skewness, output distribution kurtosis, extreme volatility, and electricity quantity are all set as intermediate type indexes, and the intermediate optimal value is the index value of each item in the original scenario set. The closer to the intermediate optimal value, the better the index effect.
[0067] The present invention has the following beneficial effects:
[0068] 1. The comprehensive similarity measurement method proposed by the present invention can effectively identify different types of photovoltaic output curves from three aspects: the magnitude of electricity quantity, the morphological trend, and the fluctuation position.
[0069] 2. The two-stage centroid optimization extraction method proposed by the present invention can take into account both the morphological representativeness and the numerical magnitude representativeness of the centroid curve.
[0070] 3. The photovoltaic typical scenario set index evaluation system proposed by the present invention can effectively reflect the photovoltaic output characteristics, comprehensively evaluate the scenario set, and solve the result evaluation error caused by different algorithm curve distance measurement methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The present invention will be further described below in conjunction with the drawings and embodiments.
[0072] Figure 1 is a flowchart of the photovoltaic output scenario clustering method of the present invention;
[0073] Figure 2 is a schematic diagram of the similarity of output area (electricity quantity);
[0074] Figure 3 is a schematic diagram of the sliding process;
[0075] Figure 4 is a schematic diagram of the similarity of fluctuation position;
[0076] Figure 5 is a quality evaluation diagram of typical scenario sets of different algorithms;
[0077] Figure 6 is a quality evaluation diagram of typical scenario sets in different intervals. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0079] Embodiment 1:
[0080] As Figure 1 shown, a photovoltaic typical output scenario clustering method considering comprehensive similarity measurement proposed by the present invention includes the following steps:
[0081] Step 1, comprehensive similarity measurement of photovoltaic curves:
[0082] Considering the characteristics of photovoltaic power generation, the distance of the curves is comprehensively measured by the similarity of the electricity quantity size, the similarity of the morphological trend, and the similarity of the fluctuation position of the photovoltaic output curves.
[0083] Step 2, constructing a photovoltaic scenario aggregation and reduction model:
[0084] Taking the dissimilarity matrix to select the initial center points, using the comprehensive measurement distance as the basis for sample division, extracting typical scenarios by a two-stage centroid optimization method, and determining the optimal number of clusters by an effectiveness evaluation index, respectively improving the traditional K-means algorithm.
[0085] Step 3, verification of the typical scenario set:
[0086] Select 6 indexes including the average number of fluctuations, the average fluctuation amplitude, the skewness of the output distribution, the kurtosis of the output distribution, and the extreme volatility electricity quantity to comprehensively evaluate the results of the typical scenario set by the entropy weight Topsis method.
[0087] Furthermore, the detailed steps of the comprehensive similarity measurement of the photovoltaic curves in Step 1 are as follows:
[0088] Assume that two photovoltaic output curves are x = (x1, x2,..., x m ) and y = (y1, y2,..., y m ), where m represents the total number of curve sampling points; x1, x2,..., x m represent the output of curve x at times 1, 2,..., m, with the unit of: MW; y1, y2,..., y m represent the output of curve y at times 1, 2,..., m, with the unit of: MW;
[0089] Step 1.1, taking the difference in the area enclosed by the output curves x and y and the coordinate axes as the similarity of the electricity quantity size between the two, as Figure 2As shown, the area enclosed by curves x and y represents the difference in electricity quantity. The formula for the similarity distance of electricity quantity magnitudes is expressed as follows:
[0090]
[0091] In the formula, ΔE(x,y) represents the similarity distance of electricity quantity magnitudes between curves x and y; t represents the t-th sampling point; m represents the total number of sampling points of the curve; x t and y t respectively represent the outputs of curves x and y at the t-th sampling point, with the unit: MW; represents the absolute value of the difference in electricity quantity between curves x and y; represents the larger value of the electricity quantities of curves x and y. The smaller the value of ΔE(x,y), the more similar the two electricity quantities are;
[0092] Step 1.2: To ensure the invariance of curve scaling, the Z-Score normalization method is selected to normalize the two curves to reduce the impact of the change in output amplitude on the correlation analysis;
[0093] Step 1.3: When calculating the cross-correlation of the two curves, keep one curve fixed. As Figure 3 shown, slide y along the horizontal axis from the initial position with a step size of 1 to the position aligned with x. Replace the data at the original position of curve y with 0 at the positions where curve y is not aligned with curve x. Thus, the cross-correlation sequence function of curves x and y can be obtained as follows:
[0094]
[0095]
[0096] In the formula, represents the cross-correlation sequence of curves x and y, where, represents taking the cross-correlation value of curves x and y when; represents the length of the sequence when the moving step size is s, s represents the moving step size of curve y along the horizontal axis, s ∈ [-m, m]; m represents the total number of sampling points of the curve; R s (x,y) represents the sum of the dot products of the two curves each time curve y moves along the horizontal axis; h represents the position subscript where curve y is aligned with x. The position corresponding to the maximum cross-correlation value can be obtained through the cross-correlation sequence And then according to the displacement of curve y and the optimal displacement curve of curve y at this time can be obtained;
[0097] Step 1.4. To ensure the invariance of curve displacement, a coefficient normalization method is used to process the cross-correlation sequence so that the autocorrelation coefficient at zero lag is equal to 1. The coefficient normalization formula is as follows:
[0098]
[0099] In the formula, NCC c (x,y) represents the cross-correlation sequence after coefficient normalization of curves x and y. The data value range in the NCC c (x,y) sequence is [-1, 1]; represents the cross-correlation sequence of curves x and y; represents the length of the sequence when the moving step is s; s represents the moving step of curve y along the horizontal axis, s ∈ [-m, m]; R0(x,x) represents the autocorrelation value of curve x; R0(y,y) represents the autocorrelation value of curve y. When the data in the NCC c (x,y) sequence is 1 or -1, it means that the two curve forms are positively and negatively correlated at this time. For the convenience of calculation, the morphological trend similarity distance between curves x and y is defined as follows:
[0100] SBD(x,y) = 1 - max(NCC c (x,y)) (5)
[0101] In the formula, SBD(x,y) represents the morphological trend similarity distance between curves x and y, and its range is [0, 2]; max(NCC c (x,y)) represents the maximum value in the cross-correlation sequence after coefficient normalization of curves x and y. The smaller the value of SBD(x,y), the more similar the morphological trends of curves x and y;
[0102] Step 1.5. Take the deviation of the positions where the fluctuations of the two curves appear as the fluctuation position similarity of the two curves. As Figure 4 shown, curves x and y have the same morphological trend, but there are certain differences in the fluctuation positions. Therefore, the fluctuation position similarity is defined as follows:
[0103]
[0104] In the formula, δ(x,y) represents the fluctuation position similarity distance between curves x and y; represents the absolute value of the displacement when curve y slides to the most similar morphological trend to x; m represents the total number of curve sampling points. The smaller the value of δ(x,y), the more similar the fluctuation positions of curves x and y;
[0105] Step 1.6, to reflect the comprehensive similarity between power generation curves, on the basis of the similarity of electricity quantity, the similarity of morphological trend and the similarity of fluctuation position are considered to obtain a comprehensive similarity metric distance applicable to photovoltaic power generation curves, and the formula is expressed as follows:
[0106] ESD(x,y)=ΔE(x,y)·SBD(x,y)·δ(x,y) (7)
[0107] In the formula, ESD(x,y) represents the comprehensive similarity distance between curves x and y; ΔE(x,y) represents the similarity distance of electricity quantity between curves x and y; SBD(x,y) represents the similarity distance of morphological trend between curves x and y; δ(x,y) represents the similarity distance of fluctuation position between curves x and y. The smaller the comprehensive similarity metric distance ESD(x,y), the more similar the two curves are;
[0108] Furthermore, the detailed steps for constructing the photovoltaic scenario clustering reduction model in step 2 are as follows:
[0109] Step 2.1, data preprocessing: Use the Lagrange interpolation method to complete the missing photovoltaic output data, remove the data with large deviations, and complete the data cleaning;
[0110] Step 2.2, initialize the clustering parameters: Set the maximum number of clusters, the number of iterations, and the initial number of clusters;
[0111] Step 2.3, determine the initial cluster centers: Based on the construction of the dissimilarity matrix, two metrics of mean dissimilarity and overall dissimilarity are defined to determine the initial cluster centers according to this criterion. The dissimilarity matrix is an object structure that stores the dissimilarities formed between several curves. The dissimilarity between two curves is defined by the comprehensive similarity distance. Assume that several photovoltaic output curves are represented by the matrix X = [x 1 ,x 2 ,...,x i ,...,x n T where, x 1 ,x 2 ,...,x i ,...,x n represent the 1st, 2nd,..., i,..., nth curves, represents the output of the i-th curve at the 1st, 2nd,..., mth moments, with the unit of: MW, i represents the curve number, n represents the total number of curves, and m represents the total number of curve sampling points. Thus, the dissimilarity matrix between curves is obtained as follows:
[0112]
[0113] Wherein, disM represents the dissimilarity matrix of n curves; n represents the total number of curves; dis(x i , x j ) represents the dissimilarity between curve x i and x j ; x i represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers;
[0114] The definition of the mean dissimilarity is as follows:
[0115]
[0116] Wherein, Adis(x i ) represents the mean dissimilarity of curve x i ; dis(x i , x j ) represents the dissimilarity between curve x i and x j ; xi represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers; n represents the total number of curves; Adis(x i ) reflects the position of curve x i , the larger its value, the sparser the distribution of the curves around x i and the higher the degree of dispersion from other curves, and vice versa, the denser;
[0117] The definition of the overall dissimilarity is as follows:
[0118]
[0119] Wherein, Tdis represents the overall dissimilarity of n curves; dis(x i , x j ) represents the dissimilarity between curve x i and x j ; n represents the total number of curves; x i represents the i-th curve; x j represents the j-th curve; i and j represent the curve numbers; Tdis is related to the distribution of all curves and reflects the sparsity degree of the curve distribution;
[0120] Step 2.4, similarity measurement division: First, calculate the three measurement distances from each curve in the sub-interval to the initial clustering center according to Step 1.1, Step 1.4, and Step 1.5 respectively, and then synthesize the three measurement methods to calculate the comprehensive similarity distance according to Step 1.6. Then, divide the curves into the category that is most similar to them (i.e., the comprehensive similarity distance is the smallest) according to the comprehensive similarity between each curve in the interval and the clustering centroid;
[0121] Step 2.5, Update the cluster center: First, since cross-correlation calculates the similarity between curve forms, the optimization objective in the first stage is set to maximize the sum of the squares of the comprehensive similarity distances between the form centroid and other curves within the cluster. The formula for the form centroid is expressed as follows:
[0122]
[0123] In the formula, represents the form centroid of the k-th class; represents the set of original curves where the centroid to be solved is located; represents the 1st, 2nd,..., i-th,..., n k curves; n k represents the number of curves in the k-th class; μ k represents the actual centroid of the k-th class; k represents the class label; represents the sum of the squares of the cross-correlations of all curves in the cluster;
[0124] Secondly, the idea of equal ratio amplification is adopted to amplify the form centroid as the second stage of centroid solution. The form centroid is obtained by optimizing after standardizing the curve set. The equal ratio amplification base is set as the ratio of the mean of the original cluster sample curves to the mean of the form centroid. The formula for the actual centroid is expressed as follows:
[0125]
[0126] In the formula, μ k represents the actual centroid of the k-th class; n k represents the number of curves in the k-th class; represents the output of the i-th curve in the k-th class at time t, with the unit: MW; k represents the class label; i represents the curve number; t represents the t-th sampling point; m represents the total number of curve sampling points; represents the value of the form centroid of the k-th class at time t; represents the form centroid of the k-th class;
[0127] Step 2.6, Cluster validity evaluation: Based on the principle of minimizing the sum of the comprehensive similarity distances from the samples within the class to the centroid, a cluster validity index W is proposed to determine the optimal number of clusters. The formula for W is expressed as follows:
[0128]
[0129] In the formula, W represents the sum of the comprehensive similarity distances from the centroid curve to the curves within the class; K represents the number of clusters; n k represents the number of curves in the k-th class; k represents the class label; i represents the curve number; represents the i-th curve in the k-th class; represents the curve and μ kThe comprehensive similarity distance. W will gradually decrease as the number of clusters increases. The inflection point of the W curve indicates that after this point, when the number of clusters is further increased, the decrease in the sum of the centroid to the comprehensive similarity distance of the intra-class curves is very small. That is, the K value corresponding to the inflection point of the W curve is the optimal number of clusters;
[0130] After the clustering with the current K value is completed, calculate the clustering validity index W according to step 2.6 of the formula. Then, let K = K + 1, and return to step 2.3. When K reaches the maximum, end the loop, and take the K value corresponding to the inflection point of the W curve as the optimal number of clusters.
[0131] Furthermore, the detailed steps for verifying the typical scenario set in step 3 are as follows:
[0132] Step 3.1, time interval division: The twenty-four solar terms, as a very important time division method in traditional Chinese culture, reflect the position of the sun on the earth and the change of light intensity. The power generation power of the photovoltaic system is directly related to the solar light intensity. At the same time, taking the twenty-four solar terms as intervals can also consider the influence of seasonal and climatic factors. For example, the temperature gradually rises in spring, reaches a peak in summer, gradually drops in autumn, and reaches a low in winter. These factors will also affect the output of the photovoltaic power generation system. Therefore, using the twenty-four solar terms as the clustering interval can better reflect the output characteristics and change laws of the photovoltaic system in different time periods:
[0133] Step 3.2, generation of scenario set replacement: According to the principle that the clustering label corresponds to the original date label, replace the obtained typical output curves of the clusters with the original position output curves for several typical output scenarios that match the optimal number of clusters within the solar term interval, so as to obtain a new typical solar term output scenario;
[0134] Step 3.3, selection of scenario set evaluation indicators: Based on the photovoltaic power generation characteristics, establish an index system including curve fluctuation characteristics and electricity quantity characteristics to verify the quality of the typical scenario set, specifically including 6 indicators: average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility rate, and electricity quantity. The indicators are introduced as follows:
[0135] 1) Average number of fluctuations. This indicator reflects the average number of output reversals during the research period. The formula is expressed as follows:
[0136]
[0137] In the formula, K fluc represents the average number of fluctuations, with the unit of times; G represents the total number of days in the research period; g represents the gth day in the research period; t represents the tth sampling point; m represents the total number of curve sampling points; P t-1 、P t 、P t+1They respectively represent the outputs at the (t-1)-th, t-th, and (t+1)-th moments, with the unit being MW. Among them, if (P t -P t-1 )(P t -P t+1 ) > 0 indicates that the output reverses once. If (P t -P t-1 )(P t -P t+1 ) ≤ 0 indicates that the output does not reverse;
[0138] 2) The average fluctuation amplitude refers to the average fluctuation of the PV output curves in adjacent time periods. This index can reflect the overall fluctuation amplitude of the PV power generation curve during the research period, with the unit being %.
[0139] 3) The skewness of the output distribution. Skewness is a measure of the direction and degree of skewness of the statistical data distribution and is a numerical characteristic of the asymmetry of the statistical data distribution. This index can reflect the overall severe fluctuation degree of the PV power generation curve. The more severe the curve fluctuation, the larger the skewness value.
[0140] 4) The kurtosis of the output distribution. Kurtosis is a characteristic number that characterizes the peak height of the probability density distribution curve at the mean value. This index can reflect the sharpness of the peak of the PV power distribution curve. The sharper the peak, the faster the curve change rate and the greater the fluctuation.
[0141] 5) The extreme volatility. This index reflects the proportion of the number of days when the PV power generation exceeds or is lower than the specified extreme fluctuation threshold to the total number of days in the research period. The formula is expressed as follows:
[0142]
[0143] In the formula, PFR represents the extreme volatility, with the unit being %. ε% represents the extreme fluctuation threshold. G represents the total number of days in the research period. G + (ε%) represents the number of days with a fluctuation amplitude greater than the maximum threshold. G - (ε%) represents the number of days with a fluctuation amplitude less than the minimum threshold.
[0144] 6) The electricity quantity reflects the total power generation of the PV power station during the research period, with the unit being MW·h.
[0145] Step 3.4, comprehensive evaluation of indicators: The entropy weight Topsis method is used to comprehensively analyze the six indicators of average fluctuation number, average fluctuation range, output distribution skewness, output distribution kurtosis, extreme volatility and electricity. The entropy weight method is first used to calculate the weight of each indicator of the scenario set, and then the Topsis analysis method is used for comprehensive scoring. Because the quality of the typical scenario set needs to be judged based on the various indicators of the original scenario set, the six indicator types of average fluctuation number, average fluctuation range, output distribution skewness, output distribution kurtosis, extreme volatility and electricity are set as intermediate indicators. The intermediate optimal value is the value of each indicator of the original scenario set. The closer to the intermediate optimal value, the better the indicator effect.
[0146] Embodiment 2:
[0147] The measured output of a 50MW photovoltaic power station in Yunnan in 2018 was used as the research object to verify the photovoltaic output scene clustering method proposed in the present invention. The output data was collected every 15 minutes, and each daily output curve had 96 points. The Lagrange interpolation method was used to complete the missing photovoltaic output data, and the data with large deviations were removed. In order to verify the advantages of the algorithm of the present invention compared with the K-means algorithm, the KESD_means algorithm based on comprehensive similarity measurement, the KED_shape algorithm based on optimized centroid extraction, the KSC (K-shape) algorithm, the KDTW-shape algorithm based on dynamic bending distance, and the K-medoids algorithm. The evaluation results of the typical scene set indicators are shown in Tables 1 and Figure 5As shown in the figure, the analysis shows that: ① For the algorithms of the present invention including the centroid extraction rule and the mean method for centroid extraction, the K-means algorithm, the KESD_means algorithm, the KED_shape algorithm, the KSC algorithm, and the KDTW-shape algorithm, the power is equal to the original power of 857.80 MW·h. However, the power of the K-medoids algorithm is 866.70 MW·h. The reason is that the K-medoids algorithm uses the sample with the minimum sum of distances to other samples in the class as the centroid, but this centroid cannot accurately represent the power of each curve within the entire cluster. Therefore, there is a difference in the power index between the typical scenario set and the original scenario set; ② The extreme volatility index has the largest difference from the original scenario set. The extreme volatility index of the original scenario set is 0.03. Among them, the extreme volatility of the KDTW_Shape algorithm is 8.3 times that of the original scenario set, which is much higher than the algorithms of the present invention, the K-means algorithm, the KESD_means algorithm, the KED_shape algorithm, the KSC algorithm, and the K-medoids algorithm. Since the extreme volatility reflects the proportion of the number of days when the fluctuation exceeds or is lower than the extreme fluctuation threshold, if it differs too much from the original, it will seriously affect the result quality. Compared with the K-means algorithm, the KESD_means algorithm, the KED_shape algorithm, the KSC algorithm, the KDTW_Shape algorithm, and the K-medoids algorithm, when calculating the morphological trend similarity, the algorithm of the present invention uses the Z-score method to standardize the curve, reducing the error caused by extreme output. Therefore, the extreme volatility index of the algorithm of the present invention is 0.01, which is lower than the original index and is closest to the original index; ③ The fluctuation coefficient, the fluctuation amplitude, the power distribution skewness, and the power distribution kurtosis are relatively close. However, when using the same similarity metric distance for division, the algorithms of the present invention, the KED_shape algorithm, and the KDTW-shape algorithm that adopt two-stage centroid extraction perform better. For example, when comparing the algorithm of the present invention with the KESD_means algorithm, the index values of the algorithm of the present invention are all higher than those of the KESD_means algorithm. When the centroid extraction method is the same, the comprehensive similarity metric method adopted is superior to the Euclidean distance and the dynamic warping DTW distance in dividing the scenario, indicating that the algorithm of the present invention has advantages in both curve category division and centroid extraction. Finally, through the comprehensive evaluation analysis of entropy weight Topsis, it can be seen that the KDTW-shape algorithm and the KSC algorithm based on the dynamic warping DTW distance have poor effects when applied to the clustering of photovoltaic output with large fluctuations. When comprehensively evaluating and reducing the impact of power difference, the score of the K-medoids algorithm is second only to the algorithm of the present invention. The reason is that the average fluctuation times and the average fluctuation amplitude index adopt the annual average results, and the representativeness of the K-medoids centroid will be weakened within a single solar term interval, while the quality of the typical scenario set generated by the algorithm of the present invention has good performance in both single-index and comprehensive scores.
[0148] Using the algorithm of the present invention, the comprehensive index evaluations of the typical scenario sets generated by different intervals divided by solar terms, seasons, months, and the whole year are shown in Table 2 and Figure 6 as follows. It can be seen that the typical scenario set generated by taking the twenty-four solar terms as the interval has a higher score and the result is the most representative. Through analysis, when taking the whole year as the clustering interval, the error between the results of each index and the original scenario is the largest, and the law characteristics of photovoltaic output cannot be reflected. As the interval division becomes denser, the results of each index obtained are closer to the original output scenario set. When dividing the time period by month and by season, the number of sub-intervals is 12 and 4 respectively, that is, the number of month intervals is 3 times that of the season interval. However, the comprehensive score by month is only 0.4 higher than the comprehensive score by season of 0.13. When divided by solar terms into 24 sub-intervals, the number of intervals is only 2 times that of the month division, but the comprehensive score has reached 0.24, which is 0.7 higher than that by month, indicating that the photovoltaic output between solar terms is more regular, and thus the quality of the typical scenario set obtained by clustering is the best.
[0149] Table 1: Comparison of the quality of scenario sets of different algorithms
[0150]
[0151] Remarks: K fluc represents the average number of fluctuations, unit: times; A avg represents the average fluctuation amplitude, unit: %; S avg represents the skewness of the output distribution; PFR represents the extreme volatility (taking 0.8 for %), unit: %; K avg represents the kurtosis of the output distribution; E represents the electricity quantity, unit: MW·h.
[0152] Table 2: Comparison of the quality of scenario sets of different intervals
[0153]
[0154] Remarks: K fluc represents the average number of fluctuations, unit: times; A avg represents the average fluctuation amplitude, unit: %; S avg represents the skewness of the output distribution; PFR represents the extreme volatility (taking 0.8 for %), unit: %; K avg represents the kurtosis of the output distribution; E represents the electricity quantity, unit: MW·h.
Claims
1. A photovoltaic typical output scenario clustering method considering comprehensive similarity measurement, characterized in that: First, consider the characteristics of photovoltaic power generation. The distance between curves is comprehensively measured by the similarity of the electricity quantity, the similarity of the morphological trend, and the similarity of the fluctuation position of the photovoltaic output curve. Then, a clustering reduction model for photovoltaic scenarios is constructed. The initial center point is selected based on the dissimilarity matrix, the comprehensive measurement distance is used as the basis for sample division, the typical scenarios are extracted by the two-stage centroid optimization method, and the optimal number of clusters is determined by the effectiveness evaluation index. The traditional K-means algorithm is improved respectively. Finally, six indicators, namely the average number of fluctuations, the average fluctuation amplitude, the skewness of the output distribution, the kurtosis of the output distribution, the extreme volatility, and the electricity quantity, are selected to comprehensively evaluate the results of the typical scenario set through the entropy weight TOPSIS method; The specific operation steps are as follows: Step 1, comprehensive similarity measurement of photovoltaic curves: Consider the characteristics of photovoltaic power generation. The distance between curves is comprehensively measured by the similarity of the electricity quantity, the similarity of the morphological trend, and the similarity of the fluctuation position of the photovoltaic output curve; Step 2, construct a clustering reduction model for photovoltaic scenarios: The initial center point is selected based on the dissimilarity matrix, the comprehensive measurement distance is used as the basis for sample division, the typical scenarios are extracted by the two-stage centroid optimization method, and the optimal number of clusters is determined by the effectiveness evaluation index. The traditional K-means algorithm is improved respectively; Step 3, verification of the typical scenario set: Six indicators, namely the average number of fluctuations, the average fluctuation amplitude, the skewness of the output distribution, the kurtosis of the output distribution, the extreme volatility, and the electricity quantity, are selected to comprehensively evaluate the results of the typical scenario set through the entropy weight TOPSIS method; The detailed steps of constructing the clustering reduction model for photovoltaic scenarios in Step 2 are as follows: Step 2.1, data preprocessing: The Lagrange interpolation method is used to complete the missing photovoltaic output data, and the data with large deviations are removed to complete the data cleaning; Step 2.2, initialize the clustering parameters: Set the maximum number of clusters, the number of iterations, and the initial number of clusters; Step 2.3, determine the initial clustering centers: Based on the constructed dissimilarity matrix, two measurement criteria, namely the mean dissimilarity and the overall dissimilarity, are defined, and the initial clustering centers are determined according to these criteria. The dissimilarity matrix is an object structure that stores the dissimilarities formed between several curves. The dissimilarity between two curves is defined by the comprehensive similarity distance. Suppose several photovoltaic output curves are represented by the matrix where, , represents the th curve, represents the th curve's output at moment, with the unit of MW, represents the curve number, n represents the total number of curves, m represents the total number of curve sampling points. Thus, the dissimilarity matrix between curves is obtained as follows: (8) In the formula, represents n the curve dissimilarity matrix; n represents the total number of curves; represents curve and dissimilarity; represents the th curve; represents the th curve; and represent the curve numbers; The definition of mean dissimilarity is as follows: (9) In the formula, represents the mean dissimilarity of the curve ; represents the dissimilarity between the curves and ; represents the th curve; represents the th curve; and represent the curve numbers; represents the total number of curves; reflects the position of the curve . The larger its value, the sparser the distribution of the curves around and the higher the degree of discreteness from other curves, and vice versa, the denser. The definition of overall dissimilarity is as follows: (10) In the formula, represents n the overall dissimilarity of the curves; represents the dissimilarity between curve and ; n represents the total number of curves; represents the -th curve; represents the -th curve; and represent the curve numbers; is related to the distribution of all curves and reflects the sparsity of the curve distribution; Step 2.4, similarity measurement and division: First, according to Steps 1.1, 1.4, and 1.5, calculate the three measurement distances from each curve in the sub-interval to the initial clustering center respectively. Then, combine the three measurement methods and calculate the comprehensive similarity distance according to Step 1.
6. Then, according to the comprehensive similarity between each curve in the interval and the clustering centroid, divide the curves into the category that is most similar to them; Step 2.5, update the clustering center: First, because cross-correlation is used to calculate the similarity between curve morphologies, the optimization goal in the first stage is set to maximize the sum of the squares of the comprehensive similarity distances between the morphological centroid and other curves in the cluster. The formula for the morphological centroid is expressed as follows: (11) In the formula, represents the morphological centroid of the k th class; represents the set of original curves where the centroid to be calculated is located; represents the th curve; represents the k number of curves in the th class; represents the actual centroid of the k th class; k represents the class label; represents the sum of the squares of the cross-correlations of all curves in the cluster; Secondly, the idea of multiplying by the same ratio is used to amplify the morphological centroid as the second stage of centroid solution. The morphological centroid is obtained by optimizing the standardized curve set. The ratio of the mean value of the original sample curves in the cluster and the mean value of the morphological centroid is set as the multiplying factor for the same ratio amplification. The formula for the actual centroid is expressed as follows: (12) Wherein, represents the actual centroid of the k th class; represents the number of curves in the k th class; represents the output of the k th curve in the i th class at the t th moment, with the unit of MW; k represents the class label; represents the curve number; t represents the t th sampling point; m represents the total number of curve sampling points; represents the value of the morphological centroid of the k th class at the t th moment; represents the morphological centroid of the k th class; Step 2.6, Cluster Validity Evaluation: Based on the principle of minimizing the sum of the comprehensive similarity distances from the samples within the class to the centroid, a cluster validity index is proposed W to determine the optimal number of clusters, W The formula is expressed as follows: (13) Wherein, W represents the sum of the comprehensive similarity distances from the centroid curve to the intra-class curves; represents the number of clusters; represents the k number of curves in the k th class; represents the class label; represents the k th curve in the i th class; represents the comprehensive similarity distance between the and curves; W will gradually decrease as the number of clusters increases; W The inflection point of the curve indicates that the decrease in the sum of the comprehensive similarity distances from the centroid to the intra-class curves is very small when the number of clusters is increased after this point, that is, W the K value corresponding to the inflection point of the curve is the optimal number of clusters; Current K After the value clustering is completed, calculate the clustering validity index according to Equation in Step 2.6 W , then, let , return to Step 2.3, when reaches the maximum, end the loop, and take W The K value corresponding to the inflection point of the curve is the optimal number of clusters.
2. A photovoltaic typical output scenario clustering method considering comprehensive similarity measurement according to claim 1, characterized in that The detailed steps of the comprehensive similarity measurement of photovoltaic curves in Step 1 are as follows: Suppose two photovoltaic output curves are and , where m represents the total number of curve sampling points; represents the output of curve x at moment, with the unit of MW; represents the output of curve y at moment, with the unit of MW; Step 1.1, using the output curve x and y The difference in the area enclosed by the coordinate axes is used as the similarity of the electricity quantities of the two. The distance formula for the similarity of the electricity quantities is expressed as follows: (1) In the formula, represents the similarity distance of the power quantity between curve and ; t represents the t th sampling point; m represents the total number of curve sampling points; and respectively represent the output of the curve of the th sampling point for curve and , with the unit of MW; represents the absolute value of the power quantity difference between curve and ; represents the larger value of the power quantities of curve and . The smaller the value, the more similar the power quantities of the two are; Step 1.2, to ensure the invariance of curve scaling, the Z-Score standardization method is selected to standardize the two curves to reduce the impact of the change in output amplitude on the correlation analysis; Step 1.3, when calculating the cross-correlation of two curves, fix one curve and slide it along the horizontal axis from the initial position in steps of 1 to the position aligned with and replace the data at the original position of the curve with 0 at the positions where the curve is not aligned with the curve . Thus, the cross-correlation sequence function expressions of the curves and and are as follows: (2) (3) Wherein, represents the cross-correlation sequence and ; where , wherein represents taking when the cross-correlation value between curve and ; represents the length of sequence when the moving step is , ; represents the moving step of curve along the horizontal axis, ; m represents the total number of sampling points of the curve; represents the sum of the dot products of the two curves each time curve moves along the horizontal axis; h represents the position subscript where curve and are aligned. The position corresponding to the maximum cross-correlation value can be obtained through the cross-correlation sequence , and then according to , the displacement of curve at this time and the optimal displacement curve of curve can be obtained; Step 1.
4. To ensure the invariance of curve displacement, the coefficient normalization method is used to process the cross-correlation sequence so that the autocorrelation coefficient at zero lag is equal to 1. The coefficient normalization formula is as follows: (4) In the formula, represents the curve and the cross-correlation sequence after coefficient normalization, where the data value range in the sequence is [-1, 1]; represents the cross-correlation sequence of the curves and ; represents the length of the sequence when the moving step is ; represents the moving step of the curve along the horizontal axis, ; represents the autocorrelation value of the curve ; represents the autocorrelation value of the curve ; when the data in the sequence is 1 or -1, it means that the shapes of the two curves are positively and negatively correlated at this time. For the convenience of calculation, the morphological trend similarity distance between the curves and is defined as follows: (5) In the formula, represents the morphological trend similarity distance between the curves and , and its range is [0, 2]; represents the maximum value in the cross-correlation sequence after coefficient normalization of the curves and . The smaller the value, the more similar the morphological trends of the curves and are; Step 1.
5. The deviation of the fluctuation occurrence positions of the two curves is used as the similarity of the fluctuation positions of the two curves. The formula for the similarity distance of the fluctuation positions is expressed as follows: (6) In the formula, represents the similarity distance of the fluctuation positions of the curves and ; represents the absolute value of the displacement when the curve y slides to the position where its morphological trend is most similar to that of x ; m represents the total number of sampling points of the curve. The smaller the value, the more similar the fluctuation positions of the curves and . Step 1.
6. To reflect the comprehensive similarity between the power generation curves, based on the similarity of the electricity quantity, the morphological trend similarity and the similarity of the fluctuation positions are considered to obtain a comprehensive similarity metric distance applicable to the photovoltaic power generation curves. The formula is expressed as follows: (7) In the formula, represents the comprehensive similarity distance and ; represents the similarity distance of the electricity quantity magnitude and ; represents the similarity distance of the morphological trend and ; represents the similarity distance of the fluctuation position and , and the smaller the comprehensive similarity metric distance , the more similar the two curves are. 3. A photovoltaic typical output scenario clustering method considering comprehensive similarity measurement according to claim 1, characterized in that, The detailed steps for verifying the typical scenario set in Step 3 are as follows: Step 3.
1. Time interval division: The twenty-four solar terms, as a very important time division method in traditional Chinese culture, reflect the position of the sun on the earth and the change of light intensity. The power generation power of the photovoltaic system is directly related to the solar light intensity. At the same time, taking the twenty-four solar terms as the interval can also consider the influence of seasonal and climatic factors. Specifically, it includes: the temperature gradually rising in spring, the temperature peak in summer, the temperature gradually decreasing in autumn, and the temperature trough in winter. These factors will also affect the output of the photovoltaic power generation system. Therefore, using the twenty-four solar terms as the clustering interval can better reflect the output characteristics and change rules of the photovoltaic system in different time periods; Step 3.
2. Generation of scenario set replacement: According to the principle that the clustering label corresponds to the original date label, several typical output scenarios matching the optimal number of clusters within the solar term interval are used to replace the typical output curves obtained by clustering with the original position output curves, so as to obtain a new typical solar term output scenario; Step 3.
3. Selection of scenario set evaluation indicators: Based on the characteristics of photovoltaic power generation, an index system including curve fluctuation characteristics and electricity quantity characteristics is established to verify the quality of the typical scenario set. Specifically, it includes 6 indicators: average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility, and electricity quantity. The formulas for the average number of fluctuations and extreme volatility are as follows: Average number of fluctuations, this indicator reflects the average number of output reversals within the research period: (14) In the formula, represents the average number of fluctuations, with the unit of times; represents the total number of days in the research period; g represents the g th day within the research period; t represents the t th sampling point; m represents the total number of curve sampling points; P t-1 、P t 、P t+1 respectively represent the outputs at the t - 1, t, t + 1 th moment, with the unit of MW. Among them, represents that the output reverses once, represents that the output does not reverse; Extreme volatility, this indicator reflects the proportion of the number of days when the photovoltaic power generation power exceeds or is lower than the specified extreme fluctuation threshold to the total number of days in the research period: (15) In the formula, represents the extreme volatility, unit: %; represents the extreme volatility threshold; represents the total number of days in the research period; represents the number of days with a volatility amplitude greater than the maximum threshold; represents the number of days with a volatility amplitude less than the minimum threshold; Step 3.
4. Comprehensive evaluation of indicators: The entropy weight Topsis method is used to comprehensively analyze the 6 indicators of average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility, and electricity quantity. First, the entropy weight method is used to calculate the weights of each indicator of the scenario set, and then the Topsis analysis method is used for comprehensive scoring. Since the quality of the typical scenario set needs to be judged based on the indicators of the original scenario set, the 6 indicator types of average number of fluctuations, average fluctuation amplitude, skewness of output distribution, kurtosis of output distribution, extreme volatility, and electricity quantity are all set as intermediate type indicators. The intermediate optimal value is the value of each indicator of the original scenario set. The closer to the intermediate optimal value, the better the indicator effect.
Citation Information
Patent Citations
Cluster-based wind power output fluctuation interval prediction method and device and storage medium
CN112052981A
Device for testing adhesion performance of soft coating in pipe and characterization method
CN114088621A