A method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering
Patent Information
- Application Number
- CN202610924822.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
AI Technical Summary
[0005]有鉴于此,本发明的目的在于提供一种基于多维综合相似距离聚类的光伏典型场景提取方法,解决传统方法难以兼顾宏观形态与微观数值差异及计算开销大的问题
首先,本发明引入多维综合相似距离,有效克服了单一欧氏距离度量难以兼顾时间序列绝对数值差异与动态波动规律的局限,能够全面刻画光伏出力的幅值、波动时序及总能量特征。
Smart Images

Figure CN122778089A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of stochastic programming and data processing technology for new energy power systems, and relates to a method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering. Background Technology
[0002] With the normalization of high-proportion photovoltaic grid connection, accurately extracting representative operating scenarios has become the core foundation for stochastic planning and reliable decision-making in power systems. However, photovoltaic output is greatly affected by meteorological factors, exhibiting strong randomness, volatility, and intermittency.
[0003] Currently, the main methods for extracting representative photovoltaic scenarios are still K-menas clustering, as well as dimensionality reduction of photovoltaic scenarios using autoencoded time series clustering and Wasserstein distance. K-menas clustering is efficient, but it cannot simultaneously take into account the shape and trend of numerical values and photovoltaic time series curves, as well as the consistency of daily power generation. Although autoencoded time series clustering and Wasserstein distance can reduce the photovoltaic time series scenarios, they are computationally efficient and have a significant impact on model accuracy when facing extreme weather.
[0004] Therefore, how to construct a lightweight clustering method and create a more reasonable model for extracting representative photovoltaic scenarios has become a key issue to be addressed. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method for extracting typical photovoltaic scenes based on multidimensional comprehensive similarity distance clustering, which solves the problems of traditional methods being unable to take into account the differences between macroscopic morphology and microscopic numerical values and having large computational overhead.
[0006] Because multi-granularity spheres can significantly compress data volume and improve clustering efficiency, and gray clustering can effectively capture photovoltaic power fluctuation characteristics and is more adaptable to noisy and uncertain data, the combination of the two can efficiently handle photovoltaic time series clustering problems. Therefore, the method of this invention considers the trend of curve shape and the similarity of generated electricity in clusters from multiple dimensions of photovoltaic power generation. A multi-dimensional comprehensive similarity distance is introduced in conjunction with spheres to perform multi-granularity partitioning and data compression of photovoltaic time series data, thereby extracting representative sphere centers and laying the foundation for subsequent efficient clustering. Gray relational clustering is introduced into the clustering module to cluster the sphereized photovoltaic time series data. A novel architecture of multi-dimensional comprehensive similarity distance sphere gray clustering is constructed.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering specifically includes the following steps: S1: Collect historical power data from multiple photovoltaic sites, and reconstruct the daily photovoltaic sequence data into a single sample by sequentially connecting them in series, thereby preserving the spatial coordination characteristics between different sites and their own temporal evolution characteristics; construct a multi-dimensional comprehensive similarity distance matrix that integrates Euclidean distance, morphological trend similarity, and power distance, and use particle swarm optimization algorithm to determine the optimal weight coefficients; S2: Calculate the similarity between samples based on the multidimensional comprehensive similarity distance and construct an initial similarity queue; set the similarity threshold and the control range of the particle size to generate multi-granularity particles; calculate the sample mean inside each particle as the particle center to realize coarse-grained reconstruction and compression of photovoltaic data; S3: Input the set of sphere center points obtained by sphere calculation into the gray relational clustering model; use the K-means algorithm to determine the initial clustering points, and calculate the gray Dunk correlation degree and gray distance between each sphere center point and all initial clustering points; S4: Select the initial cluster point with the smallest gray distance to assign the spheres, iteratively update the cluster center to the average value of the currently assigned samples, until the cluster center no longer changes or the maximum number of iterations is reached, and finally output the cluster label of each sample and the final set of cluster centers to extract the representative photovoltaic scenarios.
[0008] Furthermore, in step S1, photovoltaic data is collected at certain intervals, and multiple power stations can also be extracted in series.
[0009] Furthermore, in step S1, the formula for calculating the multidimensional comprehensive similarity distance is:
[0010] in, Represents the multidimensional comprehensive similarity distance; , , These are the weights used to constrain the multidimensional distance, ranging from 0 to 1, and the sum of the weights equals 1; given a set of photovoltaic time series sequences. ,in Indicates the first The multi-station cascaded output vector of a sample, or the first sample... A photovoltaic time series, ; This represents the Euclidean distance metric, used to measure the difference in absolute amplitude between two photovoltaic curves. The calculation formula is as follows:
[0011] in, For the total dimension of the sample, Indicates the first The sample at the th Dimension (or the) The output value at each point in time; Electricity difference measurement: The area of the photovoltaic output curve represents the total daily power generation. The electricity similarity distance is introduced to constrain samples with similar total output scale to be clustered into one class. The formula for representing the electrical similarity distance is:
[0012] in, For the first The total power generation of each sample; (3) Indicates the first The first sample and the first The formula for calculating the morphological trend distance between samples is:
[0013] in, Cross-correlation sequences and The normalized cross-correlation coefficient has a range of [-1, 1]. The range is from 0 to 2. When the morphological trend distance approaches 0, it indicates that the morphological trends of the two time series are the same; when it approaches 2, it indicates that the morphological trends of the samples are opposite.
[0014] Furthermore, in step S1, the cross-correlation sequence and The formula for calculating the normalized cross-correlation coefficient is:
[0015] in, , For sequence Total energy; L The sequence length; Let sequence Compared to produce The sequence of cross-correlation inner products of the sliding overlap portion after a translation of each time step. Defined as:
[0016] Among them, translation lag .
[0017] Furthermore, in step S1, the optimal weight coefficients are determined using the particle swarm optimization algorithm, keeping both the morphological trend and the coefficient of variation below 0.4. The morphological trend metric is the morphological trend distance, and the coefficient of variation... The calculation formula is:
[0018] in, This represents the standard deviation of daily power generation for all samples in the cluster. This represents the average daily power generation of all samples in the cluster.
[0019] Furthermore, in step S2, the similarity between samples... The calculation formula is: The value is used to measure the similarity between two photovoltaic samples; the closer it is to 1, the more similar they are.
[0020] Furthermore, in step S2, the center point of each sphere is used to replace the sample in the sphere, where the center point of the sphere is... The calculation formula is:
[0021] in, express Internal sample size Represented as the first z Each ball.
[0022] Furthermore, in step S3, the centers of the spheres are clustered using a grey relational clustering model, and the grey relational degree of the spheres is defined as:
[0023] in, For the first The cluster centers at the in Wei and Di The center of each grain is in the... Grey relational coefficient of dimension, For the first Cluster centers The Values of each dimension For the first The center of the grain Values for each dimension; The resolution coefficient is usually taken as... This is used to adjust the size of the comparison environment. , No. Each ball With the Cluster centers Overall grey relational degree The calculation formula is:
[0024] in, The total dimension of the sample, i.e., the length of the photovoltaic time series. The value range is from 0 to 1. The larger the size, the more likely it is to be a pellet. and cluster center The more relevant the relationship, the more we define the gray distance. Let the sphere be and cluster center gray distance The calculation formula is: The smaller the gray distance, the more likely it is to be a pellet. and cluster center The closer it is.
[0025] The beneficial effects of this invention are as follows: First, this invention introduces a multi-dimensional comprehensive similarity distance, which effectively overcomes the limitation that a single Euclidean distance metric cannot take into account both the absolute numerical differences and dynamic fluctuation patterns of time series, and can comprehensively characterize the amplitude, fluctuation time series and total energy characteristics of photovoltaic power output.
[0026] Secondly, this invention introduces multi-granularity particle sphere computation, which uses multi-dimensional comprehensive similarity distance to construct particle spheres, greatly reducing and compressing the original high-dimensional full data to at least a few particle sphere centers, effectively avoiding the computational bottleneck of massive data, significantly improving clustering efficiency, and having natural robustness to complex meteorological disturbances and local noise points.
[0027] Finally, this invention constructs a novel architecture combining granular clustering and grey relational clustering, effectively integrating the dimensionality reduction capability of granular clustering with the strong adaptability of grey clustering to uncertain information. The typical scenarios extracted by this method demonstrate excellent performance in probabilistic power flow simulation, morphological trend similarity distance, and coefficient of variation, providing solid theoretical and data support for capacity configuration and bidding decisions in new power systems.
[0028] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0029] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This diagram illustrates the modeling process of the photovoltaic typical scene extraction model based on multidimensional comprehensive similarity distance and granular gray relational clustering, as presented in this invention. Detailed Implementation
[0030] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0031] Please see Figure 1 This invention provides a method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance and granular-spherical gray relational clustering, specifically including the following steps: S1: Retrieve historical power generation data of wind and solar power generation equipment at specified time intervals.
[0032] The acquired historical power generation data represents the power generation output of a specific power generation device in a certain area within a certain time interval. The specified time interval can be determined based on forecasted demand (e.g., 1 hour), or data from multiple sites can be chained together.
[0033] S2: Based on multidimensional comprehensive distance Construct a similarity matrix, where , , It is used to constrain the weight ratio of multidimensional distance. The weight ranges from 0 to 1, and the weights are added together to equal 1.
[0034] (1) Euclidean distance metric: used to measure the difference in absolute amplitude between two photovoltaic curves, the calculation formula is: ,in This represents the total dimension of the sample.
[0035] (2) Measurement of power disparity: The area under the photovoltaic output curve represents the total daily power generation. A power similarity distance is introduced to constrain samples with similar total output to cluster into one class. Let the sample... Total power generation is The formula for calculating the electrical similarity distance is:
[0036] (3) Morphological trend measurement: Let the two photovoltaic characteristic sequences be... and Their sequence lengths are all To explore the waveform similarity of two sequences under different time sliding windows, a cross-correlation function is introduced. Let the sequences... Compared to produce The sequence of cross-correlation inner products of the sliding overlap portion after a translation of each time step. Defined as:
[0037] Among them, translation lag This cross-correlation sequence comprehensively reflects the states of the two feature vectors under all possible time lags, and the maximum cross-correlation value and its corresponding position can be obtained through different sliding operations.
[0038] To eliminate absolute numerical magnitudes in the sequences and improve data comparability, the cross-correlation sequences are normalized to obtain the normalized cross-correlation coefficient: ,in After normalization, The value range is [-1, 1]. Therefore, the morphological trend distance between samples is... Based on the maximum cross-correlation coefficient, the following is constructed:
[0039] in, The range is from 0 to 2. When the morphological trend distance approaches 0, it indicates that the morphological trends of the two time series are the same; when it approaches 2, it indicates that the morphological trends of the samples are opposite.
[0040] Then, the optimal weight coefficients are determined using the particle swarm optimization algorithm, keeping both the morphological trend and the coefficient of variation below 0.4. The formula for measuring the morphological trend is as described above, and the formula for the coefficient of variation is: ,in This represents the standard deviation of daily power generation for all samples in the cluster. This represents the average daily power generation of all samples in the cluster.
[0041] S3: The process for generating spheres through multidimensional comprehensive similarity distance is as follows: S31: Based on data from five photovoltaic sites and multidimensional similarity distance Construct a similarity matrix, where For multidimensional comprehensive similarity distance, For Euclidean distance, For electrical distance, The three factors (shape, trend, and distance) are weighted and combined. The aim is to enable the model to better filter out representative scenarios with similar values, equal daily electricity consumption, and relatively consistent fluctuation patterns, thereby improving the scientific rigor of photovoltaic sequence data clustering.
[0042] S32: Each day's photovoltaic sequence data is considered as one sample, totaling 164 samples; all samples exceeding a threshold are selected using a similarity matrix. The samples are added to the similarity queue; S33: Generate spheres using a similarity queue. When the proportion of a candidate sample that is similar to existing samples in the current sphere exceeds a threshold... Only then is it included in the particle. Parameters This determines the internal density and similarity of the generated granules; S34: Control the range of particle size and calculate the sample mean of particles. Replaces the center of the grain.
[0043] S4: The gray relational clustering model process is as follows: S41: Generate spheres by converting multidimensional comprehensive similarity distance into sample similarity, where the sample similarity formula is: The value is used to measure the similarity between two photovoltaic samples; the closer it is to 1, the more similar they are.
[0044] S42: Use grey relational clustering to cluster the centers of spheres.
[0045]
[0046] for , No. Each ball With the Cluster centers Overall grey relational degree The calculation formula is:
[0047] in, The total dimension of the sample is the length of the photovoltaic time series. The value range is from 0 to 1. The larger the size, the more likely it is to be a pellet. and cluster center The more relevant it is.
[0048] S43: Define gray distance, assuming a particle size distribution. and cluster center gray distance The calculation formula is: The smaller the gray distance, the more likely it is to be a pellet. and cluster center The closer the points are, the better. Based on the initial cluster centers, clustering is performed using the gray distance from the center points.
[0049] Example 1: This embodiment provides a basic photovoltaic scene extraction process: First, data acquisition and preprocessing are performed. Taking a photovoltaic power generation cluster in a coastal area as an example, its historical power generation output data is collected. Given a set of photovoltaic time series sequences... A similarity matrix is constructed based on the multidimensional comprehensive similarity distance.
[0050] Example 2: This embodiment provides a pellet generation process: The basic model formula for generating spheres is as follows: (Based on the similarity matrix, spheres are generated.)
[0051]
[0052] in, This represents the number of samples in the set. In the formula of Example 2, Set a threshold for the model. These are different photovoltaic samples from the same sphere. express The total number of all particles generated on the surface. A This is the complete set of original photovoltaic samples. B For in granules In, with the sample The similarity is greater than or equal to the threshold The sample set, n The total number of samples. The multidimensional comprehensive similarity distance is defined. For those in the granules There are any satisfy , will meet the conditions denoted as set ,but This indicates that a sample in a sphere has a similarity greater than or equal to any different sample in the same sphere. The sample set. Set similarity threshold For any unassigned sample Search The initial sphere is generated from the neighboring nodes of the sample, and an initial similarity queue is constructed for its neighboring nodes. Subsequent spheres are generated based on the similarity queue. A sample is considered sufficiently similar to a sample in a sphere, meaning the similar samples in that sphere account for a certain percentage of the total samples in the sphere. Only after this process is complete should the sample be added to the sphere. To ensure sphere quality and sample coverage, the sphere size is set. , This limits excessive granule growth, and the granules close when certain conditions are met.
[0053] Example 3: This embodiment provides an application of granular-spheroid gray relational clustering, specifically: Clustering is performed by replacing the original sample points with the centroids of the spheres. The set of centroids is then input into the clustering model, and K-means is used to determine the initial clustering points. Let the first centroid be... The target cluster centers are: The center point of the particle is introduced into the gray Dunk correlation degree, and the particle-sphere gray correlation degree is defined as follows: Let
[0054] Then there is
[0055] Among them, for , No. Each ball With the Cluster centers Overall grey relational degree The definition is as follows:
[0056] in, The total dimension of the sample, i.e., the length of the photovoltaic time series. The value range is from 0 to 1. The larger the size, the more likely it is to be a pellet. and cluster center The more relevant it is, the gray distance is defined for this purpose; let the particle be a sphere. and cluster center gray distance The calculation formula is: The smaller the gray distance, the more likely it is to be a pellet. and cluster center The closer it is.
[0057] Therefore, after the photovoltaic dataset is processed to obtain a set of sphere center points, it is input into a grey relational clustering model to cluster the sphere center points. k-means is used to determine the initial cluster points, and the grey distance from each sphere center point to all initial cluster points is calculated. The initial cluster point with the smallest grey distance is selected and added to that cluster. This process continues until the sphere is covered by all clusters. Finally, the mean of all clusters is calculated to replace the cluster center points.
[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering, characterized in that, The method specifically includes the following steps: S1: Collect historical power data from multiple photovoltaic sites, and reconstruct the daily photovoltaic sequence data into a single sample by sequentially connecting them in series, preserving the spatial coordination characteristics between different sites and their own temporal evolution characteristics; construct a multi-dimensional comprehensive similarity distance matrix that integrates Euclidean distance, morphological trend similarity, and power distance, and use particle swarm optimization algorithm to determine the optimal weight coefficients; S2: Calculate the similarity between samples based on the multidimensional comprehensive similarity distance and construct an initial similarity queue; set the similarity threshold and the control range of the particle size to generate multi-granularity particles; calculate the sample mean inside each particle as the particle center to realize coarse-grained reconstruction and compression of photovoltaic data; S3: Input the set of sphere center points obtained by sphere calculation into the gray relational clustering model; use the K-means algorithm to determine the initial clustering points, and calculate the gray Dunk correlation degree and gray distance between each sphere center point and all initial clustering points; S4: Select the initial cluster point with the smallest gray distance to assign the spheres, iteratively update the cluster center to the average value of the currently assigned samples, until the cluster center no longer changes or the maximum number of iterations is reached, and finally output the cluster label of each sample and the final set of cluster centers to extract the representative photovoltaic scenarios.
2. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 1, characterized in that, In step S1, the formula for calculating the multidimensional comprehensive similarity distance is: in, Represents the multidimensional comprehensive similarity distance; , , These are the weights used to constrain the multidimensional distance, ranging from 0 to 1, and the sum of the weights equals 1; given a set of photovoltaic time series sequences. ,in Indicates the first The multi-station cascaded output vector of a sample, or the first sample... A photovoltaic time series, ; This represents the Euclidean distance metric, used to measure the difference in absolute amplitude between two photovoltaic curves. The calculation formula is as follows: in, For the total dimension of the sample, Indicates the first The sample at the th The output value of the dimension; The formula for representing the electrical similarity distance is: in, For the first The total power generation of each sample; (3) Indicates the first The first sample and the first The formula for calculating the morphological trend distance between samples is: in, Cross-correlation sequences and The normalized cross-correlation coefficient has a range of [-1, 1]. The range is from 0 to 2. When the morphological trend distance approaches 0, it indicates that the morphological trends of the two time series are the same; when it approaches 2, it indicates that the morphological trends of the samples are opposite.
3. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 1, characterized in that, In step S1, the cross-correlation sequence and The formula for calculating the normalized cross-correlation coefficient is: in, , For sequence The total energy is used for normalization. L The sequence length; Let sequence Compared to produce The sequence of cross-correlation inner products of the sliding overlap portion after a translation of each time step. Defined as: Among them, translation lag .
4. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 2, characterized in that, In step S1, the optimal weight coefficients are determined using the particle swarm optimization algorithm, keeping both the morphological trend and the coefficient of variation below 0.
4. The morphological trend metric is the morphological trend distance, and the coefficient of variation... The calculation formula is: in, This represents the standard deviation of daily power generation for all samples in the cluster. This represents the average daily power generation of all samples in the cluster.
5. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 2, characterized in that, In step S2, the similarity between samples The calculation formula is: The value is used to measure the similarity between two photovoltaic samples; the closer it is to 1, the more similar they are.
6. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 2, characterized in that, In step S2, the center point of each sphere is used to replace the sample in the sphere, where the center point of the sphere is... The calculation formula is: in, express Internal sample size Indicates the first z Each ball.
7. The method for extracting typical photovoltaic scenarios based on multidimensional comprehensive similarity distance clustering according to claim 1, characterized in that, In step S3, the gray relational clustering model is used to cluster the centers of the spheres, and the gray relational degree of the spheres is defined as: in, For the first The cluster centers at the in Wei and Di The center of each grain is in the... Grey relational coefficient of dimension, For the first Cluster centers The Values of each dimension For the first The center of the grain Values for each dimension; For the resolution coefficient; , No. Each ball With the Cluster centers Overall grey relational degree The calculation formula is: in, The total dimension of the sample is the length of the photovoltaic time series. The value range is from 0 to 1. The larger the size, the more likely it is to be a pellet. and cluster center The more relevant the relationship, the more we define the gray distance. Let the sphere be and cluster center gray distance The calculation formula is: The smaller the gray distance, the more likely it is to be a pellet. and cluster center The closer it is.