Method for generating optical-charge typical scenario set based on integrated clustering and frequent item set tree

Through the light-load typical scene set generation method based on integrated clustering and frequent item set trees, the problem of single influencing factors and poor correlation of load scenarios in the photovoltaic scene planning of the distribution network is solved, and the comprehensiveness and scientificity of the distribution network planning is achieved, and the photovoltaic on-site consumption capacity and the stability and reliability of the power system are improved.

CN115659191BActive Publication Date: 2025-07-01GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211289091.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2025-07-01
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

In the photovoltaic scene planning of distribution networks, the existing technology considers the single influencing factors and poor correlation with the load scenario, which leads to a large deviation between the planned operation scenario and the actual scenario, affecting the on-site photovoltaic absorption capacity and the safety and reliability of the power system.

Method used

The light-load typical scene set generation method based on integrated clustering and frequent item set trees is adopted. By acquiring and preprocessing load and photovoltaic data, multiple clustered scene sets are generated using integrated clustering method, the most representative typical scenes are selected, and the photovoltaic-load typical scenes are generated in combination with the meteorological association rule database to perform comprehensive distribution network planning.

Benefits of technology

It improves the comprehensiveness and scientificity of distribution network planning, improves the on-site consumption capacity of photovoltaic and the stability and reliability of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659191B_ABST
    Figure CN115659191B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set trees, which relates to the technical field of distribution network load and photovoltaic output scenario analysis. First, historical original load data and photovoltaic output data are obtained and preprocessed and classified. The integrated clustering method is used to cluster multiple data sets to obtain multiple clustering scenario sets. Then, the photovoltaic typical scenario set and the load typical scenario set are screened out. Considering the influence of different typical meteorological conditions at the same moment in the same area, the frequent item set tree algorithm is used to generate a meteorological association rule base, thereby establishing the correlation between the photovoltaic typical scenario set and the load typical scenario set. Finally, based on the meteorological association rule base, a photovoltaic-load typical association scenario set is generated. Under this scenario set, the integrated planning of the distribution network with photovoltaic is carried out, which is comprehensive and scientific, and effectively improves the local photovoltaic consumption capacity, as well as the stability and reliability of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network load and photovoltaic output scenario analysis, and more specifically, to a method for generating a typical light-load scenario set based on integrated clustering and frequent item set trees. Background Art

[0002] With more and more photovoltaic power stations being continuously connected to the distribution network, how to generate reasonable photovoltaic planning scenarios for the distribution network, so as to improve the local photovoltaic consumption capacity and the safe and reliable operation ability of the power system, has become a difficult problem.

[0003] Currently, there is a large deviation between the planned operation scenario and the actual photovoltaic scenario. The main reasons are as follows: (1) The high penetration of photovoltaic power is gradually changing the structure and operation mode of the distribution network, resulting in a large difference in the power distribution of the distribution network compared with the traditional distribution network; (2) Photovoltaic power generation is greatly affected by environmental factors such as solar radiation intensity and meteorological conditions, resulting in its volatility and randomness being much greater than that of traditional energy generation forms, thus increasing the fluctuation of the distribution network power flow distribution; (3) The photovoltaic volatility affects the distribution network load situation, resulting in a certain fluctuation of the load curve. It is necessary to consider the comprehensive impact of the superposition of photovoltaic and load volatility on the distribution network; (4) Due to safety considerations, the distribution network planning is too conservative, and the photovoltaic output is equivalent to a simple curve, resulting in problems such as increased investment costs or insufficient excavation of photovoltaic power generation potential.

[0004] In order to fully consider the uncertainties brought about by the comprehensive influence of photovoltaic and load, it is necessary to form a typical comprehensive operation scenario of the distribution network, and conduct a comprehensive planning of the distribution network with photovoltaic power under this scenario to improve the comprehensiveness and scientificity of the entire planning.

[0005] At present, the consideration in the generation of distribution network planning scenarios is still not comprehensive enough. In the current photovoltaic scenario, on the one hand, only the influence of a single temperature is considered, while ignoring that the photovoltaic scenario is the result of the combined action of multiple influencing factors. For example, a temperature and photovoltaic scenario generation method based on time-series correlation feedback correction is disclosed in the prior art. This method generates a relevant temperature set and the corresponding photovoltaic scenario according to the illumination scenario obtained by the Monte Carlo method, takes the illumination intensity scenario value and the temperature prediction value as inputs, uses the Kalman filter to obtain the temperature output under each scenario, takes it as the reference value for each scenario, and focuses on considering the mutual correlation between temperature and illumination intensity and the time-series autocorrelation of temperature when generating the uncertainty scenario set. On the other hand, it focuses on scenario application and simply generates typical photovoltaic scenarios, without reflecting the actual situation of typical scenarios from all aspects. For example, a power system voltage control method for collaborative multi-reactive power equipment based on photovoltaic scenarios is disclosed in the prior art. First, according to historical operation data, an initial scenario is generated, and then a typical scenario and an extreme scenario are obtained. Then, based on the generated typical scenario and extreme scenario, a voltage control optimization model is established. In the load scenario, the method of generating using a generative adversarial network is becoming more and more common. This method analyzes the statistical laws of historical data for prediction to generate new scenarios, has relatively high requirements for real data, and is actually difficult to operate. Summary of the Invention

[0006] To solve the problems of single consideration of influencing factors and poor correlation with the load scenario in the current photovoltaic scenario planning of the distribution network, the present invention proposes a method for generating a light-load typical scenario set based on integrated clustering and frequent item set trees, generates a typical comprehensive operation scenario of the distribution network, and conducts a comprehensive planning of the distribution network with photovoltaic under this scenario set, which has strong comprehensiveness and scientificity, and improves the in-situ consumption ability of photovoltaic and the reliability of the power system.

[0007] To achieve the above technical effects, the technical solution of the present invention is as follows:

[0008] A method for generating a light-load typical scenario set based on integrated clustering and frequent item set trees, the method comprising the following steps:

[0009] S1. Obtain the original load data of the distribution network to be planned and the photovoltaic output data connected to it within a certain period of time, and preprocess and classify the data to obtain multiple data sets;

[0010] S2. Use the integrated clustering method to cluster multiple data sets into different clusters in turn, thereby generating multiple clustering scenario sets;

[0011] S3. Use the comprehensive distance formula to screen out the most representative typical scenarios from each clustering scenario set, use the most representative typical scenarios as the labels of the corresponding clustering scenario sets, and finally convert all clustering scenario sets into a photovoltaic typical scenario set and a load typical scenario set;

[0012] S4. For the typical photovoltaic scenario set and the typical load scenario set, considering different meteorological influencing factors at the same moment in the same region, use the frequent item set tree algorithm to generate a meteorological association rule base. Based on the meteorological association rule base, generate a typical photovoltaic-load association scenario set.

[0013] In this technical solution, first, obtain the historical original load data and photovoltaic output data, perform preprocessing and classification, use the integrated clustering method to cluster multiple data sets to obtain multiple clustering scenario sets, then screen the typical photovoltaic scenario set and the typical load scenario set, consider the influence of different typical meteorological conditions at the same moment in the same region, use the frequent item set tree algorithm to generate a meteorological association rule base, thereby establishing the correlation between the typical photovoltaic scenario and the typical load scenario. Finally, based on the meteorological association rule base, generate a typical photovoltaic-load association scenario set. Under this scenario set, conduct a comprehensive distribution network planning with photovoltaic, which is highly comprehensive and scientific, effectively improving the local photovoltaic consumption capacity, as well as the stability and reliability of the power system.

[0014] Preferably, in step S1, when preprocessing the obtained original load data and photovoltaic output data, for the time dates with a small amount of missing original load data and photovoltaic output data, use the cubic spline interpolation method to fill them, and discard the time dates with a large amount of missing data.

[0015] Preferably, during classification, classify the load data into 3 data sets corresponding to weekdays, weekends, and holidays other than weekends, and classify the photovoltaic output data into 4 data sets corresponding to spring, summer, autumn, and winter according to the four seasons, thereby considering different electricity consumption days such as weekdays and holidays.

[0016] Preferably, after classification, perform feature extraction on the classified daily data sets, and select the clustering feature vectors of the photovoltaic data set and the clustering feature vectors of the load data set. Among them,

[0017] The clustering feature vector of the photovoltaic data set is expressed as:

[0018] F pv ={P d_max P d_sum P d_mean P d_std P d_difmax P d_difmin P d_difmean}

[0019] Among them, P d_max is the maximum value of the daily photovoltaic output; P d_sum is the total sum of the daily photovoltaic output throughout the day; P d_mean is the average value of the daily photovoltaic output; P d_std is the standard deviation of the photovoltaic output;d_difmax is the maximum value of the first-order difference of the PV output sequence within one day; P d_difmin is the minimum value of the first-order difference of the PV output sequence within one day, P d_difmean is the average value of the first-order difference of the PV output sequence within one day;

[0020] The clustering feature vector of the load dataset is expressed as:

[0021] F load ={L d_max L d_min L d_mean L d_std L d_difmax L d_difmin L d_difmean}

[0022] wherein, L d_max is the daily maximum load; L d_min is the daily minimum load; L d_mean is the daily average load; L d_std is the standard deviation of the daily load; L d_difmax is the maximum value of the first-order difference of the daily load; L d_difmin is the minimum value of the first-order difference of the daily load; L d_difmean is the average value of the first-order difference of the daily load, reducing the dimension of the initial clustering data.

[0023] Preferably, the integrated clustering method described in step S2 is the HDBSCAN algorithm, and the process of generating multiple clustering scenario sets is as follows:

[0024] S21. Measure the distance between data points in the dataset using the mutual reachability distance, traverse and calculate the mutual reachability distance between two data points in the dataset without repetition, and obtain the distance table of all data;

[0025] S22. Take any two points in each dataset as vertices, connect the vertices to get edges, and use the corresponding distance in the distance table as the weight of the edge. The entire dataset is transformed into a data distance weighted graph;

[0026] S23. Use the Prim algorithm to construct the minimum spanning tree of the data distance weighted graph to connect all distance points with the minimum distance;

[0027] S24. Sort all the edges in the minimum spanning tree in ascending order of distance, then select each edge in turn, classify the sub-datasets included in the two sub-graphs connected by the edge into one category each. After classification through the union-find set, a new category corresponding to each edge is obtained, and a clustering hierarchy is constructed;

[0028] S25. Determine the minimum number of clusters, traverse the clustering hierarchy from top to bottom. When classifying each subgraph, determine whether the number of the two sub-datasets generated by the classification is greater than the minimum number of clusters. If so, classify them into one category; otherwise, mark the classification category as scattered points and delete it. After traversing the entire clustering hierarchy, obtain a compressed clustering tree with a small number of categories.

[0029] S26. Assign a class label to each compressed category in the compressed clustering tree. Traverse the compressed clustering tree from bottom to top, and determine whether the stability of the parent category of each category is greater than the sum of the stabilities of the child nodes of this category. If so, all child nodes belong to this category, and output the clustering result; otherwise, set the stability of this category to the sum of the stabilities of its child nodes.

[0030] Here, the HDBSCAN algorithm converts DBSCAN into a hierarchical clustering algorithm, and then uses a stable clustering technique to extract a flat clustering to extend DBSCAN.

[0031] Preferably, in step S26, declare all leaf nodes in the compressed clustering tree as selected clusters. Define λ as a value measuring the persistence of a cluster. For a given cluster, define λ birth and λ death as the λ when the corresponding cluster splits and becomes its own cluster, and the λ value when the cluster splits into smaller clusters respectively; for each node in the cluster, define λ p as the λ value of the outlier of this point, a value between λ birth and λ death . For each cluster, calculate the stability as:

[0032]

[0033] Preferably, in step S3, assume that there are a total of Q scenarios in each clustering scenario set. The comprehensive distance formula includes cosine distance and Euclidean distance. The cosine distance is:

[0034]

[0035] where, a i (t) represents the direction vector of scenario i at time t, a j (t) represents the direction vector of scenario j at time t, t = 1, 2, 3,..., 24, i, j ∈ Q, i ≠ j;

[0036] The Euclidean distance is:

[0037]

[0038] where, b i (t) represents the data value of scenario i at time t; b j(t) represents the data value of scenario j at time t;

[0039] When selecting the most representative typical scenario, the average value is used to represent the distance index of a certain scenario. The distance mean formula for scenario i is:

[0040]

[0041] Using the idea of normalization, evaluate the typical situation of the scenario and measure the most representative typical scenario. The expression is:

[0042]

[0043] where D is the set of {D1, D2,..., D Q}, min(D) and max(D) are the minimum and maximum values in the set respectively, Cos is the set of {Cos1, Cos2,..., Cos Q}, min(Cos) and max(Cos) are the minimum and maximum values in the set respectively.

[0044] Here, considering that each scenario has two distance means, the larger the cosine similarity, the better, and the smaller the Euclidean distance, the better. At the same time, the difference between the two distance values is relatively large, and it is difficult to select the most typical scenario using a simple formula. Based on the idea of normalization, the most representative typical scenario is measured.

[0045] Preferably, in step S4, the frequent item set tree algorithm is the FP-growth algorithm. The process of generating a meteorological association rule library using the frequent item set tree algorithm and generating a typical photovoltaic-load association scenario set based on the meteorological association rule library is as follows:

[0046] S41. Perform feature extraction and association analysis on the meteorological impact factors corresponding to the time of the typical scenario set;

[0047] S42. According to the correspondence between the dates of the photovoltaic typical scenario set and the meteorological data monitoring dates, each photovoltaic typical scenario has corresponding meteorological feature data;

[0048] S43. Take the date of each photovoltaic typical scenario as an item set, including the photovoltaic typical scenario and its corresponding meteorological features. Use 1 photovoltaic typical scenario label and the corresponding n meteorological feature data as an item set, so that each item set contains n + 1 items. Collect several item sets to establish an item set database;

[0049] S44. Traverse the item set database, count the frequencies of the meteorological features of all item sets in the item set database, delete the item sets that do not meet the minimum support count, and sort the item sets in descending order of frequency to obtain a frequent item list;

[0050] S45. Create an FP-tree with an empty node as the root node. Insert the item sets in the frequent item list into the FP-tree in sequence. If paths can be shared, share them and record the number of nodes at that point. After inserting the entire list, obtain the FP-tree.

[0051] S46. Mine the frequent item sets on the FP-tree. Starting from the bottom item in the frequent item list, find the corresponding conditional pattern bases upwards one by one. Recursively mine the frequent item sets using the conditional pattern bases. The obtained frequent item sets meet the requirements of the minimum support and the minimum confidence, which are the strong association rules.

[0052] S47. Use the strong association rules as the association rule base for meteorological characteristic data and photovoltaic typical scenarios, complete the association analysis of meteorological characteristic data and photovoltaic typical scenarios. Match the meteorological characteristic of the corresponding date in the load typical scenario set with the association rule base to obtain the corresponding photovoltaic typical scenario set, and finally obtain the photovoltaic-load association typical scenario set.

[0053] Here, in the FP-growth algorithm, after sorting each transaction data item in the transaction data table according to the support degree, the data items in each transaction are inserted into a tree with NULL as the root node in descending order one by one, and the support degree of the occurrence of this node is recorded at each node. At the same time, meteorological factors will affect the photovoltaic output and load output at the same moment. Therefore, it is considered to use meteorological factors as the association factors, generate photovoltaic-load association scenarios according to the meteorological association rule base, and construct typical distribution network operation scenarios with strong description and good representativeness, which can provide a more substantial scientific basis for distribution network planning under typical distribution network scenarios.

[0054] Preferably, in step S41, the meteorological influence factors are selected as the temperature, light, and atmospheric pressure values at the corresponding time every day, and the meteorological characteristic F w is extracted by the following formula:

[0055]

[0056] where, T d_max is the maximum temperature, T d_min is the minimum temperature, T d_mean is the average temperature, T d_difmean is the average first-order difference of temperature, S d_time is the solar light time S d_mean is the average solar radiation amount during the light time, S d_difmean is the average first-order difference of solar radiation amount, B d_difmean is the average first-order difference of atmospheric pressure, B d_difmax is the maximum first-order difference of atmospheric pressure.

[0057] Preferably, when performing correlation analysis on meteorological impact factors, quantiles are used for analysis, and all features included in meteorological feature F w are classified according to quantiles.

[0058] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0059] The present invention proposes a method for generating a photovoltaic-load typical scenario set based on integrated clustering and frequent item set trees. First, historical original load data and photovoltaic output data are obtained and preprocessed and classified. Multiple data sets are clustered using the integrated clustering method to obtain multiple clustering scenario sets. Then, the photovoltaic typical scenario set and the load typical scenario set are screened. Considering the influence of different typical meteorological conditions at the same time in the same area, the frequent item set tree algorithm is used to generate a meteorological association rule library, thereby establishing the correlation between the photovoltaic typical scenario and the load typical scenario. Finally, based on the meteorological association rule library, a photovoltaic-load typical association scenario set is generated. Under this scenario set, a comprehensive distribution network planning with photovoltaic is carried out, which is highly comprehensive and scientific, effectively improving the local photovoltaic consumption capacity and the stability and reliability of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It represents a schematic flow chart of the method for generating a photovoltaic-load typical scenario set based on integrated clustering and frequent item set trees proposed in Embodiment 1 of the present invention;

[0061] Figure 2 It represents a schematic flow chart of the integrated clustering method proposed in Embodiment 2 of the present invention, which successively clusters multiple data sets into different clusters to generate multiple clustering scenario sets;

[0062] Figure 3 It represents a schematic flow chart of using the frequent item set tree algorithm to generate a meteorological association rule library and generating a photovoltaic-load typical association scenario set based on the meteorological association rule library proposed in Embodiment 2 of the present invention;

[0063] Figure 4 It represents a schematic diagram after mapping the features to two dimensions by TSNE after obtaining the clustering scenario set by clustering the photovoltaic data set in spring using the HDBSDAN algorithm proposed in Embodiment 3 of the present invention;

[0064] Figure 5 It represents a schematic diagram of the changes in the CHI index and the DBI index during the process of adjusting the hyperparameters of the HDBSDAN algorithm proposed in Embodiment 3 of the present invention;

[0065] Figure 6 It represents a schematic diagram of 5 scenarios clustered in Embodiment 3 of the present invention;

[0066] Figure 7Schematic diagram showing a partial association rule library proposed in Embodiment 3 of the present invention;

[0067] Figure 8 Schematic diagram showing a typical photovoltaic scenario curve proposed in Embodiment 3 of the present invention;

[0068] Figure 9 Showing the association proposed in Embodiment 3 of the present invention Figure 8 Schematic diagram of 8 associated typical load scenario curves. Detailed implementation manners

[0069] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0070] For better illustration of this embodiment, some parts of the drawings are omitted, enlarged or reduced, which do not represent the actual size;

[0071] For those skilled in the art, it is understandable that some well-known content descriptions in the drawings may be omitted.

[0072] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0073] The description of the positional relationship in the drawings is only for illustrative purposes and should not be construed as a limitation of this patent;

[0074] Embodiment 1

[0075] As Figure 1 shown in the flowchart, this embodiment proposes a method for generating a set of typical photovoltaic-load scenarios based on integrated clustering and frequent item set trees, and the method includes the following steps:

[0076] S1. Obtain the original load data of the distribution network to be planned and the photovoltaic output data connected thereto within a certain period of time, and preprocess and classify the data to obtain multiple data sets;

[0077] S2. Use the integrated clustering method to cluster multiple data sets into different clusters in turn, thereby generating multiple clustering scenario sets;

[0078] S3. Use the comprehensive distance formula to screen out the most representative typical scenarios from each clustering scenario set, use the most representative typical scenarios as the labels of the corresponding clustering scenario sets, and finally convert all clustering scenario sets into a photovoltaic typical scenario set and a load typical scenario set;

[0079] S4. For the photovoltaic typical scenario set and the load typical scenario set, considering different meteorological influence factors at the same moment in the same region, use the frequent item set tree algorithm to generate a meteorological association rule library, and based on the meteorological association rule library, generate a photovoltaic-load typical association scenario set.

[0080] In this embodiment, for step S1, when preprocessing the acquired original load data and photovoltaic output data, for the time dates with a small amount of missing original load data and photovoltaic output data, cubic spline interpolation is used for filling, and the time dates with a large amount of missing data are discarded. During classification, the load data is classified into three data sets corresponding to weekdays, weekends, and holidays other than weekends, and the photovoltaic output data is divided into four data sets corresponding to spring, summer, autumn, and winter according to the four seasons. Thus, the situations of different electricity consumption days such as weekdays and holidays are considered, and the influence strengths of different dates on the load and photovoltaic output are considered. At the same time, it can also prevent errors caused by excessive clustering numbers in subsequent S2.

[0081] To reduce the dimension of the initial clustering data, after classification, the classified daily data sets are subjected to feature extraction, and the clustering feature vectors of the photovoltaic data set and the clustering feature vectors of the load data set are selected. In this embodiment, a total of 7 variables are taken for the clustering feature vectors of the photovoltaic data set, and the clustering feature vectors of the photovoltaic data set are expressed as:

[0082] F pv ={P d_max P d_sum P d_mean P d_std P d_difmax P d_difmin P d_difmean}

[0083] Among them, P d_max is the maximum value of the daily photovoltaic output; P d_sum is the total sum of the daily photovoltaic output throughout the day; P d_mean is the average value of the daily photovoltaic output; P d_std is the standard deviation of the photovoltaic output; P d_difmax is the maximum value of the first-order difference of the photovoltaic output sequence within a day; P d_difmin is the minimum value of the first-order difference of the photovoltaic output sequence within a day, and P d_difmean is the average value of the first-order difference of the photovoltaic output sequence within a day. Here, the photovoltaic output time is limited to within the lighting time;

[0084] The clustering feature vectors of the load data set take a total of 6 variables and are expressed as:

[0085] F load ={L d_max L d_min L d_mean L d_std L d_difmax L d_difmin L d_difmean}

[0086] Among them, L d_max is the daily maximum load; Ld_min is the daily minimum load; L d_mean is the daily average load; L d_std is the standard deviation of the daily load; L d_difmax is the maximum value of the first-order difference of the daily load; L d_difmin is the minimum value of the first-order difference of the daily load; L d_difmean is the average value of the first-order difference of the daily load, reducing the dimension of the initial clustering data. Here, the statistical time of the load data is the whole day.

[0087] Embodiment 2

[0088] This embodiment describes the process of successively clustering multiple datasets into different clusters by using the integrated clustering method to generate multiple clustering scenario sets. In this embodiment, the integrated clustering method described in step S2 is the HDBSCAN algorithm. HDBSCAN is a clustering algorithm developed by Campello, Moulavi, and Sander. It extends DBSCAN by converting DBSCAN into a hierarchical clustering algorithm and then extracting a flat clustering using a stable clustering technique. The biggest difference from the traditional DBSCAN is that HDBSCAN can handle clustering problems with different densities. The process of generating multiple clustering scenario sets can be seen in Figure 2 , specifically:

[0089] S21. Measure the distance between data points in the dataset using the mutual reachability distance, traverse and calculate the mutual reachability distance between two data points in the dataset without repetition to obtain the distance table of all data;

[0090] S22. Take any two points in each dataset as vertices, connect the vertices to get edges, and use the corresponding distance in the distance table as the weight of the edge. The entire dataset is transformed into a data distance weighted graph;

[0091] S23. Use the Prim algorithm to construct the minimum spanning tree of the data distance weighted graph to connect all distance points with the minimum distance;

[0092] S24. Sort all the edges in the minimum spanning tree in ascending order of distance, then select each edge in turn. Classify the sub-datasets included in the two sub-graphs connected by the edge into one category. After classification by the union-find set, a new category corresponding to each edge is obtained, and a clustering hierarchy is constructed;

[0093] S25. Determine the minimum number of clusters, traverse the clustering hierarchy from top to bottom. When classifying each sub-graph, judge whether the number of the two sub-datasets generated by the classification is greater than the minimum number of clusters. If so, classify them into one category. Otherwise, mark the classified category as scattered points and delete it. After traversing the entire clustering hierarchy, a compressed clustering tree with a small number of categories is obtained;

[0094] S26. Assign a class label to each compressed category in the compressed clustering tree. Traverse the compressed clustering tree from bottom to top, and determine whether the stability of the parent category of each category is greater than the sum of the stabilities of the child nodes of this category. If so, all the child nodes belong to this category, and the clustering result is output; otherwise, set the stability of this category to the sum of the stabilities of its child nodes.

[0095] In step S26, declare all the leaf nodes in the compressed clustering tree as the selected clusters. Define λ as a value that measures the persistence of a cluster. For a given cluster, define λ birth and λ death as the λ when the corresponding cluster splits and becomes its own cluster, and the λ value when the cluster splits into smaller clusters respectively; for each node in the cluster, define λ p as the λ value of the outlier of this point, which is a value between λ birth and λ death . For each cluster, calculate the stability as:

[0096]

[0097] In step S3, assume that there are Q scenarios in each clustering scenario set. The comprehensive distance formula includes cosine distance and Euclidean distance. The cosine distance is:

[0098]

[0099] where, a i (t) represents the direction vector of scenario i at time t, and a j (t) represents the direction vector of scenario j at time t, t = 1, 2, 3,..., 24, i, j ∈ Q, i ≠ j;

[0100] The Euclidean distance is:

[0101]

[0102] where, b i (t) represents the data value of scenario i at time t; b j (t) represents the data value of scenario j at time t;

[0103] Assume that there are Q scenarios in a clustering scenario set in total. Then a certain scenario undergoes (Q - 1) * 2 times of scenario comparison calculations to obtain (Q - 1) * 2 distance values. Since the most representative typical scenario needs to be selected, the average value is used to represent the distance index of a certain scenario.

[0104] When selecting the most representative typical scenario, the average value is used to represent the distance index of a certain scenario. The distance mean formula of scenario i is:

[0105]

[0106] Using the normalization idea, evaluate the typical situations of the scenarios, and measure the most representative typical scenario. The expression is as follows:

[0107]

[0108] where D is the set of {D1, D2,..., D Q}, min(D) and max(D) are the minimum and maximum values in the set respectively, Cos is the set of {Cos1, Cos2,..., Cos Q}, and min(Cos) and max(Cos) are the minimum and maximum values in the set respectively.

[0109] Considering that each scenario has two distance means, the larger the cosine similarity, the better, and the smaller the Euclidean distance, the better. At the same time, the numerical difference between the two distances is relatively large, and it is difficult to select the most typical scenario with a simple formula. Based on the normalization idea, the most representative typical scenario is measured.

[0110] In step S4, the frequent item set tree algorithm is the FP-growth algorithm. The process of generating a meteorological association rule library using the frequent item set tree algorithm and generating a typical photovoltaic-load association scenario set based on the meteorological association rule library is shown in Figure 3 The specific process is as follows:

[0111] S41. Perform feature extraction and association analysis on the meteorological impact factors corresponding to the time of the typical scenario set;

[0112] S42. According to the correspondence between the dates of the typical photovoltaic scenario set and the dates of meteorological data monitoring, make each typical photovoltaic scenario have corresponding meteorological feature data;

[0113] S43. Take the date of each typical photovoltaic scenario as an item set, including the typical photovoltaic scenario and its corresponding meteorological features. Use 1 typical photovoltaic scenario label and the corresponding n meteorological feature data as an item set, so that each item set contains n + 1 items. Collect several item sets to establish an item set database;

[0114] S44. Traverse the item set database, count the frequencies of the meteorological features of all item sets in the item set database, delete the item sets that do not meet the minimum support count, and sort the item sets in descending order of frequency to obtain a frequent item list;

[0115] S45. Create an FP-tree with an empty node as the root node. Insert the item sets of the frequent item list into the FP-tree in sequence. If they can share a path, share it and record the number of nodes at that point. After the list is inserted, an FP-tree is obtained;

[0116] S46. Mining frequent item sets on the FP-tree, finding the corresponding conditional pattern bases from the bottom items of the frequent item list upwards, and recursively mining the frequent item sets using the conditional pattern bases. The obtained frequent item sets meet the minimum support and minimum confidence requirements, which are strong association rules;

[0117] S47. Use strong association rules as the association rule base between meteorological characteristic data and typical photovoltaic scenarios to complete the association analysis between meteorological characteristic data and typical photovoltaic scenarios. Use the meteorological characteristics of the corresponding date of the load typical scenario set to match the association rule base to obtain the corresponding typical photovoltaic scenario, and finally obtain the photovoltaic-load association typical scenario.

[0118] Here, the FP-growth algorithm sorts each transaction data item in the transaction data table according to the support, and then inserts the data items in each transaction in descending order into a tree with NULL as the root node, and records the support of the node at each node. At the same time, meteorological factors will affect the photovoltaic output and load output, so meteorological factors are considered as correlation factors, and photovoltaic-load correlation scenarios are generated according to the meteorological correlation rule base, and typical distribution network operation scenarios with strong descriptiveness and good representativeness are constructed, which can provide a more substantial scientific basis for distribution network planning under typical distribution network scenarios.

[0119] In step S41, the selected photovoltaic time and load time are almost a data set of one year at the same local time. At the same time, meteorological factors will affect the photovoltaic output and load output. Therefore, the meteorological factors are considered as correlation factors, and the photovoltaic scene and the load scene are combined to obtain an existing scene that is closer to the distribution network planning.

[0120] The meteorological influencing factors are selected as temperature, light, atmospheric pressure at the corresponding time of the day, and meteorological characteristics F w Extracted by the following formula:

[0121]

[0122] Among them, T d_max is the maximum temperature, T d_min is the minimum temperature, T d_mean is the average temperature, T d_difmean is the first-order difference average of temperature, S d_time is the sunlight time S d_mean is the average solar radiation during the illumination period, S d_difmean The first-order difference average of solar radiation, B d_difmean is the first-order difference average of atmospheric pressure, B d_difmax is the maximum value of the first-order difference of atmospheric pressure.

[0123] When performing correlation analysis on meteorological impact factors, quantiles are used for analysis, and meteorological feature F w All the features it contains are classified according to quantiles.

[0124] Example 3

[0125] This example combines specific applications to more specifically illustrate the method proposed by the present invention. Suppose that using the HDBSDAN algorithm to cluster the photovoltaic dataset in spring can obtain 5 clustering scenario sets. After mapping the features to two dimensions using TSNE, we get Figure 4 As shown, different colors represent different clustering scenario sets, and different shades represent the degree of the scenario's distance from the clustering center.

[0126] To test the effectiveness of the selection of hyperparameters of the HDBSCAN algorithm, a clustering evaluation method combining CHI and DBI is proposed. The effectiveness of clustering scenarios has always been a research hotspot in the field of clustering scenarios. One of the difficulties in scenario verification indicators lies in the lack of indicators for guidance, but many clustering effectiveness indicators that have been proposed can be used:

[0127] (1) DBI index

[0128] The calculation formula of the DBI index is:

[0129]

[0130] Among them, d(X k ) and d(X j ) are the internal distances of the matrix; d(c k , c j ) is the distance between vectors. The smaller I DBI is, the better the clustering effect.

[0131] (2) CHI index

[0132] The CHI index comprehensively considers the dispersion between classes (represented by B) and the compactness within classes (represented by W), where:

[0133]

[0134]

[0135] Among them, x is the mean of all objects; w k,i represents the membership relationship of the i-th object to the k-th cluster, that is:

[0136] Then the calculation formula of the CHI index is

[0137]

[0138] It can be seen that the larger the I CHI , the better the dispersion between clusters and the compactness within clusters.

[0139] (3) Comprehensive evaluation index I DC

[0140] I DC = I CHI - I DBI

[0141] When I DC reaches the maximum value, the clustering effect is the best and the hyperparameters are selected most accurately;

[0142] See Figure 5 , in the specific embodiment, when debugging each hyperparameter, a certain hyperparameter can be fixed by using the comprehensive evaluation index, and then the result can be obtained by debugging another parameter. Fixing other parameters, when the minimum clustering number parameter value is 7, the comprehensive evaluation index reaches the highest, the clustering scenario quality is the best, and its value is I DC = 225.8 - 48.5 = 177.3.

[0143] In the following, the scene typical index is used to select the most representative typical scene from multiple different clustering scene sets. The index of this embodiment is based on the cosine distance and the Euclidean distance, and can measure the typical degree of a certain scene, that is, the degree of representativeness compared with other scenes. Through the typical scene index, the scene with the maximum typical value can be measured and used as the typical scene. In this embodiment, see Figure 6 , there are a total of 5 clustering scene sets in the figure. Among them, the dark line in the clustering scene set is the selected typical scene, which is located in the middle of the clustering scene set and is relatively smooth, and can be used as a typical scene with strong representativeness of the set to participate in the generation of the associated scene next.

[0144] To facilitate the generation of the meteorological association rule base using the FP-growth algorithm, the meteorological features need to be further analyzed and processed for association. This embodiment uses quantiles for hierarchical processing. A quantile is a point in a continuous distribution function, and this point corresponds to a probability p. If the probability 0 < p < 1, the quantile Za of the random variable X or its probability distribution is a real number that satisfies the condition p(X ≤ Za) = α. The quintile is a type of quantile in statistics, that is, all numerical values are arranged from small to large and divided into four equal parts, and the numerical values at the four dividing points are the quartiles.

[0145] All the features included in the meteorological feature F w are hierarchically processed according to the quantiles, taking T d_max as an example:

[0146] 1) The first quintile TQ1 , which is equal to the number at the 20% position after arranging all the values in the sample in ascending order;

[0147] 2) The second quintile T Q2 , which is equal to the number at the 40% position after arranging all the values in the sample in ascending order;

[0148] 3) The third quintile T Q3 , which is equal to the number at the 60% position after arranging all the values in the sample in ascending order.

[0149] 4) The fourth quintile T Q4 , which is equal to the number at the 80% position after arranging all the values in the sample in ascending order.

[0150]

[0151] FP-growth, that is, the association pattern mining using the FP-tree, is to generate item sets with a support count greater than or equal to the set minimum support count, and then obtain a strong association rule library based on the minimum confidence. The main work of association rule mining lies in mining all frequent item sets.

[0152] (1) Support

[0153] For the association rule R: X => Y, where, And I is an item set, and X and Y are associated elements. If the proportion of the item sets containing both X and Y in the item set database T is s, it is said that the support degree of the association rule R in T is s, and it can also be expressed as the probability P(X U Y), that is, the ratio of the number of times X and Y appear in T to the total number of times, as shown in the following formula:

[0154]

[0155] (2) Confidence

[0156] For the association rule R: X => Y, where, And I is an item set, and X and Y are associated elements. Then the confidence of the rule R refers to the possibility of containing Y in the item sets containing X in the item set database T, and it can be expressed by the conditional probability P(Y|X). The formula is the ratio of the number of item sets containing both X and Y to the number of item sets containing X, as shown in the following formula:

[0157]

[0158] In this embodiment, the FP-growth algorithm is used to mine the correlation between meteorological features and typical photovoltaic scenarios, and strong association rules, i.e., association rules with minimum support and minimum confidence, are set, so as to obtain the meteorological association rule base corresponding to each typical photovoltaic scenario set. As Figure 7 shown, a partial association rule base is presented. The typical load scenarios can be associated with the corresponding typical photovoltaic scenarios by matching the association rule base based on meteorological features, thereby generating photovoltaic-load association scenarios. Figure 8 Shown are the curves of the selected typical photovoltaic scenarios, Figure 9 and shown are the curves of 8 associated typical load scenarios, that is, all possible load scenarios corresponding to the typical photovoltaic scenarios are obtained. By using them together, the association scenarios composed of the two can scientifically describe the current situation of the distribution network planning scenarios, have a certain ability to summarize typical scenarios, and can provide a reasonable scientific basis for distribution network planning.

[0159] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set trees, characterized in that, The method includes the following steps: S1. Obtain the original load data of the distribution network to be planned and the photovoltaic output data connected thereto within a certain period of time, and perform preprocessing and classification on the data to obtain multiple data sets; S2. Use the integrated clustering method to cluster multiple data sets into different clusters in sequence, thereby generating multiple clustering scenario sets; S3. Use the comprehensive distance formula to screen out the most representative typical scenarios from each clustering scenario set, use the most representative typical scenarios as the labels of the corresponding clustering scenario sets, and finally convert all clustering scenario sets into a photovoltaic typical scenario set and a load typical scenario set; S4. For the photovoltaic typical scenario set and the load typical scenario set, considering different meteorological influence factors at the same moment in the same region, use the frequent item set tree algorithm to generate a meteorological association rule base, and based on the meteorological association rule base, generate a photovoltaic-load typical association scenario set; In step S4, the frequent item set tree algorithm is the FP-growth algorithm. The process of using the frequent item set tree algorithm to generate a meteorological association rule base and generating a photovoltaic-load typical association scenario set based on the meteorological association rule base is as follows: S41. Perform feature extraction and association analysis on the meteorological influence factors corresponding to the time of the typical scenario set; S42. According to the correspondence between the dates of the photovoltaic typical scenario set and the meteorological data monitoring dates, each photovoltaic typical scenario has corresponding meteorological feature data; S43. Take the date of each photovoltaic typical scenario as an item set, including the photovoltaic typical scenario and its corresponding meteorological features. Use 1 photovoltaic typical scenario label and the corresponding n meteorological feature data as the item set, so that each item set contains n + 1 items, and collect several item sets to establish an item set database; S44. Traverse the item set database, count the frequencies of the meteorological features of all item sets in the item set database, delete the item sets that do not meet the minimum support count, and sort the item sets in descending order of frequency to obtain a frequent item list; S45. Create an FP-tree with an empty node as the root node, insert the item sets of the frequent item list on the FP-tree in sequence. If the paths can be shared, share them and record the number of nodes. After the list is inserted, the FP-tree is obtained; S46. Mine the frequent item sets on the FP-tree, find the corresponding conditional pattern bases from the bottom items of the frequent item list upwards in sequence, and recursively mine the frequent item sets using the conditional pattern bases. The obtained frequent item sets meet the requirements of the minimum support and the minimum confidence, which are strong association rules; S47. Use the strong association rules as the association rule base between the meteorological feature data and the photovoltaic typical scenarios, complete the association analysis between the meteorological feature data and the photovoltaic typical scenarios, match the meteorological features corresponding to the dates of the load typical scenario set with the association rule base to obtain the corresponding photovoltaic typical scenarios, and finally obtain the photovoltaic-load associated typical scenarios.

2. The method for generating the optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 1, characterized in that In step S1, when preprocessing the obtained original load data and photovoltaic output data, for the time dates with a small amount of missing original load data and photovoltaic output data, the cubic spline interpolation method is used for filling, and the time dates with a large amount of missing data are discarded.

3. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 2, wherein When classifying, the load data is classified into three data sets corresponding to weekdays, weekends, and holidays other than weekends, and the output data of the photovoltaic is classified into four data sets corresponding to spring, summer, autumn, and winter according to the four seasons.

4. The method for generating the optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 3, characterized in that After classification, the classified daily data sets are subjected to feature extraction, and the clustering feature vectors of the photovoltaic data set and the clustering feature vectors of the load data set are selected. Among them, The clustering feature vector of the photovoltaic data set is expressed as: Among them, is the maximum value of the daily photovoltaic output; is the total sum of the daily photovoltaic output throughout the day; is the average value of the daily photovoltaic output; is the standard deviation of the photovoltaic output; is the maximum value of the first-order difference of the photovoltaic output sequence within a day; is the minimum value of the first-order difference of the photovoltaic output sequence within a day, is the average value of the first-order difference of the photovoltaic output sequence within a day; The clustering feature vector of the load data set is expressed as: Among them, is the daily maximum load; is the daily minimum load; is the daily average load; is the standard deviation of the daily load; is the maximum value of the first-order difference of the daily load; is the minimum value of the first-order difference of the daily load; is the average value of the first-order difference of the daily load.

5. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 4, wherein The integrated clustering method described in step S2 is the HDBSCAN algorithm, that is, the integrated clustering algorithm based on density and hierarchy. The process of generating multiple clustering scenario sets is as follows: S21. Use the mutual reachability distance to measure the distance between data points in the data set, and traverse and calculate the mutual reachability distance between two data points in the data set without repetition to obtain the distance table of all data; S22. Take any two points in each data set as vertices, and connect the vertices to get edges. Take the corresponding distance in the distance table as the weight of the edge, and the entire data set is transformed into a data distance weighted graph; S23. Use the Prim algorithm to construct the minimum spanning tree of the data distance weighted graph to connect all distance points with the minimum distance; S24. Sort all the edges in the minimum spanning tree in ascending order of distance, and then select each edge in turn. Classify the sub-data sets included in the two sub-graphs connected by the edge into one category. After classification by the union-find set, a new category corresponding to each edge is obtained, and a clustering hierarchy is constructed; S25. Determine the minimum number of clusters, traverse the clustering hierarchy from top to bottom. When classifying each sub-graph, judge whether the number of the two sub-data sets generated by the classification is greater than the minimum number of clusters. If so, classify them into one category; otherwise, mark the classified category as scattered points and delete them. After traversing the entire clustering hierarchy, a compressed clustering tree with a small number of categories is obtained; S26. Assign a class label to each compressed category in the compressed clustering tree. Traverse the compressed clustering tree from bottom to top, and judge whether the stability of the parent category of each category is greater than the sum of the stabilities of the child nodes of this category. If so, all child nodes belong to this category, and the clustering result is output; otherwise, set the stability of this category to the sum of the stabilities of its child nodes.

6. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set trees according to claim 5, wherein In step S26, all the compressed nodes in the compressed clustering tree are declared as selected clusters. Define λ as a value measuring the persistence of a cluster. For a given cluster, define λ birth and λ death to be the λ when the corresponding cluster splits and becomes this selected cluster, and the λ value when the cluster splits into smaller clusters respectively; for each node in the cluster, define λ p as the λ value of the outlier of this point, which is a value between λ birth and λ death For each cluster, calculate the stability as: 。 7. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 6, wherein In step S3, assume that there are a total of Q scenarios in each clustering scenario set. The comprehensive distance formula includes the cosine distance and the Euclidean distance. The cosine distance is: Among them, represents t the direction vector of the scene at time i , represents t the direction vector of the scene at time j , , , ; The Euclidean distance is: Among them, represents t the data value of the moment scene i ; represents t the data value of the moment scene j ; When selecting the most representative typical scenarios, the average value is used to represent the distance metric of a certain scenario. The i distance mean formula for the scenario is as follows: Using the idea of normalization, evaluate the typical situation of the scenario, and measure the most representative typical scenario. The expression is: Among them, is a set of , are respectively the minimum value and the maximum value in the set, is a set of , are respectively the minimum value and the maximum value in the set.

8. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 7, characterized in that, In step S41, the meteorological influence factors are selected as the temperature, light, and atmospheric pressure values at the corresponding time of each day, and the meteorological characteristics are extracted by the following formula: Among them, is the maximum temperature, is the minimum temperature, is the average temperature, is the average first-order difference of temperature, is the sunshine duration is the average solar radiation during the sunshine duration, is the average first-order difference of solar radiation, is the average first-order difference of atmospheric pressure, is the maximum first-order difference of atmospheric pressure.

9. The method for generating an optical-charge typical scenario set based on integrated clustering and frequent item set tree according to claim 8, wherein When performing correlation analysis on meteorological impact factors, quantiles are used for analysis, and all features contained in the meteorological characteristics are classified according to quantiles.