A power distribution operation scene construction method and device based on load clustering

By extracting features from historical load data of the distribution network and generating synthetic data using a generative adversarial network model, and then performing secondary cluster analysis, the problems of insufficient initial cluster center selection and dynamic characteristics consideration in the construction of distribution network operation scenarios are solved, thereby improving the accuracy of scenario construction and data support capabilities.

CN120995150BActive Publication Date: 2026-02-06STATE GRID JIANGSU ECONOMIC RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511529093.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-06
Estimated Expiration
2045-10-24

AI Technical Summary

Technical Problem

Existing technologies struggle to select suitable initial cluster centers in the construction of power distribution network operation scenarios and fail to fully consider the overall dynamic characteristics of curves, resulting in insufficient accuracy in the construction of power distribution operation scenarios and an inability to meet the demands for high-quality and high-reliability power supply.

Method used

By acquiring multiple sets of historical load data from power distribution operations, feature extraction and outlier correction are performed. Combined with a generative adversarial network model, synthetic load data is generated, and secondary clustering analysis is conducted to determine the target cluster center matrix, thereby mining load characteristics and constructing power distribution operation scenarios.

Benefits of technology

It improves the accuracy of power distribution operation scenario construction, provides data foundation support for high-quality and high-reliability power supply, and provides data support for demand-side response policies and high-precision load forecasting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995150B_ABST
    Figure CN120995150B_ABST
Patent Text Reader

Abstract

The application discloses a power distribution operation scene construction method and device based on load clustering, and the method comprises the following steps: acquiring multiple groups of historical load data; performing feature extraction on each group of historical load data respectively to obtain a load feature matrix; combining a preset initial cluster number to determine an initial cluster center set from the load feature matrix; performing clustering analysis based on each initial cluster center in the initial cluster center set to obtain a historical clustering center matrix; acquiring a synthetic clustering center matrix, fusing the historical clustering center matrix, and giving a mixed clustering center matrix; performing secondary clustering analysis based on the mixed clustering center matrix to obtain a target clustering center matrix, so as to complete the construction of the power distribution operation scene. Through clustering of historical load data of a power distribution network, load characteristic analysis is realized, a load feature matrix is obtained through feature extraction, and potential change rules are mined based on the load feature matrix, so that the target clustering center matrix is given, and the accuracy of the construction of the power distribution operation scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of power system analysis, specifically relating to a method and apparatus for constructing distribution operation scenarios based on load clustering. Background Technology

[0002] Classifying distribution network operation scenarios and extracting typical scenarios is a crucial foundation for distribution network operation and planning. The distribution network is the hub of electrical energy, directly serving tens of thousands of end users, including residential, commercial, and industrial users, whose electricity demands vary significantly. How to uncover the potential patterns of change in grid operation and provide customized, high-quality, and highly reliable power services to distribution network users is a problem that needs to be solved.

[0003] Patent CN115526267A provides a method and apparatus for extracting power distribution network operation scenarios, applicable to photovoltaic power generation or other fields. The method includes: using a predetermined initial clustering number as the iterative clustering number; increasing the iterative clustering number to perform iterative operations to obtain clustering results and iterative losses at each iterative clustering number, until the iterative clustering number reaches a preset maximum clustering number; calculating initial cluster centers, where the number of initial cluster centers equals the iterative clustering number; calculating clustering results and iterative losses based on the initial cluster centers; calculating the clustering loss at the corresponding iterative clustering number based on the iterative losses generated by the clustering results at each iterative clustering number; and selecting the clustering result corresponding to the minimum clustering loss as the scenario extraction result. This method achieves automatic probability calculation to select suitable initial cluster centers and fully considers the overall dynamic characteristics of the curve and the correlation between photovoltaic power generation and load power consumption during clustering.

[0004] In existing technologies, typical operating scenarios of power distribution networks are mainly obtained through clustering. Commonly used clustering methods include k-means clustering, fuzzy C-means clustering, and hierarchical clustering. However, when performing clustering, existing technologies make it difficult to select suitable initial cluster centers. Furthermore, when calculating the similarity between sample points and cluster centers, they only consider the distribution characteristics of the curves and do not take into account the overall dynamic characteristics of the curves.

[0005] Therefore, how to uncover the potential patterns of change in power distribution operation, improve the accuracy of power distribution operation scenario construction, and provide a data foundation for customized power, high-quality power supply, and high-reliability power supply services for power users in the power distribution network is a problem that needs to be solved. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, this invention provides a method and apparatus for constructing power distribution operation scenarios based on load clustering. The method includes: acquiring multiple sets of historical load data for power distribution operation; extracting features from each set of historical load data to obtain a load feature matrix; determining an initial cluster center set from the load feature matrix based on a preset initial cluster number; performing cluster analysis on each initial cluster center in the initial cluster center set to obtain a historical cluster center matrix and assigning a first label; acquiring a synthetic cluster center matrix and fusing it with the historical cluster center matrix to obtain a hybrid cluster center matrix, wherein the synthetic cluster center matrix is ​​obtained from synthetic load data based on historical load data; and performing secondary cluster analysis on the hybrid cluster center matrix to obtain a target cluster center matrix, thereby completing the construction of the power distribution operation scenario, wherein the target cluster center matrix includes features of various scenarios during power distribution operation. By clustering historical load data of the distribution network, load characteristic analysis is achieved. Through feature extraction, common intrinsic characteristics are found to obtain a load characteristic matrix. Based on this, potential change patterns are explored, and a target cluster center matrix is ​​given. This enables the distribution network to provide customized power, high-quality power supply, and high-reliability power supply services to power users. It can also provide data support for the formulation of demand-side response policies and high-precision load forecasting.

[0007] In a first aspect, the present invention provides a method for constructing power distribution operation scenarios based on load clustering, specifically including the following steps:

[0008] Acquire multiple sets of historical load data for power distribution operation;

[0009] Feature extraction was performed on each set of historical load data to obtain a load feature matrix;

[0010] Based on the preset initial number of clusters, the initial cluster center set is determined from the load feature matrix;

[0011] Based on the initial cluster centers in the initial cluster center set, cluster analysis is performed to obtain the historical cluster center matrix and give the first label;

[0012] Obtain the synthetic cluster center matrix and merge it with the historical cluster center matrix to give the mixed cluster center matrix. The synthetic cluster center matrix is ​​obtained based on the synthetic load data given by the historical load data.

[0013] Based on the hybrid cluster center matrix, a secondary clustering analysis is performed to obtain the target cluster center matrix, which completes the construction of power distribution operation scenarios. The target cluster center matrix includes the characteristics of various scenarios in the power distribution operation process.

[0014] Furthermore, multiple sets of historical load data for power distribution operation are obtained through the following steps:

[0015] Acquire multiple sets of initial historical load data, wherein each set of initial historical load data includes the historical active power of transformers and / or transmission lines within a unit time period;

[0016] Based on the analysis of each set of initial historical load data, the median power corresponding to each set of initial historical load data is given.

[0017] By combining the median power, power outliers in each set of initial historical load data are identified and corrected.

[0018] By combining the difference method with the normal power values ​​in each group of initial historical load data, the missing power values ​​in each group of initial historical load data are filled in to obtain the historical load data for each group.

[0019] Furthermore, by combining the median power, power outliers in each set of initial historical load data are identified and corrected, specifically including:

[0020] Based on the median power, calculate the difference between the historical active power and the median power in each set of initial historical load data to obtain the absolute deviation value;

[0021] Based on the absolute deviation value corresponding to each set of initial historical load data, the median of the absolute deviation is given;

[0022] Based on the median absolute deviation and the number of data points in each set of initial historical load data, the corresponding reasonable data range is obtained;

[0023] Based on a reasonable data range, the initial historical load data is filtered to identify power anomalies;

[0024] By combining the median power, median absolute deviation, and number of data points, the correction of power outliers is completed.

[0025] Furthermore, by combining the interpolation method with the normal power values ​​in each group of initial historical load data, the missing power values ​​in each group of initial historical load data are filled in to obtain the historical load data for each group, specifically including:

[0026] Determine the location of missing power values ​​in each set of initial historical load data;

[0027] Based on the missing location, select the corresponding normal power value from the initial historical load data;

[0028] Using the Lagrange interpolation method and combining it with the corresponding normal power values, power completion data is given, and historical load data for each group is obtained.

[0029] Furthermore, feature extraction is performed on each set of historical load data to obtain a load feature matrix, specifically including:

[0030] Based on the historical load data of each group, a load sample matrix is ​​constructed;

[0031] The load sample matrix is ​​preprocessed to obtain the standard load sample matrix and the corresponding correlation coefficient matrix;

[0032] Based on the correlation coefficient matrix and the identity matrix, a load characteristic equation is constructed, and multiple load characteristic values ​​corresponding to the load characteristic equation are given;

[0033] By combining multiple load characteristic values ​​and the unit load characteristic vectors corresponding to the load characteristic values, the cumulative contribution rate corresponding to each group of historical load data is given, and the characteristic load matrix is ​​obtained.

[0034] The load characteristic matrix is ​​obtained by fusing the characteristic load matrix and the standard load sample matrix.

[0035] Furthermore, based on the preset initial number of clusters, the initial cluster center set is determined from the load feature matrix, specifically including:

[0036] Randomly select any load feature from the load feature matrix and give the center of the first cluster;

[0037] Based on the distance between the first cluster center and each load feature in the load feature matrix, analyze the probability that each load feature is the next cluster center, and give the second cluster center;

[0038] Repeat the above steps until you obtain the cluster centers of all the initial clusters, thus giving the initial cluster center set.

[0039] Further, the cluster center matrix is ​​synthesized through the following steps:

[0040] Based on the preset label information and combined with the pre-built generative adversarial network model, synthetic payload data is generated.

[0041] The composite load data is projected onto the space of the load feature matrix to obtain the composite feature matrix;

[0042] Pre-clustering is performed on the synthetic feature matrix to obtain the synthetic cluster center matrix, and a second label is given.

[0043] Furthermore, the construction of the generative adversarial network model is determined through the following steps:

[0044] Initialize the model parameters of the generative adversarial network model, which include the number of neurons in each layer, loss function, penalty coefficient, and generator parameter update step size;

[0045] The conditional labels corresponding to the historical load data are concatenated with the noise vector through a fully connected layer in a generative adversarial network model.

[0046] In the generative adversarial network model, the generator outputs reference load data based on historical load data and the corresponding condition labels.

[0047] In the generative adversarial network model, the discriminator outputs a gradient penalty and updates the discriminator parameters based on historical load data and corresponding reference load data, combined with a penalty coefficient.

[0048] The generator parameters are updated once every preset number of times the discriminator parameters are updated.

[0049] Based on the similarity between the reference load data and historical load data, the generative adversarial network model is trained until convergence.

[0050] Furthermore, based on the mixed cluster center matrix, a secondary clustering analysis is performed to obtain the target cluster center matrix, which specifically includes:

[0051] Obtain the historical within-cluster standard deviation of each cluster in the historical cluster center matrix and the composite within-cluster standard deviation of each cluster in the composite cluster center matrix;

[0052] The splitting threshold adjustment function is determined based on the historical intra-class standard deviation and the composite intra-class standard deviation;

[0053] Obtain the historical cluster spacing of each cluster in the historical cluster center matrix and the synthetic cluster spacing of each cluster in the synthetic cluster center matrix;

[0054] The merging threshold adjustment function is determined based on historical cluster spacing and synthetic cluster spacing;

[0055] Based on the hybrid cluster center matrix, and combining the split threshold adjustment function and the merge threshold adjustment function, the division of various clusters in the secondary clustering process is adjusted, and the target cluster center matrix is ​​given.

[0056] Furthermore, the splitting threshold adjustment function is determined through the following steps:

[0057] Based on the maximum value of the historical within-class standard deviation and the maximum value of the composite within-class standard deviation, calculate the difference between the maximum value of the historical within-class standard deviation and the maximum value of the composite within-class standard deviation to obtain the within-class difference.

[0058] Based on the pre-set split adjustment coefficient, the intra-class difference is adjusted to provide a split adjustment term;

[0059] Based on the difference between the current splitting threshold and the splitting adjustment term, a splitting threshold adjustment function is given.

[0060] Furthermore, the threshold adjustment function is merged, specifically determined through the following steps:

[0061] Based on the minimum historical cluster spacing and the minimum synthetic cluster spacing, the difference between the minimum historical cluster spacing and the minimum synthetic cluster spacing is calculated to obtain the cluster spacing difference.

[0062] Based on the pre-set merging adjustment coefficient, the cluster spacing difference is adjusted, and the merging adjustment item is given;

[0063] Based on the difference between the current merging threshold and the merging adjustment term, a merging threshold adjustment function is given.

[0064] Furthermore, based on the hybrid cluster center matrix, and combining the splitting threshold adjustment function and the merging threshold adjustment function, the partitioning of various clusters in the secondary clustering process is adjusted to give the target cluster center matrix, which specifically includes:

[0065] Based on each cluster center in the mixed cluster center matrix and the cluster corresponding to each cluster center, give the maximum intra-cluster standard deviation of each cluster and the inter-cluster distance between each cluster;

[0066] If the maximum intra-class standard deviation is greater than the current classification threshold, the cluster corresponding to the maximum intra-class standard deviation will be split into two clusters.

[0067] If the cluster spacing is less than the current merging threshold and the difference in the number of samples is less than the difference in the number of samples threshold, the two clusters corresponding to the cluster spacing will be merged.

[0068] The current split threshold and the current merge threshold are updated based on the split threshold adjustment function and the merge threshold adjustment function;

[0069] Based on the splitting and / or merging scenarios, the individual cluster centers are redefined;

[0070] If the number of clusters changes or the number of iterations reaches the corresponding threshold, the second-order clustering is completed, and the target cluster center matrix is ​​given.

[0071] Furthermore, the cluster splits into splits along the principal component axes.

[0072] Furthermore, the sample difference number is the ratio of the difference in the number of clusters before and after merging and / or splitting to the number of clusters before merging and / or splitting.

[0073] Secondly, the present invention also provides a power distribution operation scenario construction device based on load clustering, employing a power distribution operation scenario construction method based on load clustering as described above, comprising:

[0074] The data acquisition module is used to acquire multiple sets of historical load data for power distribution operation;

[0075] The feature extraction module is used to extract features from each group of historical load data to obtain a load feature matrix;

[0076] The initial cluster determination module is used to determine the initial cluster center set from the load feature matrix by combining the preset initial cluster number;

[0077] The first clustering module is used to perform cluster analysis based on each initial cluster center in the initial cluster center set, obtain the historical cluster center matrix, and give the first label;

[0078] The hybrid cluster center determination module is used to obtain the synthetic cluster center matrix and merge it with the historical cluster center matrix to give the hybrid cluster center matrix. The synthetic cluster center matrix is ​​obtained based on the synthetic load data given by the historical load data.

[0079] The second clustering module is used to perform secondary clustering analysis based on the hybrid clustering center matrix to obtain the target clustering center matrix, so as to complete the construction of the power distribution operation scenario. The target clustering center matrix includes the characteristics of various scenarios in the power distribution operation process.

[0080] The present invention provides a method and apparatus for constructing power distribution operation scenarios based on load clustering, which has at least the following beneficial effects:

[0081] (1) By clustering the historical load data of the distribution network, load characteristic analysis is realized. Through feature extraction, its common inherent characteristics are found, and the load characteristic matrix is ​​obtained. Based on this, its potential change pattern is explored, and the target cluster center matrix is ​​given, which improves the accuracy of the construction of the distribution operation scenario and provides data support for the formulation of demand-side response policies and high-precision load forecasting.

[0082] (2) By adopting a generative adversarial network model, synthetic load data is generated, which improves the authenticity of synthetic load data and expands the data range of cluster analysis, so that the target cluster center matrix obtained by integrating the synthetic cluster center matrix covers more operating scenarios. Attached Figure Description

[0083] Figure 1 A flowchart illustrating the method for constructing a power distribution operation scenario based on load clustering, as provided in an embodiment of the present invention;

[0084] Figure 2 A flowchart for obtaining the load characteristic matrix provided in an embodiment of the present invention;

[0085] Figure 3 A flowchart for determining the initial cluster center set provided in an embodiment of the present invention;

[0086] Figure 4 This is a flowchart for determining the synthetic cluster center matrix provided in an embodiment of the present invention;

[0087] Figure 5 A flowchart for constructing a generative adversarial network model provided in an embodiment of the present invention;

[0088] Figure 6 A model architecture diagram of the adversarial network model provided in the embodiments of the present invention;

[0089] Figure 7 This is a flowchart for obtaining the target cluster center matrix provided in an embodiment of the present invention;

[0090] Figure 8 This is a flowchart of cluster update in secondary clustering provided in an embodiment of the present invention;

[0091] Figure 9 The structural block diagram of the power distribution operation scenario construction device based on load clustering provided in the embodiments of the present invention.

[0092] Among them, 201 is the data acquisition module; 202 is the feature extraction module; 203 is the initial cluster determination module; 204 is the first clustering module; 205 is the mixture center determination module; and 206 is the second clustering module. Detailed Implementation

[0093] To better understand the above technical solutions, a detailed description of the solutions will be provided below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0094] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0095] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that an article or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device that includes said element.

[0096] The problem of power distribution operation optimization exists in multiple research areas, such as distribution network planning, control, and evaluation, and forms the foundation of power system analysis. In the operation and development of power systems, the power grid is particularly vulnerable due to its wide geographical distribution and direct susceptibility to natural environmental influences. In particular, the high sensitivity of power equipment to meteorological conditions such as temperature, humidity, snow, and lightning means that meteorological factors have a direct impact on the safe and stable operation of power grid equipment. Furthermore, meteorological resources such as wind speed and solar irradiance exhibit extremely strong random fluctuations, making it very difficult to predict their trends.

[0097] With the integration of large-scale renewable energy power generation facilities into the grid, the uncertainty of renewable energy output and its absorption have become key constraints. In addition, meteorological conditions affect people's perception of environmental comfort, which in turn influences electricity consumption behavior; different meteorological factors have varying sensitivities to electricity load at different times.

[0098] This invention provides a method for constructing a power distribution operation scenario based on load clustering. The method includes acquiring multiple sets of historical load data for power distribution operation; extracting features from each set of historical load data to obtain a load feature matrix; determining an initial cluster center set from the load feature matrix based on a preset initial number of clusters; performing cluster analysis on each initial cluster center in the initial cluster center set to obtain a historical cluster center matrix and assigning a first label; acquiring a synthetic cluster center matrix and fusing it with the historical cluster center matrix to obtain a hybrid cluster center matrix, wherein the synthetic cluster center matrix is ​​obtained based on synthetic load data given from the historical load data; and performing secondary cluster analysis on the hybrid cluster center matrix to obtain a target cluster center matrix, thereby completing the construction of the power distribution operation scenario, wherein the target cluster center matrix includes features of various scenarios during power distribution operation.

[0099] This invention clusters historical load data of the distribution network to achieve load characteristic analysis. Through feature extraction, it finds common inherent characteristics to obtain a load feature matrix. Based on this, it explores potential change patterns and provides a target cluster center matrix. This enables the distribution network to provide customized power, high-quality power supply, and high-reliability power supply services to power users. It can also provide data support for the formulation of demand-side response policies and high-precision load forecasting.

[0100] like Figure 1 As shown in the figure, this embodiment of the invention provides a method for constructing a power distribution operation scenario based on load clustering, and the specific steps are as follows:

[0101] S101: Obtain multiple sets of historical load data for power distribution operation.

[0102] Furthermore, multiple sets of historical load data for power distribution operation are obtained through the following steps:

[0103] Acquire multiple sets of initial historical load data, wherein each set of initial historical load data includes the historical active power of transformers and / or transmission lines within a unit time period;

[0104] Based on the analysis of each set of initial historical load data, the median power corresponding to each set of initial historical load data is given.

[0105] By combining the median power, power outliers in each set of initial historical load data are identified and corrected.

[0106] By combining the difference method with the normal power values ​​in each group of initial historical load data, the missing power values ​​in each group of initial historical load data are filled in to obtain the historical load data for each group.

[0107] Furthermore, by combining the median power, power outliers in each set of initial historical load data are identified and corrected, specifically including:

[0108] Based on the median power, calculate the difference between the historical active power and the median power in each set of initial historical load data to obtain the absolute deviation value;

[0109] Based on the absolute deviation value corresponding to each set of initial historical load data, the median of the absolute deviation is given;

[0110] Based on the median absolute deviation and the number of data points in each set of initial historical load data, the corresponding reasonable data range is obtained;

[0111] Based on a reasonable data range, the initial historical load data is filtered to identify power anomalies;

[0112] By combining the median power, median absolute deviation, and number of data points, the correction of power outliers is completed.

[0113] In one specific implementation, each set of initial historical load data represents the historical active power of transformers and / or transmission lines within a unit of time. In a specific example, the unit of time is a day. For instance, a set of initial historical data might consist of active power collected within a day at a sampling frequency of 15 minutes per sampling, i.e., 96 time points per day. The data format is a time series, including timestamps and the active power value at the corresponding timestamp. It is understood that each set of initial historical data represents daily load data.

[0114] The median absolute deviation (MAD) method is used to identify power anomalies in each set of initial historical data. MAD is a method that detects anomalies by calculating the sum of the distances between each sampled value and the average value. For daily load data of a certain day, i.e., a set of initial historical data X={X1,X2,…,X…},…n The specific process is as follows:

[0115] First, calculate the power median of the daily load data. Then, the absolute deviation of each sampling point from the median in the daily load data is calculated. And give the median absolute deviation. The reasonable data range is a preset data range that can be set according to different scenarios and values. For example, by combining the number of data points n in an initial set of historical data, the reasonable data range corresponding to that initial set of historical data can be obtained. .

[0116] Based on a reasonable data range, data outside this range from the initial historical data are filtered out to obtain power anomalies. The correction for these power anomalies is specifically expressed as follows:

[0117]

[0118] in, X is the data obtained after correcting the i-th power anomaly. i For the i-th power anomaly, denoted as median power, n as the number of data points, and MAD as median absolute deviation.

[0119] Furthermore, by combining the interpolation method with the normal power values ​​in each group of initial historical load data, the missing power values ​​in each group of initial historical load data are filled in to obtain the historical load data for each group, specifically including:

[0120] Determine the location of missing power values ​​in each set of initial historical load data;

[0121] Based on the missing location, select the corresponding normal power value from the initial historical load data;

[0122] Using the Lagrange interpolation method and combining it with the corresponding normal power values, power completion data is given, and historical load data for each group is obtained.

[0123] In one specific implementation, Lagrange interpolation (LI) is used to complete the missing power values ​​in each set of initial historical load data. Lagrange interpolation is a data recovery algorithm that utilizes the strong continuity and autocorrelation between sampling points on the same day, assuming that a small number of consecutive load data points exhibit a continuous variation pattern. It primarily uses several load points before and after the missing data points to calculate the Lagrange interpolation formula and obtain the corresponding polynomial. The function value corresponding to the missing point in the polynomial is then used as the repaired value for that missing data.

[0124] Missing data locations are categorized as either first / last missing or middle missing. If the missing data is at the first / last position, several normal load data points closest to the first / last position are selected, and Lagrange interpolation is performed. If the missing data is in the middle, several normal load data points before and after that point are selected, and Lagrange interpolation is then performed. For example, given a set of initial historical load data (2,3,4,5,6,7,8,9,10), where the first data point is missing (first missing in a first / last missing data point), (2,3,4,5) is selected for Lagrange interpolation to determine the corresponding power completion data. Similarly, given a set of initial historical load data (1,2,3,4,5,6,7,8,9), where the last data point is missing (last missing in a first / last missing data point), (6,7,8,9) is selected for Lagrange interpolation to determine the corresponding power completion data. In another specific example, there is a set of initial historical load data, including 10 data points (1,2,3,4,6,7,8,9,10). The last data point is missing, which is considered a missing data point in the middle. Therefore, (3,4,6,7) is selected to calculate the Lagrange interpolation formula to determine the corresponding power completion data.

[0125] By correcting and supplementing the initial historical load data to ensure its accuracy, completeness, consistency, and timeliness, a reliable data foundation is provided for subsequent feature extraction, cluster analysis, and other processes, thereby ensuring the rationality and effectiveness of the power distribution operation scenario construction.

[0126] S102: Extract features from each set of historical load data to obtain the load feature matrix.

[0127] Reference Figure 2 The load characteristic matrix is ​​obtained, specifically including:

[0128] Based on the historical load data of each group, a load sample matrix is ​​constructed;

[0129] The load sample matrix is ​​preprocessed to obtain the standard load sample matrix and the corresponding correlation coefficient matrix;

[0130] Based on the correlation coefficient matrix and the identity matrix, a load characteristic equation is constructed, and multiple load characteristic values ​​corresponding to the load characteristic equation are given;

[0131] By combining multiple load characteristic values ​​and the unit load characteristic vectors corresponding to the load characteristic values, the cumulative contribution rate corresponding to each group of historical load data is given, and the characteristic load matrix is ​​obtained.

[0132] The load characteristic matrix is ​​obtained by fusing the characteristic load matrix and the standard load sample matrix.

[0133] In one specific implementation, there are n sets of historical load data (i.e., the historical active power of transformers and transmission lines at different unit times), and each set of historical load data is a p-dimensional variable. The corresponding load sample matrix X is then obtained as follows:

[0134]

[0135] Where, x ij Let be the value of the j-th dimension in the i-th group of historical load data.

[0136] The load sample matrix is ​​preprocessed to obtain the standard load sample matrix and the corresponding correlation coefficient matrix, as shown below:

[0137]

[0138] in, This is the value obtained by standardizing the j-th dimension of the i-th group of historical load data. Let s be the average of all data in the j-th column of the load sample matrix X. j Let x be the standard deviation of all data in the j-th column of the loading sample matrix X. kj This refers to the data in the k-th row of the historical load data in column j.

[0139] The correlation coefficient matrix R of the standardized matrix is ​​specifically represented as:

[0140]

[0141] Each correlation coefficient r in the correlation coefficient matrix R ij Let x represent the correlation between the features in the i-th row and the j-th column of the load sample matrix. ki This represents the k-th dimension of the historical load data in the i-th row.

[0142] Based on the correlation coefficient matrix and the identity matrix, a load characteristic equation is constructed, specifically: |λI-R|=0. Multiple load characteristic values ​​are then obtained from this equation, specifically represented as (λ1, λ2, ..., λ...). p The load characteristic values ​​are then set according to λ1≥λ2≥…≥λ p Arranged in order of ≥0, where I is the identity matrix and λ is the load eigenvalue vector.

[0143] In a specific example, the Jacobi method can be used to solve the load characteristic equation, obtaining multiple load characteristic values ​​and their relationship with the load characteristic value λ. i The corresponding unit load characteristic vector is represented as: e i =[e i1 ,e i2,…,e ip ] T ,i=1,2,…,p.

[0144] Combining multiple load characteristic values ​​and the corresponding unit load characteristic vectors, the cumulative contribution rate for each set of historical load data is given, resulting in the characteristic load matrix L, specifically represented as follows:

[0145]

[0146] Among them, l ij Let λ be the characteristic load factor corresponding to the load eigenvalue in the i-th row and j-th column of the load sample matrix X. i For the i-th load characteristic value, e ij For λ i The j-th value in the corresponding unit load characteristic vector.

[0147] Understandably, load characteristic values ​​are generally selected based on the cumulative contribution rate of each load characteristic value. In this example, load characteristic values ​​with a cumulative contribution rate of 85% to 95% are selected, corresponding to m principal components.

[0148] Wherein, the contribution rate α of the i-th load characteristic value i for:

[0149]

[0150] The cumulative contribution rate β of the i-th load characteristic value i for:

[0151]

[0152] By fusing the load matrix and the standard load sample matrix, the load characteristic matrix Z is obtained. hist Specifically, it is expressed as:

[0153]

[0154] in, To The data after dimensionality reduction, where L is the feature load matrix. for The m-th principal component, where N is the number of historical samples.

[0155] In the embodiments provided by this invention, load characteristics can be determined based on the specific details of historical load data. For example, load characteristics can be determined by filtering from factors such as daily load peak value, load duration, and load fluctuation intensity.

[0156] Principal component analysis was used to perform dimensionality reduction on the preprocessed historical load data to obtain the load feature matrix. The aim was to reduce the computational complexity of subsequent cluster analysis by extracting the main features of the historical load data.

[0157] S103: Determine the initial cluster center set from the load feature matrix based on the preset initial cluster number.

[0158] Reference Figure 3 Determine the initial cluster centroid set, specifically including:

[0159] Randomly select any load feature from the load feature matrix and give the center of the first cluster;

[0160] Based on the distance between the first cluster center and each load feature in the load feature matrix, analyze the probability that each load feature is the next cluster center, and give the second cluster center;

[0161] Repeat the above steps until you obtain the cluster centers of all the initial clusters, thus giving the initial cluster center set.

[0162] In one specific implementation, the feature load matrix after feature extraction is used as the cluster in the clustering model. In this example, the K-means++ method is used to obtain the initial cluster centers, and then the iterative self-organizing data analysis technique (ISODATA) is used for clustering. The specific steps are as follows:

[0163] First, a load feature is randomly selected from the load feature matrix. This feature could be the daily load peak, load duration, or load fluctuation intensity, and it will be used as the first cluster center. Then, the distance D(x) between each load feature and the first cluster center is calculated, along with the probability P(x) of each load feature being selected as the next cluster center. Specifically, the probability of each load feature being the next cluster center is expressed as follows:

[0164]

[0165] Where χ represents the sample set corresponding to the load feature matrix. In this example, D(x) is the Euclidean distance between each load feature and the center of the first cluster.

[0166] The load feature corresponding to the maximum probability is taken as the next cluster center. The above process is repeated until the cluster center corresponding to the initial number of clusters is selected, thus obtaining the initial cluster center set.

[0167] S104: Based on each initial cluster center in the initial cluster center set, perform cluster analysis to obtain the historical cluster center matrix and give the first label.

[0168] In the embodiments provided by this invention, based on the initial cluster center set, the ISODATA clustering algorithm is used for cluster analysis to obtain the historical cluster center matrix C. hist ∈R K×m .

[0169] It's important to understand that each cluster center in the historical cluster center matrix corresponds to a power distribution operation mode, which can also be understood as a power distribution operation scenario. Different cluster center morphological features can be labeled with corresponding first tags. The specific correspondence between the first tag and the cluster center morphological features—that is, the first tag being determined by the morphology of each cluster center in the historical cluster center matrix—can be as follows in a particular implementation scenario:

[0170] Based on the daily load curve (96 time points) of each cluster center, its morphological characteristics, including peak values, valley values, and fluctuations, are extracted. Typical time periods (peak, flat, and valley) are divided according to power distribution operation experience. For example, 9:00-12:00 and 15:00-18:00 are peak periods, 23:00-5:00 are valley periods, and other times are flat periods.

[0171] In a specific example, if the overlap rate between the peak load period and the "peak period" of a cluster is ≥80%, then the first label is marked as "peak load"; if the overlap rate between the valley load period and the "valley period" is ≥80%, then the second label is marked as "valley load".

[0172] It is understandable that the first label will not change as the data is processed. Therefore, the first label can be labeled in the initial historical data, or after obtaining the historical load data and load feature matrix, or after obtaining the historical cluster center matrix. There are no restrictions on this.

[0173] S105: Obtain the synthetic cluster center matrix and merge it with the historical cluster center matrix to give the mixed cluster center matrix.

[0174] The synthetic cluster center matrix is ​​obtained based on synthetic load data derived from historical load data. Specifically, refer to... Figure 4 The cluster center matrix is ​​synthesized and determined through the following steps:

[0175] Based on the preset label information and combined with the pre-built generative adversarial network model, synthetic payload data is generated.

[0176] The composite load data is projected onto the space of the load feature matrix to obtain the composite feature matrix;

[0177] Pre-clustering is performed on the synthetic feature matrix to obtain the synthetic cluster center matrix.

[0178] After obtaining the synthetic cluster center matrix, the historical cluster center matrix and the synthetic cluster center matrix are merged to give the hybrid cluster center matrix. .

[0179] The historical cluster center matrix and the synthetic cluster center matrix are horizontally concatenated, that is, the columns of the historical cluster center matrix are appended to the left side of the synthetic cluster center matrix to form a new matrix, namely the hybrid cluster center matrix.

[0180] Furthermore, referring to Figure 5 The construction of the generative adversarial network model is determined through the following steps:

[0181] Initialize the model parameters of the generative adversarial network model, which include the number of neurons in each layer, loss function, penalty coefficient, and generator parameter update step size;

[0182] The conditional labels corresponding to the historical load data are concatenated with the noise vector through a fully connected layer in a generative adversarial network model.

[0183] In the generative adversarial network model, the generator outputs reference load data based on historical load data and the corresponding condition labels.

[0184] In the generative adversarial network model, the discriminator outputs a gradient penalty and updates the discriminator parameters based on historical load data and corresponding reference load data, combined with a penalty coefficient.

[0185] The generator parameters are updated once every preset number of times the discriminator parameters are updated.

[0186] Based on the similarity between the reference load data and historical load data, the generative adversarial network model is trained until convergence.

[0187] The model architecture diagram of the adversarial network model is as follows: Figure 6As shown. In one specific implementation, firstly, the model parameters of the Generative Adversarial Network (GAN) model are initialized. These parameters include the number of neurons in each layer, activation functions, loss functions, gradient penalty terms, penalty coefficients, and generator parameter update step sizes. The GAN model includes a generator and a discriminator. Then, conditional labels (including weather and season) are concatenated with a noise vector through a fully connected layer. Real load data x and corresponding conditional labels y are obtained from historical load data. The generator G in the GAN model is used to generate fake data, i.e., reference load data G(z|y), where z is random noise. The discriminator in the GAN model is used to distinguish between real historical load data and fake reference load data D(x|y) and D(G(z|y)), calculating the gradient penalty term to ensure the discriminator satisfies 1-Lipschitz continuity. The discriminator parameters are updated once each reference load data is generated. After this, the generator parameters are updated once. The generator and discriminator in the generative adversarial network model are trained through the above process until the loss function converges, completing the training.

[0188] Understandably, after training is complete, the trained generative adversarial network model is used to generate daily operating load data for potential scenarios based on preset label information, thus obtaining synthetic load data X. gen ∈R M×96 Where M is the number of generated samples. The synthetic load data is projected onto the space of the load feature matrix to obtain the synthetic feature matrix Z. gen ∈R M×m Pre-clustering is performed on the synthetic feature matrix to obtain the synthetic cluster center matrix. The number of clusters is K ' .

[0189] The tag information corresponds to the first and second tags and is used to restrict the power distribution operation scenario corresponding to the obtained composite load data. For example, if the tag information is "typhoon high load scenario," the data requirement for the typhoon high load scenario is a scenario where the wind speed exceeds level 10 and the load exceeds a preset first load threshold. That is, the composite load data generated by the typhoon high load scenario must meet the data requirement of "wind speed exceeding level 10 and load exceeding a preset first load threshold." Another example is "extreme low temperature load scenario," which requires the data requirement of a scenario where the temperature is below -10℃ and the load fluctuation rate exceeds 30%. The tag information can also be for scenarios such as high temperature high load scenario, heavy rain high load scenario, nighttime low load scenario, industrial area high load scenario, and residential area low load scenario.

[0190] Projecting the synthetic load data into the space of the load feature matrix specifically includes:

[0191] Assign an initial weight to each feature in the load feature matrix, where the initial weight represents the importance of each feature in the synthetic load data.

[0192] The initial weights are adjusted by optimization methods (such as least squares) to minimize the error when the composite load data is represented by a combination of the load feature matrix and the weights.

[0193] Using the adjusted weights, each feature in the load feature matrix is ​​combined according to its weight to obtain the projection of the synthetic load data into the load feature matrix space.

[0194] By projecting composite load data into the space of the load feature matrix, we can capture the main features of the composite load data, better understand the characteristics and patterns of the composite load data, and provide data support for power distribution operation and management.

[0195] It is understandable that the composite feature matrix and the load feature matrix have a corresponding relationship. The load feature matrix is ​​obtained from real active power data, while the composite feature matrix is ​​obtained from virtual data. The load feature matrix has a corresponding second label, and the composite feature matrix also has a corresponding second label. That is, the second label is determined by the shape of each cluster center in the composite feature matrix.

[0196] In one specific implementation, the 99th percentile value Q_99 is calculated for each time point in the historical load data. If the load value at any time point in the synthetic feature matrix exceeds Q_99, the second label is marked as "extreme high load scenario".

[0197] Depending on the actual situation, the second label can be further marked. For example, the "extreme high load scenario" can be further labeled based on meteorological conditions.

[0198] If the wind speed is greater than level 10, the scenario with wind speed greater than level 10 and load exceeding Q_99 is marked as "Typhoon High Load". If the temperature label T is less than -10℃ and the load fluctuation rate is greater than 30%, it is marked as "Extreme Low Temperature Load".

[0199] S106: Based on the hybrid cluster center matrix, perform secondary cluster analysis to obtain the target cluster center matrix, thereby completing the construction of the power distribution operation scenario.

[0200] The target cluster center matrix includes features of various scenarios during power distribution operation.

[0201] Reference Figure 7 Based on the mixed cluster center matrix, a secondary clustering analysis is performed to obtain the target cluster center matrix, which specifically includes:

[0202] Obtain the historical within-cluster standard deviation of each cluster in the historical cluster center matrix and the composite within-cluster standard deviation of each cluster in the composite cluster center matrix;

[0203] The splitting threshold adjustment function is determined based on the historical intra-class standard deviation and the composite intra-class standard deviation;

[0204] Obtain the historical cluster spacing of each cluster in the historical cluster center matrix and the synthetic cluster spacing of each cluster in the synthetic cluster center matrix;

[0205] The merging threshold adjustment function is determined based on historical cluster spacing and synthetic cluster spacing;

[0206] Based on the hybrid cluster center matrix, and combining the split threshold adjustment function and the merge threshold adjustment function, the division of various clusters in the secondary clustering process is adjusted, and the target cluster center matrix is ​​given.

[0207] Furthermore, the splitting threshold adjustment function is determined through the following steps:

[0208] Based on the maximum value of the historical within-class standard deviation and the maximum value of the composite within-class standard deviation, calculate the difference between the maximum value of the historical within-class standard deviation and the maximum value of the composite within-class standard deviation to obtain the within-class difference.

[0209] Based on the pre-set split adjustment coefficient, the intra-class difference is adjusted to provide a split adjustment term;

[0210] Based on the difference between the current splitting threshold and the splitting adjustment term, a splitting threshold adjustment function is given.

[0211] In one specific implementation, the historical within-cluster standard deviation corresponding to each cluster is extracted from the historical cluster center matrix, and the maximum value σ is taken. max-hist Simultaneously, the composite within-cluster standard deviation of each cluster in the composite cluster center matrix is ​​calculated, and the maximum value σ is taken. max-gen The splitting threshold adjustment function is specifically expressed as:

[0212]

[0213] in, θ is the adjusted current splitting threshold. split α is the current splitting threshold, and α is the splitting adjustment coefficient.

[0214] By setting a splitting threshold adjustment function, the current splitting threshold is gradually reduced to avoid over-splitting of complex patterns in the generated data; in this example, α=0.1. It is understandable that when the current splitting threshold is adjusted for the first time, θ... split The pre-set splitting threshold.

[0215] Furthermore, the threshold adjustment function is merged, specifically determined through the following steps:

[0216] Based on the minimum historical cluster spacing and the minimum synthetic cluster spacing, the difference between the minimum historical cluster spacing and the minimum synthetic cluster spacing is calculated to obtain the cluster spacing difference.

[0217] Based on the pre-set merging adjustment coefficient, the cluster spacing difference is adjusted, and the merging adjustment item is given;

[0218] Based on the difference between the current merging threshold and the merging adjustment term, a merging threshold adjustment function is given.

[0219] In one specific implementation, the historical cluster spacing of each cluster is extracted from the historical cluster center matrix, and the minimum value D is taken. min_hist Meanwhile, the inter-cluster spacing of each cluster is extracted from the synthetic cluster center matrix, and the minimum value D is taken. min_gen The merging threshold adjustment function is specifically expressed as follows:

[0220]

[0221] in, θ is the adjusted current merging threshold. merge β is the current merging threshold, and β is the merging adjustment coefficient.

[0222] By setting a merging threshold adjustment function, the current merging threshold is gradually reduced to avoid erroneous merging of historical scenes and synthesized scenes. In this example, β=0.05. It can be understood that when the current merging threshold is adjusted for the first time, θ... merge The pre-set merging threshold.

[0223] In a hybrid clustering center matrix formed by power grid load data in a specific implementation scenario, for the historical clustering center matrix, there are N clustering centers with typical load patterns within a certain time period. For the synthetic clustering center matrix, synthetic clustering centers with load patterns similar to those of the historical clustering center matrix but with certain changes are generated within the same time period, thereby increasing the diversity of samples.

[0224] The historical cluster center matrix and the synthetic cluster center matrix are merged to obtain a hybrid cluster center matrix. A splitting threshold adjustment function is set based on the standard deviation of each cluster. For example, if the standard deviation of a cluster exceeds a preset initial splitting threshold, it is considered that different load patterns exist within the cluster, requiring splitting. The splitting threshold adjustment function can dynamically adjust the strictness of splitting based on the magnitude of the standard deviation; the larger the standard deviation, the lower the splitting threshold should be, making it easier to trigger splitting operations for more detailed load pattern segmentation. Similarly, for some adjacent clusters, their similarity is calculated; when the similarity is higher than a merging threshold, they are considered for merging. The merging threshold adjustment function can be set according to the actual situation of power distribution operation and the objectives of cluster analysis to simplify the clustering results and highlight the main load patterns.

[0225] Furthermore, referring to Figure 8 Based on the hybrid cluster center matrix, and combining the splitting threshold adjustment function and the merging threshold adjustment function, the partitioning of various clusters in the secondary clustering process is adjusted, and the target cluster center matrix is ​​given, specifically including:

[0226] Based on each cluster center in the mixed cluster center matrix and the cluster corresponding to each cluster center, give the maximum intra-cluster standard deviation of each cluster and the inter-cluster distance between each cluster;

[0227] If the maximum intra-class standard deviation is greater than the current classification threshold, the cluster corresponding to the maximum intra-class standard deviation will be split into two clusters.

[0228] If the cluster spacing is less than the current merging threshold and the difference in the number of samples is less than the difference in the number of samples threshold, the two clusters corresponding to the cluster spacing will be merged.

[0229] The current split threshold and the current merge threshold are updated based on the split threshold adjustment function and the merge threshold adjustment function;

[0230] Based on the splitting and / or merging scenarios, the individual cluster centers are redefined;

[0231] If the number of clusters changes or the number of iterations reaches the corresponding threshold, the second-order clustering is completed, and the target cluster center matrix is ​​given.

[0232] In one specific implementation, cluster affiliation is determined based on the cluster centers in the mixed cluster center matrix and the clusters corresponding to each cluster center. Then, for each cluster, the maximum intra-cluster standard deviation σ is obtained. max ,like The data is split into two sub-clusters along the principal component axis. For example, in a certain application scenario, cluster A contains a large number of transformer load data points. Calculations show that the standard deviation of this cluster is large, indicating significant differences in the load patterns of transformers within the cluster. Principal component analysis (PCA) is performed on the transformer load data in cluster A. The principal component axis of this cluster is calculated, revealing that the first principal component (PC1) represents the load difference between weekdays and weekends, and the second principal component (PC2) represents the load difference between high-temperature weather and normal weather. These two principal components can explain most of the variance of the load data within the cluster. PC1 can be selected as the principal splitting axis, and the transformer load data in cluster A can be split into two sub-clusters along the PC1 axis according to a set splitting threshold (the splitting threshold can be determined based on the distribution of PC1 scores, serving as the splitting boundary). In subsequent construction of distribution operation scenarios, these sub-clusters can be analyzed and predicted separately.

[0233] For the inter-cluster spacing between each cluster, obtain the inter-cluster spacing between each cluster. Furthermore, if the difference in sample size is less than the sample difference threshold, the two clusters corresponding to the inter-cluster distance are merged. Here, the difference in sample size N... ij Specifically, it is expressed as:

[0234]

[0235] Where, N ij N represents the difference in the number of samples between the i-th cluster and the j-th cluster. i N is the number of samples in the i-th cluster. j Let be the number of samples in the j-th cluster.

[0236] After completing one cluster merging and / or splitting operation, the cluster centers are recalculated. If the change in the number of clusters or the number of iterations reaches a corresponding threshold, a second clustering operation is performed, and the target cluster center matrix is ​​provided. For example, regarding the change in the number of clusters, if the change is less than a preset threshold, the update of the cluster centers is stopped, and the target cluster center matrix is ​​provided. Regarding the number of iterations, if the maximum number of iterations is reached, iteration is stopped, and the target cluster center matrix is ​​provided. It covers new scenarios (such as extreme low temperature load scenarios and typhoon load scenarios), among which K * This represents the final number of clusters.

[0237] In a specific example, the number of clusters changes N o / n Specifically, it is expressed as:

[0238]

[0239] Where, N n N represents the number of clusters after splitting and / or merging. oThe number of clusters before splitting and / or merging.

[0240] Based on the above descriptions of specific implementation scenarios, the division of various clusters in the secondary clustering process is adjusted to obtain more reasonable clustering results that better reflect the actual operation of the power grid. This provides a more accurate load pattern division for the construction of distribution operation scenarios, which in turn helps with power grid planning, operation scheduling, and reliability assessment.

[0241] Reference Figure 9 This invention provides a device for constructing power distribution operation scenarios based on load clustering, comprising:

[0242] Data acquisition module 201 is used to acquire multiple sets of historical load data for power distribution operation;

[0243] Feature extraction module 202 is used to extract features from each group of historical load data to obtain a load feature matrix;

[0244] The initial cluster determination module 203 is used to determine the initial cluster center set from the load feature matrix by combining the preset initial cluster number;

[0245] The first clustering module 204 is used to perform cluster analysis based on each initial cluster center in the initial cluster center set, obtain the historical cluster center matrix, and give the first label;

[0246] The hybrid cluster center determination module 205 is used to obtain the synthetic cluster center matrix and merge it with the historical cluster center matrix to give the hybrid cluster center matrix. The synthetic cluster center matrix is ​​obtained based on the synthetic load data given by the historical load data.

[0247] The second clustering module 206 is used to perform secondary clustering analysis based on the hybrid clustering center matrix to obtain the target clustering center matrix, so as to complete the construction of the power distribution operation scenario. The target clustering center matrix includes the characteristics of various scenarios in the power distribution operation process.

[0248] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0249] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0250] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0251] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0252] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0253] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0254] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and variations of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and variations.

Claims

1. A power distribution operation scenario construction method based on load clustering, characterized in that, The method comprises the following steps: obtaining a plurality of groups of historical load data of power distribution operation; performing feature extraction on each group of historical load data respectively to obtain a load feature matrix; determining an initial cluster center set from the load feature matrix in combination with a preset initial cluster number; performing clustering analysis based on each initial cluster center in the initial cluster center set to obtain a historical clustering center matrix and give a first label; obtaining a synthetic clustering center matrix and fusing the historical clustering center matrix to give a mixed clustering center matrix, wherein the synthetic clustering center matrix is obtained based on synthetic load data given by the historical load data, the synthetic clustering center matrix is determined by the following steps: generating synthetic load data based on preset label information in combination with a pre-constructed generative adversarial network model; projecting the synthetic load data into the space of the load feature matrix to obtain a synthetic feature matrix; performing pre-clustering on the synthetic feature matrix to obtain the synthetic clustering center matrix; fusing the historical clustering center matrix and the synthetic clustering center matrix to give the mixed clustering center matrix; horizontally splicing the historical clustering center matrix and the synthetic clustering center matrix, i.e. connecting the columns of the historical clustering center matrix to the left of the synthetic clustering center matrix to form a new matrix, i.e. the mixed clustering center matrix; performing secondary clustering analysis based on the mixed clustering center matrix to obtain a target clustering center matrix to complete the construction of the power distribution operation scene, wherein the target clustering center matrix comprises the features of various scenes in the power distribution operation process.

2. The load cluster-based power distribution operation scenario construction method according to claim 1, characterized in that, A plurality of groups of historical load data of power distribution operation are obtained by the following steps: obtaining a plurality of groups of initial historical load data, wherein each group of initial historical load data comprises historical active power of a transformer and / or a power transmission line in a unit time; giving a power median corresponding to each group of initial historical load data according to the analysis of each group of initial historical load data; identifying power abnormal values in each group of initial historical load data in combination with the power median and correcting the power abnormal values; completing the power missing values in each group of initial historical load data in combination with the difference method and the power normal values in each group of initial historical load data to obtain the historical load data of each group.

3. The load cluster-based power distribution operation scenario construction method according to claim 2, characterized in that, The power abnormal values in each group of initial historical load data are identified in combination with the power median and the power abnormal values are corrected, specifically comprising: calculating the difference between the historical active power and the power median in each group of initial historical load data according to the power median to obtain an absolute deviation value; giving an absolute deviation median based on the absolute deviation value corresponding to each group of initial historical load data; obtaining a corresponding reasonable data range according to the absolute deviation median in combination with the data quantity in each group of initial historical load data; screening the initial historical load data according to the reasonable data range to identify the power abnormal values; fusing the power median, the absolute deviation median and the data quantity to complete the correction of the power abnormal values.

4. The load cluster-based power distribution operation scenario construction method according to claim 2 or 3, characterized in that, The power missing values in each group of initial historical load data are completed in combination with the difference method and the power normal values in each group of initial historical load data to obtain the historical load data of each group, specifically comprising: determining the missing position of the power missing values in each group of initial historical load data; According to the missing position, the corresponding power normal value is selected from the initial historical load data; The power completion data is given by using the Lagrange interpolation method combined with the corresponding power normal value, and the historical load data of each group is obtained.

5. The load cluster based power distribution operational scenario construction method of claim 1, wherein, The feature extraction is performed on each group of historical load data respectively to obtain a load feature matrix, which specifically includes: According to each group of historical load data, a load sample matrix is constructed; The load sample matrix is preprocessed to obtain a standard load sample matrix and a corresponding correlation coefficient matrix; Based on the correlation coefficient matrix and the unit matrix, a load feature equation is constructed, and a plurality of load feature values corresponding to the load feature equation are given; Combined with a plurality of load feature values and a unit load feature vector corresponding to the load feature value, the cumulative contribution rate corresponding to each group of historical load data is given, and a feature load matrix is obtained. The feature load matrix and the standard load sample matrix are fused to obtain a load feature matrix.

6. The load cluster based power distribution operational scenario construction method of claim 1, wherein, The initial cluster center set is determined from the load feature matrix combined with the preset initial cluster number, which specifically includes: Any load feature is randomly selected from the load feature matrix to give a first cluster center; According to the distance between the first cluster center and each load feature in the load feature matrix, the probability of each load feature being the next cluster center is analyzed, and a second cluster center is given, until the cluster centers of all initial cluster numbers are obtained, and the initial cluster center set is given.

7. The load cluster based power distribution operational scenario construction method of claim 1, wherein, The construction of the generative adversarial network model is determined by the following steps: Initialize the model parameters of the generative adversarial network model, wherein the model parameters include the number of neurons in each layer, the loss function, the penalty coefficient, and the generator parameter update step; The conditional label corresponding to the historical load data is spliced with the noise vector through the full connection layer in the generative adversarial network model; The generator in the generative adversarial network model outputs the reference load data according to the historical load data and the conditional label corresponding to the historical load data; The discriminator in the generative adversarial network model outputs the gradient penalty and updates the discriminator parameters according to the historical load data and the corresponding reference load data combined with the penalty coefficient; The generator parameters are updated once every preset number of times of updating the discriminator parameters; Based on the similarity of the reference load data and the historical load data, the generative adversarial network model is trained until convergence.

8. The load cluster based power distribution operational scenario construction method of claim 1, wherein, Based on the hybrid clustering center matrix, secondary clustering analysis is performed to obtain a target clustering center matrix, which specifically includes: Obtain the historical intra-cluster standard deviation of each cluster in the historical clustering center matrix and the synthetic intra-cluster standard deviation of each cluster in the synthetic clustering center matrix; Determine the split threshold adjustment function based on the historical intra-cluster standard deviation and the synthetic intra-cluster standard deviation; Obtain the historical cluster distance of each cluster in the historical clustering center matrix and the synthetic cluster distance of each cluster in the synthetic clustering center matrix; Determine the merging threshold adjustment function based on the historical cluster distance and the synthetic cluster distance; Based on the hybrid clustering center matrix, the split threshold adjustment function and the merging threshold adjustment function are combined to adjust the division of various clusters in the secondary clustering process, and the target clustering center matrix is given.

9. The load cluster-based power distribution operational scenario construction method of claim 8, wherein, The split threshold adjustment function is determined by the following steps: According to the maximum value of the historical intra-class standard deviation and the maximum value of the synthetic intra-class standard deviation, a difference value between the maximum value of the historical intra-class standard deviation and the maximum value of the synthetic intra-class standard deviation is calculated to obtain an intra-class difference value; The intra-class difference value is adjusted in combination with a preset split adjustment coefficient to give a split adjustment term; A split threshold adjustment function is given based on a difference value between the current split threshold and the split adjustment term.

10. The load cluster based power distribution operational scenario construction method of claim 8, wherein, The merging threshold adjustment function is determined by the following steps: According to the minimum value of the historical cluster distance and the minimum value of the synthetic cluster distance, a difference value between the minimum value of the historical cluster distance and the minimum value of the synthetic cluster distance is calculated to obtain a cluster distance difference value; The cluster distance difference value is adjusted in combination with a preset merging adjustment coefficient to give a merging adjustment term; A merging threshold adjustment function is given based on a difference value between the current merging threshold and the merging adjustment term.

11. The load cluster based power distribution operational scenario construction method of claim 8, wherein, Based on the hybrid clustering center matrix, the split threshold adjustment function and the merging threshold adjustment function, the division of various clusters in the secondary clustering process is adjusted to give a target clustering center matrix, which specifically includes: According to each clustering center in the hybrid clustering center matrix and the cluster corresponding to each clustering center, the maximum intra-class standard deviation of each cluster and the cluster distance between each cluster are given; If the maximum intra-class standard deviation is greater than the current classification threshold, the cluster corresponding to the maximum intra-class standard deviation is split into two clusters; If the cluster distance is less than the current merging threshold and the sample number difference is less than the sample difference threshold, the two clusters corresponding to the cluster distance are merged; The current split threshold and the current merging threshold are updated based on the split threshold adjustment function and the merging threshold adjustment function; Based on the split and / or merge, each clustering center is re-determined; If the cluster number change or the number of iterations reaches the corresponding threshold, the secondary clustering is completed, and the target clustering center matrix is given.

12. The load cluster-based power distribution operational scenario construction method of claim 11, wherein, The cluster splitting is performed along the principal component axis.

13. The load cluster based power distribution operational scenario construction method of claim 11, wherein, The sample difference number is the ratio of the difference value of the cluster number before and after merging and / or splitting to the cluster number before merging and / or splitting.

14. A power distribution operation scenario construction device based on load clustering, characterized by, The power distribution operation scenario construction method based on load clustering comprises the following steps: A data acquisition module is configured to acquire a plurality of groups of historical load data of power distribution operation; A feature extraction module is configured to perform feature extraction on each group of historical load data to obtain a load feature matrix; An initial cluster determination module is configured to determine an initial cluster center set from the load feature matrix in combination with a preset initial cluster number; A first clustering module is configured to perform clustering analysis based on each initial cluster center in the initial cluster center set to obtain a historical clustering center matrix and give a first label. The mixing center determination module is configured to obtain a synthetic clustering center matrix and fuse a historical clustering center matrix to give a mixed clustering center matrix, wherein the synthetic clustering center matrix is obtained based on synthetic load data given by historical load data, the synthetic clustering center matrix is determined by the following steps: generating synthetic load data according to preset label information and in combination with a pre-constructed generative adversarial network model; projecting the synthetic load data into a space of a load feature matrix to obtain a synthetic feature matrix; pre-clustering the synthetic feature matrix to obtain a synthetic clustering center matrix; fusing the historical clustering center matrix and the synthetic clustering center matrix to give the mixed clustering center matrix; and horizontally splicing the historical clustering center matrix and the synthetic clustering center matrix, that is, connecting the columns of the historical clustering center matrix to the left of the synthetic clustering center matrix to form a new matrix, namely the mixed clustering center matrix. The second clustering module is configured to perform secondary clustering analysis based on the mixed clustering center matrix to obtain a target clustering center matrix, so as to complete construction of a power distribution operation scene, wherein the target clustering center matrix includes features of various scenes in a power distribution operation process.

Citation Information

Patent Citations

  • Power distribution network operation scene extraction method and device

    CN115526267A

  • Multi-time clustering method and system for disappearing officers

    CN118445648A

  • Power consumption demand automatic response method based on energy consumption subentry measurement

    CN119494523A