A Method for Evaluating the Operation Status of Distribution Networks Based on Interval Type-II Fuzzy Clustering Analysis
By optimizing the cluster center update formula based on interval type II fuzzy clustering analysis, the problem of evaluating unbalanced data in the distribution network is solved, enabling high-precision evaluation of the distribution network's operating status and fault prediction, thereby improving the stability and reliability of the power supply system.
Patent Information
- Application Number
- CN202111468189.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-12-03
AI Technical Summary
Existing technologies suffer from data overfitting or information loss when processing imbalance data in distribution networks, making it difficult to accurately and quickly assess the operating status of the distribution network and affecting power supply reliability and security.
A method based on interval type II fuzzy clustering analysis is adopted. By introducing the concept of rough set, boundary region data in imbalanced data are included in the clustering analysis, the cluster center update formula is optimized, and the clustering accuracy and state assessment accuracy are improved.
It improves the clustering effect of distribution network imbalance data, enhances the accuracy and speed of distribution network operation status assessment, enables timely detection of existing faults and prediction of potential faults, and improves the stability of the power supply system.
Smart Images

Figure CN114597886B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power distribution network technology, and in particular to a method for evaluating the operating status of power distribution networks based on interval type-two fuzzy clustering analysis. Background Technology
[0002] Currently, my country's power distribution network is rapidly expanding, forming a complex distribution architecture characterized by distributed systems, multiple loads, and multiple power sources. End-users are increasingly demanding higher reliability from power supply providers. However, power outages caused by distribution network failures can lead to significant economic losses and impact social stability. Therefore, the most crucial task is to accurately and quickly analyze distribution network data and implement reasonable and effective maintenance strategies to ensure the stable operation of the national power supply system. Comprehensive aggregation and analysis of distribution network data can effectively assess the network's operational status and identify changes and potential safety risks.
[0003] The distribution network is a crucial link connecting the transmission network and users, and its operational status directly affects the reliability of the power supply system. Even within the highly stable, normally operating data of the distribution network, a small amount of anomalous data still exists; these two types of data together constitute the imbalanced data of the distribution network. Existing research on processing imbalanced datasets mainly focuses on the data preprocessing level, that is, using techniques such as "oversampling" and "undersampling" to transform the imbalanced dataset into a roughly balanced dataset before using existing algorithms for analysis.
[0004] Oversampling techniques continuously interpolate data samples from minority clusters to generate new data samples, increasing the size of minority clusters and reducing cluster imbalance. Undersampling techniques randomly select samples from majority clusters, reducing the number of samples in the majority clusters and reducing cluster imbalance. While data preprocessing methods such as sampling can address cluster imbalance to some extent, these methods inevitably lead to overfitting or information loss. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and propose a distribution network operation status assessment method based on interval type II fuzzy clustering analysis. By introducing the concept of rough set, data that are difficult to distinguish in terms of status can be included in the boundary region, reminding operation and maintenance personnel to pay attention to these boundary data, thereby improving the clustering accuracy and status assessment accuracy.
[0006] The technical problem solved by this invention is achieved through the following technical solution:
[0007] The method for evaluating the operating status of distribution networks based on interval type-II fuzzy clustering analysis includes the following steps:
[0008] Step 1: Collect imbalance monitoring data of the substation. Based on the common fault index system of the power system, set threshold parameters for common fault types in the distribution imbalance monitoring data, and use the collected imbalance monitoring data as data samples.
[0009] Step 2: According to the interval type II fuzzy c-means clustering algorithm, randomly select data from the unbalanced monitoring data as the initial cluster centers in the iteration process of the interval type II fuzzy c-means clustering algorithm, and set the clustering analysis parameters according to the characteristics of historical data;
[0010] Step 3: Calculate the Euclidean distance between each data sample in the imbalanced monitoring data and the initial cluster center. Based on the Euclidean distance, divide the data samples in the imbalanced monitoring data into the lower approximate set or boundary region of the normal cluster or the abnormal cluster, and calculate the imbalance degree between the clusters.
[0011] Step 4: Substitute the cluster imbalance degree calculated in Step 3 into the optimized cluster center update formula based on interval type II fuzzy clustering analysis for iterative calculation, calculate the cluster center, and determine the cluster to which the sample belongs by calculating the membership degree.
[0012] Step 5: Compare the cluster centers and clusters calculated in Step 4 with the cluster centers and clusters of the previous iteration. If the cluster centers and clusters are no longer updated, count the samples of the approximate sets and boundary regions on each cluster to evaluate the operating status of the power distribution network; otherwise, return to Step 3.
[0013] Moreover, the specific implementation method of step 2 is as follows: randomly select two types of data from the unbalanced detection data as the initial cluster centers, one set as the initial cluster centers of the normal clusters of the distribution network, and the other set as the initial cluster centers of the abnormal clusters of the distribution network, and set the distance judgment threshold and fuzzy coefficient according to the historical data characteristics of the actual operation records of the distribution network system.
[0014] Furthermore, step 3 includes the following steps:
[0015] Step 3.1: Calculate the first Euclidean distance between each data sample in the unbalanced monitoring dataset and the cluster center of the normal cluster in Step 2, and the second Euclidean distance between each sample and the cluster center of the abnormal cluster. Determine the magnitude of the first Euclidean distance and the second Euclidean distance, and obtain the ratio of the larger value to the smaller value.
[0016] Step 3.2: Compare the ratio obtained in Step 3.1 with the distance judgment threshold. If the ratio is greater than the distance judgment threshold, the data sample is assigned to the lower approximate set of the cluster corresponding to the smaller Euclidean distance; otherwise, it is assigned to the boundary region.
[0017] Step 3.3: Calculate the ratio of the number of samples in the upper approximation set in both the normal and abnormal clusters to the total number of samples in all upper approximation sets in the imbalanced monitoring data, to obtain the imbalance degree f between the normal and abnormal clusters:
[0018]
[0019] in, This represents the number of approximate set samples on the minority class cluster of the cross-cluster in the current loop iteration. is the number of samples on the approximate set of the majority class cluster of the cross-cluster.
[0020] Furthermore, the specific implementation method for calculating the cluster centers in step 4 is as follows: substituting the imbalance degree calculated in step 3 into the optimized cluster center v based on interval type II fuzzy clustering analysis. i :
[0021]
[0022] Among them, v i Let ω be the cluster center in the i-th iteration. l ω is the approximate weighting coefficient. b Here are the approximate weighting coefficients, f is the imbalance degree, m is the fuzzy coefficient, and X is the weighting coefficient. ij For the data sample, x j For the sample, a ij For type II fuzzy membership, C i For the following approximate region dataset, This is a dataset for the boundary region.
[0023] Furthermore, the specific implementation method for calculating the membership degree to determine the cluster to which the sample belongs in step 4 is as follows:
[0024]
[0025]
[0026] Where, μ ij For membership degree, For μ ij membership degree, μ ij For μ ij The lower membership degree, distance d ji v is the cluster center in the i-th iteration. i With sample x j The distance between them, distance d zi v represents the cluster center in the i-th iteration. i With sample data sample x z The distance between them, where k is the number of clusters. C iFor the following approximate region dataset, This is a dataset representing a boundary region; for sample x j Compared to cluster C i The fuzzy membership degree is determined by the interval for:
[0027] μ i (x j )=min{μ ij (m1),μ ij (m2)}
[0028]
[0029] Where, μ ij (m1) and μ ij (m2) represent x when the fuzzy coefficients m = m1 and m = m2, respectively. j Compared to cluster C i A type of fuzzy membership measure, sample x j To cluster C i final membership degree β ij for:
[0030]
[0031] Where, N i For cluster C i The number of samples contained in the upper approximation set, where N is the total number of samples, is determined according to the final membership degree β. ij Determine the cluster to which all sample data belong.
[0032] Furthermore, step 5 includes the following steps:
[0033] Step 5.1: Based on the clustering results of the data samples in Step 3, iteratively update the cluster centers.
[0034] Step 5.2: Determine whether the cluster centers have been updated. If the cluster centers are no longer being updated, proceed to step 5.3; otherwise, return to step 3.
[0035] Step 5.3: Count the approximate set samples under the normal cluster. These samples are determined to be normal data and are marked with 0, indicating that the corresponding distribution network operation system has not experienced this type of fault. Count the approximate set samples under the abnormal cluster. These samples are determined to be fault samples and are marked with "1", indicating that the distribution network operation system has experienced this type of fault. Count the boundary area samples and mark these samples with "2", indicating that the distribution network operation system may experience this type of fault in the future.
[0036] The advantages and positive effects of this invention are:
[0037] 1. This invention employs an interval-based type-two c-means fuzzy clustering method based on local fuzzy metrics in the boundary region. This incorporates data imbalance factors into the cluster center update function, ensuring that cluster centers are correlated not only with the membership function of the imbalanced dataset but also with the degree of imbalance between clusters. The interval-based type-two fuzzy c-means clustering algorithm is used to calculate, analyze, and cluster data samples in the boundary region. An optimized cluster center update function considering local fuzzy metrics in the boundary region is introduced, improving the aggregation and clustering effect of the interval-based type-two fuzzy c-means clustering method on imbalanced operating data of distribution networks. This invention applies the improved aggregation and clustering algorithm to analyze frequently alarming imbalanced data in distribution networks, assessing the operating status of the distribution network. It can not only determine the types of faults that have already triggered alarms but also predict the probability of future alarms in the distribution network.
[0038] 2. This invention optimizes the cluster center update formula of the interval type II fuzzy c-means algorithm based on considering the local fuzzy metric of the boundary region. This reduces the adverse impact of the boundary region occupied by the majority cluster on the clustering effect of the minority cluster, thus keeping the cluster centers of small-scale clusters in a relatively ideal position. It can also suppress the phenomenon that data originally belonging to the majority cluster is misclassified into the minority cluster, thereby better preserving the data characteristics of the minority cluster. Therefore, it can improve the clustering performance of the algorithm on unbalanced data and can also improve the accuracy and speed of data aggregation in distribution networks. It not only has high academic research significance but also has strong engineering application value. Attached Figure Description
[0039] Figure 1 This is a structural diagram of the evaluation method of the present invention;
[0040] Figure 2 This is a flowchart of the evaluation method of the present invention;
[0041] Figure 3 This is a spatial distribution diagram of the dimensionality reduction data of the present invention.
[0042] Figure 4 Image of traditional fuzzy c-means clustering results
[0043] Figure 5 This is a diagram showing the results of the interval type-two fuzzy c-means clustering of the present invention. Detailed Implementation
[0044] The present invention will be further described in detail below with reference to the accompanying drawings.
[0045] A method for assessing the operating status of distribution networks based on interval type-2 fuzzy clustering analysis, such as Figure 1 and Figure 2 As shown, it includes the following steps:
[0046] Step 1: Collect imbalance monitoring data of the substation. Based on the common fault index system of the power system, set threshold parameters for common fault types in the distribution imbalance monitoring data, and use the collected imbalance monitoring data as data samples.
[0047] Based on the common fault indicator system of power systems, different types of fault data that are abnormal compared with the normal operation data of the distribution network are screened from the unbalanced power monitoring data, and the threshold parameters required for each fault are analyzed. Common alarm information of the distribution system is shown in Table 1, and the alarm thresholds of each monitoring variable of the distribution system are shown in Table 2.
[0048] Table 1. List of common alarms in power distribution systems
[0049]
[0050]
[0051] Table 2 Commonly Monitored Variables in Power Distribution Systems
[0052]
[0053]
[0054] Step 2: Using the interval type-two fuzzy c-means algorithm, two sets of data from the unbalanced monitoring data are randomly selected as the initial cluster centers in the iteration process of the interval type-two fuzzy c-means algorithm, and the parameters involved in the clustering analysis algorithm are set according to the characteristics of historical data. The initial cluster center data selection is random; that is, two sets of data from the unbalanced monitoring data are randomly selected, one set as the initial cluster center for the normal state cluster of the distribution network, and the other set as the initial cluster center for the abnormal state cluster of the distribution network. A distance judgment threshold is set based on the historical data characteristics of the actual operation records of the distribution network system. The fuzzy coefficients are m1 = 2 and m2 = 10.
[0055] Step 3: Calculate the Euclidean distance between each sample in the imbalanced data and the initial cluster center. Based on the Euclidean distance, divide the data samples in the imbalanced monitoring data into the lower approximate set or boundary region of normal or abnormal clusters, and calculate the imbalance degree between clusters. The steps are as follows:
[0056] Step 3.1: Calculate the first Euclidean distance between each data sample in the imbalanced dataset and the center of the normal cluster, and the second Euclidean distance between each data sample and the center of the abnormal cluster. Determine the magnitude of the first Euclidean distance and the second Euclidean distance, and obtain the ratio of the larger value to the smaller value.
[0057] Step 3.2: Compare the obtained ratio with the distance judgment threshold. If the ratio is greater than the distance judgment threshold, the data sample is assigned to the lower approximate set of the cluster corresponding to the smaller Euclidean distance between the first Euclidean distance and the second Euclidean distance; otherwise, it is assigned to the boundary region.
[0058] Step 3.3: Calculate the ratio of the number of samples in the upper approximation set in both the normal and abnormal clusters to the total number of samples in all upper approximation sets in the imbalanced monitoring data, to obtain the imbalance degree f between the normal and abnormal clusters:
[0059]
[0060] in, This represents the number of approximate set samples under the minority class cluster of the cross-cluster in the current loop iteration. This represents the number of samples in the approximate set of the majority class cluster of the cross-cluster.
[0061] Step 4: Substitute the cluster imbalance calculated in Step 3 into the optimized cluster center v. i The formula is updated to iteratively calculate the cluster centers v. i :
[0062]
[0063] Among them, v i Let ω be the cluster center in the i-th iteration. l ω is the approximate weighting coefficient. b Here are the approximate weighting coefficients, f is the imbalance degree, m is the fuzzy coefficient, and X is the weighting coefficient. ij For the data sample, x j For the sample, a ij For type II fuzzy membership, C i For the following approximate region dataset, This is a dataset for the boundary region.
[0064] To address the impact of imbalanced cluster sizes on the clustering results of the interval type-two fuzzy c-means algorithm, this invention first proposes a cluster center update method based on interval type-two fuzzy clustering analysis. The iterative center update formula is optimized to consider the imbalance of samples in the boundary region, suppressing the phenomenon where minority cluster centers shift due to the pull of boundary region data. In the optimized cluster center update formula, the greater the imbalance of the dataset, the smaller the iterative center coefficient f in the boundary region, and the smaller the contribution weight of the boundary region to the cluster centers, thereby achieving the goal of suppressing the shift of cluster centers towards the majority clusters.
[0065] The cluster to which a sample belongs is determined by calculating membership degrees:
[0066]
[0067]
[0068] Where, μ ij For membership degree, For μ ij membership degree, μ ij For μ ij The lower membership degree, distance d ji v is the cluster center in the i-th iteration. i With sample x j The distance between them, distance d zi v represents the cluster center in the i-th iteration. i With sample data sample x z The distance between them, where k is the number of clusters. C i For the following approximate region dataset, This is a dataset for the boundary region.
[0069] For sample x j Compared to cluster C i The fuzzy membership degree is determined by the interval for:
[0070] μ i (x j )=min{μ ij (m1),μ ij (m2)}
[0071]
[0072] Where, μ ij (m1) and μ ij (m2) represent x when the fuzzy coefficients m = m1 and m = m2, respectively. j Compared to cluster C i A type of fuzzy membership measure, sample x j To cluster C i final membership degree β ij for:
[0073]
[0074] Where, N i For cluster C i The number of samples contained in the upper approximation set, where N is the total number of samples, is determined according to the final membership degree β. ij Determine the cluster to which all sample data belong.
[0075] Step 5: Compare the cluster centers and clusters calculated in Step 4 with the cluster centers and clusters of the previous iteration. If the cluster centers and clusters are no longer updated, count the samples of the approximate sets and boundary regions on each cluster to evaluate the operating status of the power distribution network; otherwise, return to Step 3.
[0076] Step 5.1: Based on the clustering results of the data samples in Step 3, iteratively update the cluster centers.
[0077] Step 5.2: Determine whether the cluster centers have been updated. If the cluster centers are no longer being updated, proceed to step 5.3; otherwise, return to step 3.
[0078] Step 5.3: Count the approximate set samples under the normal cluster. These samples are determined to be normal data and are marked with 0, indicating that the corresponding distribution network operation system has not experienced this type of fault. Count the approximate set samples under the abnormal cluster. These samples are determined to be fault samples and are marked with "1", indicating that the distribution network operation system has experienced this type of fault. Count the boundary area samples and mark these samples with "2", indicating that the distribution network operation system may experience this type of fault in the future.
[0079] This invention employs an interval-based type-two fuzzy c-means algorithm to classify the operating status of the distribution network into normal and abnormal states. The thresholds corresponding to each alarm type in Table 1 are extracted from the original data, and then the interval-based type-two fuzzy c-means algorithm based on unbalanced data clustering proposed in this invention is used to perform cluster analysis on the distribution network imbalance monitoring data. Based on the clustering results of the data samples in step 3, the cluster centers are iteratively updated. If the cluster centers are no longer updated, the samples of the lower approximation set and boundary regions in the corresponding clusters are statistically analyzed to assess the operating status of the distribution network: The samples of the lower approximation set in the normal cluster are statistically analyzed; these samples are determined to be normal data and are marked "0", indicating that the corresponding distribution network operating system has not experienced this type of fault; the samples of the lower approximation set in the fault cluster are statistically analyzed; these samples are determined to be fault samples and are marked "1", indicating that the distribution network operating system has experienced this type of fault; the samples of the boundary region are statistically analyzed and marked "2", indicating that the distribution network operating system may experience this type of fault in the future. If the cluster centers continue to be updated, the process returns to step 3.
[0080] Expressed through a calculation formula: Let the dataset be U, U = {X} z |z=1,...,N}, divide the data object set U into k clusters; initialize the cluster centers v i Distance judgment threshold And fuzzy coefficients m1, m2.
[0081] For each object X z Calculate Xz to each cluster center v i Euclidean distance d ij Choose o = {j|d jz =min({d iz})}, and i = 1,...,k, if but and Otherwise x z ∈ C j For all clusters, Where, d iz For cluster center v i To data object X z The distance, d j,z For cluster center v j' To data object X z The distance, d jz For cluster center v j To data object X z The distance between them, where o and o' are the sets of data j and j' that satisfy the conditions, respectively. The distance threshold is used for judgment. C j and These are the boundary region, lower approximation region, and upper approximation region of the rough set, respectively.
[0082] Calculate the approximate set on each cluster Number of samples | C j Calculate the weight coefficient f of the cluster centers in the boundary region, substitute the imbalance coefficient f into the improved cluster center formula, and iteratively update the center points of the clusters. If the clusters no longer change, the algorithm terminates; otherwise, iterative updates continue.
[0083] Based on the above-mentioned method for evaluating the operating status of distribution networks using interval type-II fuzzy clustering analysis, relevant experiments were conducted using imbalance monitoring data from a certain distribution substation:
[0084] The distribution network operation status assessment method proposed in this invention fully utilizes the concept of rough sets, clustering distribution network imbalance data into a majority class of normal data and a minority class of abnormal data. If a set of monitoring data is determined to belong to the lower approximate set of the minority class of abnormal data, then this set of data is identified as an existing fault, and real-time alarm processing is performed. If the set of data belongs to the boundary region of the minority class cluster, then this set of data belongs to a possible fault, and predictive alarm is performed, thereby achieving a status assessment of the entire distribution network operation system. The proposed status assessment method can determine the time and location of existing faults and predict possible faults for data in the abnormal boundary region, making it more advanced than traditional status diagnosis methods. The method proposed in this invention can be deployed in the monitoring and analysis room of the hub substation to analyze various monitoring data of the substation in a timely manner, assisting operation and maintenance personnel in promptly discovering status changes and existing safety hazards.
[0085] For example, the test data of a certain low-voltage power distribution equipment is shown in Table 3:
[0086] Table 3 Example of monitoring data for a certain low-voltage power distribution equipment
[0087]
[0088]
[0089] To more intuitively demonstrate the algorithm's effectiveness, Principal Component Analysis (PCA) was used to identify the main feature projections in the data, removing noise and redundancy. The 6-dimensional sample data selected in Table 3 was then reduced to a 3-dimensional space, as shown in the results. Figure 3 As shown, the pattern “☆” represents minority abnormal data, and the pattern “□” represents majority safe operating data.
[0090] Experimental results:
[0091] like Figure 4 and Figure 5 The results of the traditional fuzzy c-means algorithm and the interval type II c-means algorithm are shown. In the figures, the patterns "☆" and "□" represent samples that are correctly clustered; the pattern "*" represents samples that originally belonged to the majority cluster but were incorrectly classified into approximate regions under at least a few clusters; and the pattern "o" represents samples that were classified into the boundary region.
[0092] Based on the experimental test results of 20 sets of data under the proposed state assessment model, it is evident that although the traditional fuzzy c-means clustering method incorporates the concepts of upper and lower approximations from rough set theory and assigns some unpredictable samples to the boundary space, it fails to consider the impact of uneven cluster sizes on the clustering results. Its clustering performance for data with uneven cluster sizes is not ideal. Specifically, in the traditional fuzzy c-means clustering algorithm, two majority cluster samples were incorrectly clustered into at least a few clusters, and five majority cluster samples were incorrectly clustered into the boundary space, resulting in unsatisfactory results for distribution network state assessment. The distribution network operation state assessment method based on the interval type-two fuzzy clustering algorithm proposed in this chapter optimizes the cluster center update formula while considering uneven cluster sizes. The improved algorithm is more suitable for the aggregation and clustering analysis of uneven operation data in distribution systems compared to the traditional algorithm. In the clustering results of the interval type-two fuzzy c-means clustering method, only one set of majority cluster data was incorrectly clustered into at least a few clusters, and another set of majority cluster data was incorrectly clustered into the at least a few cluster boundary region. Therefore, the evaluation method proposed in this invention can effectively aggregate and cluster the unbalanced monitoring data generated by distribution network equipment, and then evaluate the operation status of the distribution network, thereby effectively improving the speed of distribution network system alarms and the accuracy of fault prediction.
[0093] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.
Claims
1. A method for evaluating the operating status of a distribution network based on interval type-two fuzzy clustering analysis, characterized in that: Includes the following steps: Step 1: Collect imbalance monitoring data of the substation. Based on the common fault index system of the power system, set threshold parameters for common fault types in the distribution imbalance monitoring data, and use the collected imbalance monitoring data as data samples. Step 2: According to the interval type II fuzzy c-means clustering algorithm, randomly select data from the unbalanced monitoring data as the initial cluster centers in the iteration process of the interval type II fuzzy c-means clustering algorithm, and set the clustering analysis parameters according to the characteristics of historical data; Step 3: Calculate the Euclidean distance between each data sample in the imbalanced monitoring data and the initial cluster center. Based on the Euclidean distance, divide the data samples in the imbalanced monitoring data into the lower approximate set or boundary region of the normal cluster or the abnormal cluster, and calculate the imbalance degree between the clusters. Step 4: Substitute the cluster imbalance degree calculated in Step 3 into the optimized cluster center update formula based on interval type II fuzzy clustering analysis for iterative calculation, calculate the cluster center, and determine the cluster to which the sample belongs by calculating the membership degree. The specific implementation method for calculating the cluster centers in step 4 is as follows: Substitute the imbalance degree calculated in step 3 into the optimized cluster center v based on interval type II fuzzy clustering analysis. i : Among them, v i Let ω be the cluster center in the i-th iteration. l ω is the approximate weighting coefficient. b Here are the approximate weighting coefficients, f is the imbalance degree, m is the fuzzy coefficient, and X is the weighting coefficient. ij For the data sample, x j For the sample, a ij For type II fuzzy membership, C i For the following approximate region dataset, This is a dataset for the boundary region. The specific implementation method for calculating the membership degree to determine the cluster to which the sample belongs in step 4 is as follows: Where, μ ij For membership degree, For μ ij membership degree, μ ij For μ ij The lower membership degree, distance d ji v is the cluster center in the i-th iteration. i With sample x j The distance between them, distance d zi v represents the cluster center in the i-th iteration. i With sample data sample x z The distance between them, where k is the number of clusters. C i For the following approximate region dataset, This is a dataset representing a boundary region; for sample x j Compared to cluster C i The fuzzy membership degree is determined by the interval for: μ i (x j )=min{μ ij (m1),m ij (m2)} Where, μ ij (m1) and μ ij (m2) represent x when the fuzzy coefficients m = m1 and m = m2, respectively. j Compared to cluster C i A type of fuzzy membership measure, sample x j To cluster C i final membership degree β ij for: Where, N i For cluster C i The number of samples contained in the upper approximation set, where N is the total number of samples, is determined according to the final membership degree β. ij Determine the cluster to which all sample data belong; Step 5: Compare the cluster centers and clusters calculated in Step 4 with the cluster centers and clusters of the previous iteration. If the cluster centers and clusters are no longer updated, count the samples of the approximate sets and boundary regions on each cluster to evaluate the operating status of the power distribution network; otherwise, return to Step 3.
2. The method for evaluating the operating status of a distribution network based on interval type-two fuzzy clustering analysis according to claim 1, characterized in that: The specific implementation method of step 2 is as follows: randomly select two types of data from the unbalanced detection data as the initial cluster centers, one set as the initial cluster centers of the normal clusters of the distribution network, and the other set as the initial cluster centers of the abnormal clusters of the distribution network. The distance judgment threshold and fuzzy coefficient are set according to the historical data characteristics of the actual operation records of the distribution network system.
3. The method for evaluating the operating status of a distribution network based on interval type-two fuzzy clustering analysis according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Calculate the first Euclidean distance between each data sample in the unbalanced monitoring dataset and the cluster center of the normal cluster in Step 2, and the second Euclidean distance between each sample and the cluster center of the abnormal cluster. Determine the magnitude of the first Euclidean distance and the second Euclidean distance, and obtain the ratio of the larger value to the smaller value. Step 3.2: Compare the ratio obtained in Step 3.1 with the distance judgment threshold. If the ratio is greater than the distance judgment threshold, the data sample is assigned to the lower approximate set of the cluster corresponding to the smaller Euclidean distance; otherwise, it is assigned to the boundary region. Step 3.3: Calculate the ratio of the number of samples in the upper approximation set in both the normal and abnormal clusters to the total number of samples in all upper approximation sets in the imbalanced monitoring data, to obtain the imbalance degree f between the normal and abnormal clusters: in, This represents the number of approximate set samples on the minority class cluster of the cross-cluster in the current loop iteration. is the number of samples on the approximate set of the majority class cluster of the cross-cluster.
4. The method for evaluating the operating status of a distribution network based on interval type-two fuzzy clustering analysis according to claim 1, characterized in that: Step 5 includes the following steps: Step 5.1: Based on the clustering results of the data samples in Step 3, iteratively update the cluster centers. Step 5.2: Determine whether the cluster centers have been updated. If the cluster centers are no longer being updated, proceed to step 5.3; otherwise, return to step 3. Step 5.3: Count the approximate set samples under the normal cluster. These samples are determined to be normal data and are marked with 0, indicating that the corresponding distribution network operation system has not experienced the corresponding type of fault. Count the approximate set samples under the abnormal cluster. These samples are determined to be fault samples and are marked with "1", indicating that the distribution network operation system has experienced the corresponding type of fault. Count the boundary area samples and mark these samples with "2", indicating that the distribution network operation system may experience the corresponding type of fault in the future.