Method for analyzing and evaluating operation state of power distribution network

Through the improved K-means clustering algorithm and graph neural network model, combined with data cleaning and outlier detection, the accuracy and intelligence of distribution network operating status evaluation are solved, and the true reflection and real-time evaluation of distribution network operating status are realized.

CN120338280APending Publication Date: 2025-07-18TONGLING POWER SUPPLY CO OF STATE GRID ANHUI ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510479112.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing distribution network operating status evaluation methods are difficult to fully reflect the actual operating status, data processing is difficult, lack of in-depth analysis, insufficient intelligence level, and cannot accurately identify weak links and potential risks.

Method used

The improved K-means clustering algorithm is used to combine isolated forests and random forest algorithms for data cleaning, and the graph neural network model is used to generate a distribution network operating state portrait, and attribute weights are calculated through information entropy and inter-class variance, dynamic balance feature importance, outlier value detection and missing value filling are performed.

Benefits of technology

The accuracy and efficiency of outlier elimination is improved, and the accurate analysis and evaluation of the operating status of the distribution network is realized, which can reflect the real operating status, and solve the shortcomings of the existing methods in data randomness and ambiguity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338280A_ABST
    Figure CN120338280A_ABST
Patent Text Reader

Abstract

The invention relates to power distribution network management, in particular to a power distribution network operation state analysis and evaluation method, which comprises the following steps of: acquiring multi-dimensional data for analyzing and evaluating the operation state of a power distribution network from a power distribution network data acquisition and monitoring control system, and performing data cleaning on the multi-dimensional data; constructing a training data set based on the cleaned multi-dimensional data, and performing model training on a power distribution network operation state portrait generation model by using the training data set, so that the power distribution network operation state portrait generation model learns to generate a power distribution network operation state portrait according to the multi-dimensional data; acquiring multi-dimensional basic data of the power distribution network in real time, performing data cleaning on the multi-dimensional basic data, and inputting the cleaned multi-dimensional basic data into the trained power distribution network operation state portrait generation model to obtain a corresponding power distribution network operation state portrait; analyzing and evaluating the running state of the power distribution network according to the running state portrait of the power distribution network; according to the technical scheme provided by the invention, the defect that the operation state of the power distribution network is difficult to accurately analyze and evaluate in the prior art can be effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the management of distribution networks, and particularly to a method for analyzing and evaluating the operating state of a distribution network. Background Art

[0002] With the continuous development of the power system, as an important link connecting the transmission network and users, the stability and reliability of the operating state of the distribution network are directly related to the power supply quality and user satisfaction. However, due to the complex structure, numerous devices, and variable operating environment of the distribution network, its operating state will be affected by various factors such as load fluctuations, equipment failures, and natural disasters.

[0003] Currently, the analysis and evaluation of the operating state of the distribution network mainly face the following challenges:

[0004] 1) Limitations in evaluation methods: Existing evaluation methods for the operating state of distribution networks mostly focus on the equipment asset management level, such as indicators like equipment level, grid structure, and management level. Although these indicators can provide a basis for the planning of distribution networks, it is difficult to comprehensively reflect the actual operating state of the distribution network. Especially when the distribution network faces frequent and random fluctuations in its operating conditions, the practicality of existing methods is not strong;

[0005] 2) Difficulties in data processing: When the data volume is too large, existing methods are difficult to perform effective data cleaning, and the remaining outliers and missing values will have a greater impact on the accuracy of the evaluation. At the same time, the large data volume will also cause existing methods to be difficult to operate efficiently, limiting the real-time nature of the evaluation;

[0006] 3) Lack of in-depth analysis: Existing methods often lack in-depth analysis and excavation of the overall operation of the distribution network, and are unable to accurately identify weak links and potential risks in the distribution network;

[0007] 4) Insufficient intelligent level: With the development of smart grids, the intelligent level of distribution networks is constantly improving. However, existing methods still have deficiencies in intelligent applications and are unable to make full use of advanced technologies such as big data, cloud computing, and artificial intelligence for accurate and efficient evaluation. Summary of the Invention

[0008] (I) Technical Problems to be Solved

[0009] In view of the above-mentioned drawbacks of the existing technology, the present invention provides a method for analyzing and evaluating the operating state of a distribution network, which can effectively overcome the defect that it is difficult to accurately analyze and evaluate the operating state of the distribution network in the existing technology.

[0010] (II) Technical Solutions

[0011] To achieve the above object, the present invention is realized through the following technical solutions:

[0012] A method for analyzing and evaluating the operation status of a distribution network, comprising the following steps:

[0013] S1. Obtain multi-dimensional data for analyzing and evaluating the operation status of the distribution network from the distribution network data acquisition and monitoring control system, and perform data cleaning on the multi-dimensional data;

[0014] S2. Construct a training data set based on the cleaned multi-dimensional data, and use the training data set to train the distribution network operation status portrait generation model, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to the multi-dimensional data;

[0015] S3. Real-time collect the multi-dimensional basic data of the distribution network, perform data cleaning on the multi-dimensional basic data and then input it into the trained distribution network operation status portrait generation model to obtain the corresponding distribution network operation status portrait;

[0016] S4. Analyze and evaluate the operation status of the distribution network according to the distribution network operation status portrait to obtain the real-time evaluation result of the distribution network operation status;

[0017] Among them, in the process of performing data cleaning on the multi-dimensional data, the improved K-means clustering algorithm is used to perform more accurate secondary abnormal data judgment on the multi-dimensional data after preliminary outlier removal, and secondary outlier removal is performed on the multi-dimensional data based on the secondary abnormal data judgment result;

[0018] In the improved K-means clustering algorithm:

[0019] Calculate the initial attribute weights from the dual perspectives of information entropy and between-class variance, dynamically balance the importance of attribute features and class discrimination, so as to automatically identify key attribute features and improve the clustering effect of multi-dimensional data;

[0020] When calculating the distance between the attribute features of each multi-dimensional data and the clustering center, and updating the clustering center, introduce the attribute weights to improve the accuracy of the clustering boundary and accelerate the algorithm convergence;

[0021] Perform outlier detection by calculating the weighted local outlier factor and weighted density estimation of each multi-dimensional data to adapt to the differences in the importance of attribute features and improve the sensitivity of outlier detection.

[0022] Preferably, the data cleaning of the multi-dimensional data in S1 includes:

[0023] S11. Use the isolation forest algorithm to perform preliminary abnormal data judgment on the multi-dimensional data, and perform preliminary outlier removal on the multi-dimensional data based on the preliminary abnormal data judgment result;

[0024] S12. Use the improved K-means clustering algorithm to perform a secondary abnormal data judgment on the remaining multi-dimensional data, and eliminate the secondary outliers from the multi-dimensional data based on the results of the secondary abnormal data judgment;

[0025] S13. Use the improved random forest algorithm to fill in the missing values of the multi-dimensional data after removing the outliers.

[0026] Preferably, in S11, the isolation forest algorithm is used to perform a preliminary abnormal data judgment on the multi-dimensional data, and the preliminary outliers are removed from the multi-dimensional data based on the results of the preliminary abnormal data judgment, including:

[0027] S111. Construct an isolation forest model containing M isolation trees;

[0028] S112. For a data set D(N) containing N multi-dimensional data, traverse each isolation tree to obtain the path length of the multi-dimensional data D i in each isolation tree, that is, the number of edges passed from the root node to the leaf node, i ∈ [1, N];

[0029] S113. According to the path length of the multi-dimensional data D i in each isolation tree, calculate the average path length E i of the multi-dimensional data D i in all isolation trees:

[0030]

[0031] where, L i_s is the path length of the i-th multi-dimensional data D i in the s-th isolation tree;

[0032] S114. Calculate the average path length C(N) of the tree:

[0033]

[0034] where, H(N - 1) is the harmonic number, H(N - 1) ≈ In(N) + 0.5772156649, and 0.5772156649 is the Euler-Mascheroni constant;

[0035] S115. Calculate the anomaly score S(D i , N) of the multi-dimensional data D i :

[0036]

[0037] S116. Judge the anomaly score S(D i of the multi-dimensional data D i, N) is greater than the upper limit value of the preset range. If it is greater than the upper limit value of the preset range, then it is determined that the multi-dimensional data D i is abnormal data, and the abnormal data is excluded;

[0038] Among them, for the abnormal score S(D i , N) that falls within the preset range, the multi-dimensional data D i , it is impossible to determine whether it is abnormal data.

[0039] Preferably, in S12, the improved K-means clustering algorithm is used to perform a secondary abnormal data judgment on the remaining multi-dimensional data, and the multi-dimensional data is subjected to a secondary outlier rejection based on the secondary abnormal data judgment result, including:

[0040] S121. Select the initial clustering center, and calculate the initial attribute weights from the dual perspectives of information entropy and between-class variance;

[0041] S122. Calculate the weighted distance between each multi-dimensional data and the clustering center based on the attribute weights, and reassign each multi-dimensional data to the nearest clustering center;

[0042] S123. Update the clustering center according to the latest clustering result combined with the attribute weights, and update the attribute weights;

[0043] S124. Determine whether the iteration termination condition is satisfied. If the iteration termination condition is not satisfied, return to S122, otherwise enter S125;

[0044] S125. Calculate the weighted local outlier factor and weighted density estimation of each multi-dimensional data, determine whether the multi-dimensional data is abnormal data according to the calculation result, and exclude the abnormal data.

[0045] Preferably, in S121, the initial clustering center is selected, and the initial attribute weights are calculated from the dual perspectives of information entropy and between-class variance, including:

[0046] S1211. Select the initial clustering center;

[0047] S1212. Calculate the information entropy weight in the initial attribute weights using the following formula:

[0048]

[0049] Among them, is the information entropy weight of the attribute feature j, H j is the normalized information entropy value of the attribute feature j, H k is the normalized information entropy value of the attribute feature k, x ij is the value of the attribute feature j of the multi-dimensional data i, m is the number of attribute features, and n is the number of multi-dimensional data;

[0050] S1213. Calculate the between-class variance weight in the initial attribute weights using the following formula:

[0051]

[0052] Wherein, is the between-class variance weight of attribute feature j, σ j is the between-class variance of attribute feature j, σ k is the between-class variance of attribute feature k, μ cj is the mean of attribute feature j of the multi-dimensional data in class c, μ j is the mean of attribute feature j, n c is the number of multi-dimensional data in class c; b is the number of classes;

[0053] S1214. Calculate the initial attribute weights using the following formula:

[0054]

[0055] Wherein, ω j is the initial attribute weight of attribute feature j, and α is the first adjustment coefficient.

[0056] Preferably, in S122, calculating the weighted distance between each multi-dimensional data and the cluster center based on the attribute weights, and reassigning each multi-dimensional data to the nearest cluster center includes:

[0057] S1221. For the first iteration, calculate the weighted distance between each multi-dimensional data and the cluster center based on the initial attribute weights using the following formula:

[0058]

[0059] Wherein, d ω (i, C) is the weighted distance between multi-dimensional data i and cluster center C, P ij is the coordinate of attribute feature j of multi-dimensional data i, P Cj is the coordinate of attribute feature j of cluster center C, |·| represents calculating the modulus length;

[0060] For the second and subsequent iterations, calculate the weighted distance between each multi-dimensional data and the cluster center based on the updated attribute weights using the following formula:

[0061]

[0062] Wherein, is the updated attribute weight of attribute feature j;

[0063] S1222. According to the weighted distance between each multi-dimensional data and the cluster center, reassign each multi-dimensional data to the nearest cluster center.

[0064] Preferably, in S123, the cluster center is updated according to the latest clustering result in combination with the attribute weight, and the attribute weight is updated, including:

[0065] S1231. For the first iteration, the cluster center is updated according to the latest clustering result in combination with the initial attribute weight by the following formula:

[0066]

[0067] where, is the updated coordinate of the attribute feature j of the cluster center C, i∈c represents the multi-dimensional data set of multi-dimensional data i belonging to category c, and the center of category c is the cluster center C;

[0068] For the second and subsequent iterations, the cluster center is updated according to the latest clustering result in combination with the updated attribute weight by the following formula:

[0069]

[0070] S1232. The attribute weight is updated by the following formula:

[0071]

[0072] where, is the attribute weight of the attribute feature j in the previous iteration, is the weighted within-class variance of the attribute feature j of the multi-dimensional data in category c, and β is the second adjustment coefficient.

[0073] Preferably, in S125, the weighted local outlier factor and weighted density estimation of each multi-dimensional data are calculated, and it is determined whether the multi-dimensional data is abnormal data according to the calculation results, and the abnormal data is excluded, including:

[0074] S1251. The weighted local outlier factor of each multi-dimensional data is calculated by the following formula:

[0075]

[0076] where, is the weighted local outlier factor of the multi-dimensional data i, l∈S k (i) represents the k-nearest neighbor sample set S of the multi-dimensional data l belonging to the multi-dimensional data i k (i), d ω (i, l) is the weighted distance between the multi-dimensional data i and the multi-dimensional data l, is the weighted distance between the multi-dimensional data l and the multi-dimensional data of the k+1 nearest neighbor of the multi-dimensional data i, is the k-nearest neighbor sample set S of the multi-dimensional data i k(i) The number of multidimensional data in;

[0077] S1252, using the following formula to calculate the weighted density estimate of each multidimensional data:

[0078]

[0079] in, is the weighted density estimate of multidimensional data i, q∈S ε (i) indicates that the multidimensional data q belongs to the ε neighborhood sample set S of the multidimensional data i ε (i) is the updated attribute weight of multidimensional data q, d ω (i,q) is the weighted distance between multidimensional data i and multidimensional data q, is the weighted distance d between multidimensional data i and multidimensional data q ω The standardized result of (i,q);

[0080] S1253: When the weighted local outlier factor of the multidimensional data is greater than a first preset threshold and the multidimensional data of the multidimensional data is less than a second preset threshold, the multidimensional data is determined to be abnormal data, and the abnormal data is eliminated.

[0081] Preferably, in S13, an improved random forest algorithm is used to fill missing values in the multidimensional data after outliers are removed, including:

[0082] S131, interpolating the multidimensional data after removing outliers to obtain an interpolation matrix, and constructing a random forest model according to the interpolation matrix;

[0083] S132, for each column of the interpolation matrix, when a certain column is used as the target column, the remaining columns constitute a filling matrix;

[0084] S133. Use the random forest model to make predictions based on the data of the filling matrix and the target column to obtain the prediction results of the target column, and fill the missing values of the target column based on the prediction results.

[0085] Preferably, in S2, a training data set is constructed based on the cleaned multidimensional data, and the distribution network operation status portrait generation model is trained using the training data set, so that the distribution network operation status portrait generation model learns to generate a distribution network operation status portrait according to the multidimensional data, including:

[0086] S21, obtaining the cleaned multidimensional data, and using the hierarchical analysis method to perform a correlation analysis on the multidimensional data about the operation status of the distribution network, to obtain a preliminary evaluation result of the operation status of the distribution network;

[0087] S22. Adjust the preliminary evaluation result of the distribution network operation status in combination with historical experience data and actual production requirements to obtain the evaluation result of the distribution network operation status;

[0088] S23. Use the evaluation result of the distribution network operation status as the label of the corresponding multi-dimensional data, and construct a training data set based on the multi-dimensional data and its corresponding label;

[0089] S24. Input the training data set into the graph neural network model GNN for model training, so that the graph neural network model GNN learns to generate a distribution network operation status portrait according to the multi-dimensional data.

[0090] (III) Beneficial effects

[0091] Compared with the prior art, the distribution network operation status analysis and evaluation method provided by the present invention has the following beneficial effects:

[0092] 1) When removing outliers from multi-dimensional data, since the K-means clustering algorithm requires a large amount of computing resources and processing time when dealing with large-scale data, the isolation forest algorithm is first used to perform preliminary outlier removal on the multi-dimensional data, and then the improved K-means clustering algorithm is used to perform secondary outlier removal on the data that cannot be effectively determined in the multi-dimensional data after preliminary outlier removal, so as to effectively improve the accuracy and efficiency of outlier removal, and ensure that the distribution network operation status portrait generated by the finally trained distribution network operation status portrait generation model can truly reflect the distribution network operation status;

[0093] 2) In the improved K-means clustering algorithm, the initial attribute weights are calculated from the dual perspectives of information entropy and between-class variance to dynamically balance the importance of attribute features and the class discrimination degree, so as to automatically identify key attribute features and improve the clustering effect of multi-dimensional data; when calculating the distance between the attribute features of each multi-dimensional data and the clustering center, and updating the clustering center, the attribute weights are introduced to improve the accuracy of the clustering boundary and accelerate the algorithm convergence; outlier detection is performed by calculating the weighted local outlier factor and weighted density estimation of each multi-dimensional data to adapt to the differences in the importance of attribute features and improve the sensitivity of outlier detection. Through these three aspects of improvement, accurate identification and removal of abnormal data in large-scale data can be achieved, so as to accurately analyze and evaluate the distribution network operation status;

[0094] 3) Use the training data set to train the model of the distribution network operation status portrait generation model, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to multi-dimensional data, and use the trained distribution network operation status portrait generation model to obtain the distribution network operation status portrait corresponding to the real-time collected multi-dimensional basic data. Finally, obtain the real-time evaluation result of the distribution network operation status according to the distribution network operation status portrait, which can effectively solve the problem that the existing methods have great deficiencies in dealing with the randomness and ambiguity of distribution network data, and make the real-time evaluation result of the distribution network operation status reflect the real operation status of the distribution network. Brief Description of the Drawings

[0095] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0096] Figure 1 It is a flowchart of the present invention;

[0097] Figure 2 It is a flowchart of data cleaning for multi-dimensional data in the present invention. Detailed Embodiments

[0098] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0099] A method for analyzing and evaluating the operation status of a distribution network, as Figure 1 shown, includes the following steps:

[0100] S1. Obtain multi-dimensional data for analyzing and evaluating the operation status of the distribution network from the distribution network data acquisition and monitoring control system, and perform data cleaning on the multi-dimensional data;

[0101] S2. Construct a training data set based on the cleaned multi-dimensional data, and use the training data set to train the model of the distribution network operation status portrait generation model, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to multi-dimensional data;

[0102] S3. Collect the multi-dimensional basic data of the distribution network in real time, clean the multi-dimensional basic data, and input it into the trained distribution network operation state portrait generation model to obtain the corresponding distribution network operation state portrait;

[0103] S4. Analyze and evaluate the operation state of the distribution network based on the distribution network operation state portrait to obtain the real-time evaluation result of the distribution network operation state;

[0104] Among them, in the process of cleaning the multi-dimensional data, the improved K-means clustering algorithm is used to make a more accurate secondary abnormal data judgment on the multi-dimensional data after preliminary outlier removal, and the multi-dimensional data is subjected to secondary outlier removal based on the secondary abnormal data judgment result;

[0105] In the improved K-means clustering algorithm:

[0106] Calculate the initial attribute weights from the dual perspectives of information entropy and between-class variance, dynamically balance the importance of attribute features and class discrimination, automatically identify key attribute features, and improve the clustering effect of multi-dimensional data;

[0107] When calculating the distance between the attribute features of each multi-dimensional data and the clustering center, and updating the clustering center, introduce the attribute weights to improve the accuracy of the clustering boundary and accelerate the algorithm convergence;

[0108] Detect outliers by calculating the weighted local outlier factor and weighted density estimation of each multi-dimensional data to adapt to the differences in the importance of attribute features and improve the sensitivity of outlier detection.

[0109] In S1, data cleaning is performed on the multi-dimensional data, as Figure 2 shown, including:

[0110] S11. Use the isolation forest algorithm to make a preliminary abnormal data judgment on the multi-dimensional data, and perform preliminary outlier removal on the multi-dimensional data based on the preliminary abnormal data judgment result;

[0111] S12. Use the improved K-means clustering algorithm to make a secondary abnormal data judgment on the remaining multi-dimensional data, and perform secondary outlier removal on the multi-dimensional data based on the secondary abnormal data judgment result;

[0112] S13. Use the improved random forest algorithm to fill in the missing values of the multi-dimensional data after outlier removal.

[0113] ① In S11, the isolation forest algorithm is used to make a preliminary abnormal data judgment on the multi-dimensional data, and preliminary outlier removal is performed on the multi-dimensional data based on the preliminary abnormal data judgment result, including:

[0114] S111. Construct an isolation forest model containing M isolation trees;

[0115] S112. For a data set D(N) containing N multidimensional data, traverse each isolated tree to obtain the multidimensional data D i The path length in each isolated tree is the number of edges from the root node to the leaf node, i∈[1,N];

[0116] S113. According to multidimensional data D i The path length in each isolated tree is used to calculate the multidimensional data D i The average path length E among all isolated trees i :

[0117]

[0118] Among them, L i_s is the i-th multidimensional data D i The length of the path in the sth isolated tree;

[0119] S114. Calculate the average path length C(N) of the tree:

[0120]

[0121] Among them, H(N-1) is the harmonic number, H(N-1)≈In(N)+0.5772156649, 0.5772156649 is the Euler-Mascheroni constant;

[0122] S115. Calculate multidimensional data D i The anomaly score S(D i ,N):

[0123]

[0124] S116. Determine multidimensional data D i The anomaly score S(D i ,N) is greater than the upper limit of the preset range. If it is greater than the upper limit of the preset range, it is determined that the multidimensional data D i Abnormal data is identified and removed;

[0125] Among them, for the anomaly score S(D i ,N) Multidimensional data D that falls within the preset range i , it is impossible to determine whether it is abnormal data.

[0126] ②In S12, an improved K-means clustering algorithm is used to perform secondary abnormal data judgment on the remaining multidimensional data, and secondary abnormal values of the multidimensional data are removed based on the secondary abnormal data judgment results, such as Figure 2 As shown, including:

[0127] S121, select the initial cluster center, and calculate the initial attribute weights from the dual perspectives of information entropy and inter-class variance;

[0128] S122, calculating the weighted distance between each multidimensional data and the cluster center based on the attribute weight, and reallocating each multidimensional data to the nearest cluster center;

[0129] S123, updating the cluster center according to the latest clustering result combined with the attribute weight, and updating the attribute weight;

[0130] S124, judging whether the iteration termination condition is satisfied, if not, returning to S122, otherwise entering S125;

[0131] S125, calculating the weighted local outlier factor and weighted density estimation of each multidimensional data, determining whether the multidimensional data is abnormal data according to the calculation results, and eliminating the abnormal data.

[0132] 1) In S121, the initial cluster center is selected, and the initial attribute weight is calculated from the dual perspectives of information entropy and inter-class variance, including:

[0133] S1211, selecting an initial cluster center;

[0134] S1212. Calculate the information entropy weight in the initial attribute weight using the following formula:

[0135]

[0136] in, is the information entropy weight of attribute feature j, H j is the normalized information entropy value of attribute feature j, H k is the normalized information entropy value of attribute feature k, x ij is the value of attribute feature j of multidimensional data i, m is the number of attribute features, and n is the number of multidimensional data;

[0137] S1213. Calculate the inter-class variance weight in the initial attribute weight using the following formula:

[0138]

[0139] in, is the inter-class variance weight of attribute feature j, σ j is the between-class variance of attribute feature j, σ k is the inter-class variance of attribute feature k, μ cj is the mean of attribute feature j of multidimensional data in category c, μ j is the mean of attribute feature j, n c is the number of multidimensional data in category c, and b is the number of categories;

[0140] S1214. Calculate the initial attribute weights using the following formula:

[0141]

[0142] where ω j is the initial attribute weight of attribute feature j, and α is the first adjustment coefficient.

[0143] 2) In S122, calculate the weighted distances between each multi-dimensional data and the cluster centers based on the attribute weights, and reassign each multi-dimensional data to the nearest cluster center, including:

[0144] S1221. For the first iteration, calculate the weighted distances between each multi-dimensional data and the cluster centers using the following formula based on the initial attribute weights:

[0145]

[0146] where d ω (i, C) is the weighted distance between multi-dimensional data i and cluster center C, P ij is the coordinate of attribute feature j of multi-dimensional data i, P Cj is the coordinate of attribute feature j of cluster center C, and |·| represents calculating the modulus length;

[0147] For the second and subsequent iterations, calculate the weighted distances between each multi-dimensional data and the cluster centers using the following formula based on the updated attribute weights:

[0148]

[0149] where is the updated attribute weight of attribute feature j;

[0150] S1222. Reassign each multi-dimensional data to the nearest cluster center according to the weighted distances between each multi-dimensional data and the cluster centers.

[0151] 3) In S123, update the cluster centers according to the latest clustering results combined with the attribute weights, and update the attribute weights, including:

[0152] S1231. For the first iteration, update the cluster centers using the following formula according to the latest clustering results combined with the initial attribute weights:

[0153]

[0154] where is the updated coordinate of attribute feature j of cluster center C, i ∈ c represents the set of multi-dimensional data where multi-dimensional data i belongs to category c, and the center of category c is cluster center C;

[0155] For the second and subsequent iterations, the cluster centers are updated according to the latest clustering results in combination with the updated attribute weights using the following formula:

[0156]

[0157] S1232. The attribute weights are updated using the following formula:

[0158]

[0159] where is the attribute weight of the attribute feature j in the previous iteration, is the weighted within-class variance of the attribute feature j of the multi-dimensional data in the class c, and β is the second adjustment coefficient.

[0160] 4) In S125, calculate the weighted local outlier factor and weighted density estimate of each multi-dimensional data, determine whether the multi-dimensional data is abnormal data according to the calculation results, and eliminate the abnormal data, including:

[0161] S1251. Calculate the weighted local outlier factor of each multi-dimensional data using the following formula:

[0162]

[0163] where is the weighted local outlier factor of the multi-dimensional data i, l ∈ S k (i) indicates that the multi-dimensional data l belongs to the k-nearest neighbor sample set S k (i) of the multi-dimensional data i, d ω (i, l) is the weighted distance between the multi-dimensional data i and the multi-dimensional data l, is the weighted distance between the multi-dimensional data l and the multi-dimensional data of the k + 1 nearest neighbor of the multi-dimensional data i, is the number of multi-dimensional data in the k-nearest neighbor sample set S k (i) of the multi-dimensional data i;

[0164] S1252. Calculate the weighted density estimate of each multi-dimensional data using the following formula:

[0165]

[0166] where is the weighted density estimate of the multi-dimensional data i, q ∈ S ε (i) indicates that the multi-dimensional data q belongs to the ε-neighborhood sample set S ε (i) of the multi-dimensional data i, is the updated attribute weight of the multi-dimensional data q, d ω (i, q) is the weighted distance between the multi-dimensional data i and the multi-dimensional data q, The weighted distance d between the multi-dimensional data i and the multi-dimensional data q ω (i,q) normalization result;

[0167] S1253. When the weighted local outlier factor of the multi-dimensional data is greater than the first preset threshold and the multi-dimensional data of this multi-dimensional data is less than the second preset threshold, then determine that this multi-dimensional data is abnormal data and eliminate the abnormal data.

[0168] ③ In S13, the improved random forest algorithm is used to fill in the missing values of the multi-dimensional data after removing the outliers, including:

[0169] S131. Interpolate the multi-dimensional data after removing the outliers to obtain an interpolation matrix, and construct a random forest model according to the interpolation matrix;

[0170] S132. For each column of the interpolation matrix, when a certain column is used as the target column, the remaining columns form a filling matrix;

[0171] S133. Use the random forest model to predict according to the data of the filling matrix and the target column to obtain the prediction result of the target column, and fill in the missing values of the target column based on the prediction result.

[0172] In the above technical solution, when removing outliers from multi-dimensional data, since the K-means clustering algorithm requires a large amount of computing resources and processing time when dealing with large-scale data, the isolation forest algorithm is first used to preliminarily remove outliers from the multi-dimensional data, and then the improved K-means clustering algorithm is used to perform secondary outlier removal on the data that cannot be effectively determined in the multi-dimensional data after preliminary outlier removal. Thus, the accuracy and efficiency of outlier removal can be effectively improved, ensuring that the generated distribution network operation status portrait of the finally trained distribution network operation status portrait generation model can truly reflect the distribution network operation status.

[0173] At the same time, in the improved K-means clustering algorithm, the initial attribute weights are calculated from the dual perspectives of information entropy and between-class variance to dynamically balance the importance of attribute features and the class discrimination degree, so as to automatically identify key attribute features and improve the clustering effect of multi-dimensional data; when calculating the distance between the attribute features of each multi-dimensional data and the clustering center, and updating the clustering center, the attribute weights are introduced to improve the accuracy of the clustering boundary and accelerate the algorithm convergence; outlier detection is performed by calculating the weighted local outlier factor and weighted density estimation of each multi-dimensional data to adapt to the differences in the importance of attribute features and improve the sensitivity of outlier detection. Through these three aspects of improvement, accurate identification and removal of abnormal data in large-scale data can be achieved, so as to accurately analyze and evaluate the distribution network operation status.

[0174] In S2, a training data set is constructed based on the cleaned multi-dimensional data, and the model of the distribution network operation status portrait generation model is trained using the training data set, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to the multi-dimensional data, including:

[0175] S21. Obtain the cleaned multi-dimensional data, and perform correlation analysis on the multi-dimensional data regarding the operation status of the distribution network using the analytic hierarchy process to obtain a preliminary evaluation result of the distribution network operation status;

[0176] S22. Adjust the preliminary evaluation result of the distribution network operation status in combination with historical experience data and actual production requirements to obtain an evaluation result of the distribution network operation status;

[0177] S23. Use the evaluation result of the distribution network operation status as the label corresponding to the multi-dimensional data, and construct a training data set based on the multi-dimensional data and its corresponding label;

[0178] S24. Input the training data set into the graph neural network model GNN for model training, so that the graph neural network model GNN learns to generate the distribution network operation status portrait according to the multi-dimensional data.

[0179] In the technical solution of this application, the model of the distribution network operation status portrait generation model is trained using the training data set, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to the multi-dimensional data, and the distribution network operation status portrait corresponding to the real-time collected multi-dimensional basic data is obtained using the trained distribution network operation status portrait generation model. Finally, the real-time evaluation result of the distribution network operation status is obtained according to the distribution network operation status portrait, which can effectively solve the problem that the existing methods have great deficiencies in dealing with the randomness and ambiguity of the distribution network data, and enables the real-time evaluation result of the distribution network operation status to reflect the real operation status of the distribution network.

[0180] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for analyzing and evaluating the operating state of a distribution network, characterized in that: It includes the following steps: S1. Obtain multi-dimensional data for analyzing and evaluating the operation status of the distribution network from the distribution network data acquisition and monitoring control system, and perform data cleaning on the multi-dimensional data; S2. Construct a training data set based on the cleaned multi-dimensional data, and use the training data set to train the distribution network operation status portrait generation model, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to the multi-dimensional data; S3. Real-time collect the multi-dimensional basic data of the distribution network, perform data cleaning on the multi-dimensional basic data and then input it into the trained distribution network operation status portrait generation model to obtain the corresponding distribution network operation status portrait; S4. Analyze and evaluate the operation status of the distribution network according to the distribution network operation status portrait to obtain the real-time evaluation result of the distribution network operation status; Among them, in the process of performing data cleaning on the multi-dimensional data, an improved K-means clustering algorithm is used to perform more accurate secondary abnormal data judgment on the multi-dimensional data after preliminary outlier removal, and secondary outlier removal is performed on the multi-dimensional data based on the secondary abnormal data judgment result; In the improved K-means clustering algorithm: Calculate the initial attribute weights from the dual perspectives of information entropy and between-class variance, dynamically balance the importance of attribute features and the class discrimination degree, so as to automatically identify key attribute features and improve the clustering effect of multi-dimensional data; Introduce attribute weights when calculating the distance between the attribute features of each multi-dimensional data and the clustering center, and when updating the clustering center, so as to improve the accuracy of the clustering boundary and accelerate the convergence of the algorithm; Perform outlier detection by calculating the weighted local outlier factor and weighted density estimation of each multi-dimensional data to adapt to the differences in the importance of attribute features and improve the sensitivity of outlier detection.

2. The method for analyzing and evaluating the operation state of a distribution network according to claim 1, wherein: The data cleaning of the multi-dimensional data in S1 includes: S11. Use the isolation forest algorithm to perform preliminary abnormal data judgment on the multi-dimensional data, and perform preliminary outlier removal on the multi-dimensional data based on the preliminary abnormal data judgment result; S12. Use the improved K-means clustering algorithm to perform secondary abnormal data judgment on the remaining multi-dimensional data, and perform secondary outlier removal on the multi-dimensional data based on the secondary abnormal data judgment result; S13. Use the improved random forest algorithm to fill in the missing values of the multi-dimensional data after outlier removal.

3. The method for analyzing and evaluating the operation state of a distribution network according to claim 2, wherein: In S11, using the isolation forest algorithm to perform preliminary abnormal data judgment on the multi-dimensional data and performing preliminary outlier removal on the multi-dimensional data based on the preliminary abnormal data judgment result includes: S111. Construct an isolation forest model containing M isolation trees; S112. For a dataset D(N) containing N multi-dimensional data, traverse each isolation tree to obtain the multi-dimensional data D i The path length in each isolation tree, that is, the number of edges passed from the root node to the leaf node, where i ∈ [1, N]; S113. Calculate the multi-dimensional data D i based on the path lengths in each isolated tree i and calculate the average path length E i of the multi-dimensional data D in all isolated trees Among them, L i_s is the path length of the i-th multi-dimensional data D i in the s-th isolation tree; S114. Calculate the average path length C(N) of the tree: Among them, H(N - 1) is the harmonic number, H(N - 1)≈In(N)+0.5772156649, and 0.5772156649 is the Euler-Mascheroni constant; S115. Calculate the multi-dimensional data D i 's anomaly score S(D i , N): S116. Determine the multi-dimensional data D i 's anomaly score S(D i , N) is greater than the upper limit of the preset range. If it is greater than the upper limit of the preset range, then determine that the multi-dimensional data D i is abnormal data and eliminate the abnormal data; Among them, for the multi-dimensional data D i where the anomaly score S(D i , N) falls within a preset range, it is impossible to determine whether it is anomalous data.

4. The method for analyzing and evaluating the operation state of a distribution network according to claim 2, characterized in that: In S12, using the improved K-means clustering algorithm to perform secondary abnormal data judgment on the remaining multi-dimensional data and performing secondary outlier removal on the multi-dimensional data based on the secondary abnormal data judgment result includes: S121. Select the initial clustering center, and calculate the initial attribute weights from the dual perspectives of information entropy and between-class variance; S122, calculating the weighted distance between each multidimensional data and the cluster center based on the attribute weight, and reallocating each multidimensional data to the nearest cluster center; S123, updating the cluster center according to the latest clustering result combined with the attribute weight, and updating the attribute weight; S124, judging whether the iteration termination condition is satisfied, if not, returning to S122, otherwise entering S125; S125, calculating the weighted local outlier factor and weighted density estimation of each multidimensional data, determining whether the multidimensional data is abnormal data according to the calculation results, and eliminating the abnormal data.

5. The method for analyzing and evaluating the operation state of a distribution network according to claim 4, wherein: In S121, the initial cluster center is selected, and the initial attribute weights are calculated from the dual perspectives of information entropy and inter-class variance, including: S1211, selecting an initial cluster center; S1212. Calculate the information entropy weight in the initial attribute weight using the following formula: Among them, is the information entropy weight of attribute feature j, H j is the normalized information entropy value of attribute feature j, H k is the normalized information entropy value of attribute feature k, x ij is the value of attribute feature j of multi-dimensional data i, m is the number of attribute features, and n is the number of multi-dimensional data; S1213. Calculate the inter-class variance weight in the initial attribute weight using the following formula: Among them, is the between-class variance weight of attribute feature j, σ j is the between-class variance of attribute feature j, σ k is the between-class variance of attribute feature k, μ cj is the mean of attribute feature j of the multi-dimensional data in class c, μ j is the mean of attribute feature j, n c is the number of multi-dimensional data in class c; b is the number of classes. S1214. Calculate the initial attribute weight using the following formula: Among them, ω j is the initial attribute weight of attribute feature j, and α is the first adjustment coefficient.

6. The method for analyzing and evaluating the operation state of a distribution network according to claim 5, characterized in that: In S122, the weighted distance between each multidimensional data and the cluster center is calculated based on the attribute weight, and each multidimensional data is redistributed to the nearest cluster center, including: S1221. For the first iteration, the weighted distance between each multidimensional data and the cluster center is calculated based on the initial attribute weight using the following formula: where d ω (i, C) is the weighted distance between the multi-dimensional data i and the cluster center C, P ij is the coordinate of the attribute feature j of the multi-dimensional data i, P Cj is the coordinate of the attribute feature j of the cluster center C, and |·| represents the calculation of the modulus length; For the second and subsequent iterations, the weighted distance between each multidimensional data and the cluster center is calculated based on the updated attribute weights using the following formula: Among them, is the updated attribute weight of attribute feature j; S1222. Redistribute each multidimensional data to the nearest cluster center according to the weighted distance between each multidimensional data and the cluster center.

7. The method for analyzing and evaluating the operating state of a distribution network according to claim 6, wherein: In S123, the cluster center is updated according to the latest clustering result combined with the attribute weight, and the attribute weight is updated, including: S1231. For the first iteration, the cluster center is updated using the following formula based on the latest clustering result combined with the initial attribute weights: Among them, is the updated coordinate of the attribute feature j of the clustering center C. i ∈ c represents the multi-dimensional data set of multi-dimensional data i belonging to category c, and the center of category c is the clustering center C; For the second and subsequent iterations, the cluster center is updated according to the latest clustering results combined with the updated attribute weights using the following formula: S1232. Update the attribute weight using the following formula: Among them, is the attribute weight of attribute feature j in the previous iteration, is the weighted within-class variance of attribute feature j of the multi-dimensional data in class c, and β is the second adjustment coefficient.

8. The method for analyzing and evaluating the operation state of a distribution network according to claim 7, characterized in that: In S125, the weighted local outlier factor and weighted density estimation of each multidimensional data are calculated, and whether the multidimensional data is abnormal data is determined according to the calculation results, and the abnormal data is eliminated, including: S1251, using the following formula to calculate the weighted local outlier factor of each multidimensional data: Among them, is the weighted local outlier factor of the multi-dimensional data i, l ∈ S k (i) represents that the multi-dimensional data l belongs to the set S of the k-nearest neighbor samples of the multi-dimensional data i k (i), d ω (i, l) is the weighted distance between the multi-dimensional data i and the multi-dimensional data l, is the weighted distance between the multi-dimensional data l and the multi-dimensional data of the (k + 1)-th nearest neighbor of the multi-dimensional data i, is the set S of the k-nearest neighbor samples of the multi-dimensional data i k the number of multi-dimensional data in (i); S1252, using the following formula to calculate the weighted density estimate of each multidimensional data: Among them, is the weighted density estimate of the multi-dimensional data i, q ∈ S ε (i) represents that the multi-dimensional data q belongs to the ε-neighborhood sample set S ε (i), is the updated attribute weight of the multi-dimensional data q, d ω (i, q) is the weighted distance between the multi-dimensional data i and the multi-dimensional data q, is the weighted distance d ω between the multi-dimensional data i and the multi-dimensional data q; the standardized result of (i, q). S1253: When the weighted local outlier factor of the multidimensional data is greater than a first preset threshold and the multidimensional data of the multidimensional data is less than a second preset threshold, the multidimensional data is determined to be abnormal data, and the abnormal data is eliminated.

9. The method for analyzing and evaluating the operating state of a distribution network according to claim 2, wherein: In S13, an improved random forest algorithm is used to fill missing values in multidimensional data after outliers are removed, including: S131, interpolating the multidimensional data after removing outliers to obtain an interpolation matrix, and constructing a random forest model according to the interpolation matrix; S132, for each column of the interpolation matrix, when a certain column is used as the target column, the remaining columns constitute a filling matrix; S133. Use the random forest model to make predictions based on the data in the filling matrix and the target column, obtain the prediction results of the target column, and fill in the missing values of the target column based on the prediction results.

10. The method for analyzing and evaluating the operation state of a distribution network according to claim 1, wherein: In S2, a training data set is constructed based on the cleaned multi-dimensional data, and the training data set is used to train the model of the distribution network operation status portrait generation model, so that the distribution network operation status portrait generation model learns to generate the distribution network operation status portrait according to the multi-dimensional data, including: S21. Obtain the cleaned multi-dimensional data, and use the analytic hierarchy process to perform a correlation analysis of the multi-dimensional data on the distribution network operation status to obtain a preliminary evaluation result of the distribution network operation status; S22. Adjust the preliminary evaluation result of the distribution network operation status in combination with historical experience data and actual production requirements to obtain the evaluation result of the distribution network operation status; S23. Use the evaluation result of the distribution network operation status as the label corresponding to the multi-dimensional data, and construct a training data set based on the multi-dimensional data and its corresponding label; S24. Input the training data set into the graph neural network model GNN for model training, so that the graph neural network model GNN learns to generate the distribution network operation status portrait according to the multi-dimensional data.