Power distribution network equipment risk assessment method and system based on entropy weight t-SNE and Gaussian mixture model

By combining the entropy weight t-SNE and Gaussian hybrid model, the problem of weight selection impact in the risk assessment of distribution network equipment is solved, and efficient risk assessment of complex network structures and a large number of outlets to be detected is achieved, which improves the credibility and accuracy of the assessment.

CN120471440APending Publication Date: 2025-08-12CHONGQING TECH & BUSINESS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510553095.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the existing distribution network equipment risk assessment method, weight selection affects the credibility of the evaluation results, and there is a problem of order reversal, making it difficult to adapt to complex network structures and a large number of outlets to be tested.

Method used

Using a method based on the hybrid model of entropy weight t-SNE and Gaussian, the optimal cluster cluster number was determined by calculating feature weights, dimensionality reduction, visualization and cluster analysis, combining the contour coefficient and the Calinski-Harabasz score, and risk assessment was performed.

Benefits of technology

It effectively avoids the impact of weight selection on the evaluation results, is suitable for complex network structures and a large number of outlets to be tested, improves the credibility and accuracy of the evaluation, and has good development prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471440A_ABST
    Figure CN120471440A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network equipment risk assessment method and system based on an entropy weight t-SNE and a Gaussian mixture model, and the method comprises the steps: obtaining a plurality of to-be-assessed power distribution network equipment, each to-be-assessed power distribution network equipment having n features, and obtaining the weight of each feature through an entropy weight method; adding the weight of each feature into the data of each feature to obtain a weighted data set, and performing dimensionality reduction on the weighted data set to obtain dimensionality-reduced data; visualizing the dimension-reduced data, and determining the range of the cluster number; and carrying out clustering analysis on the dimension reduction data by utilizing a Gaussian mixture model, and determining an optimal clustering cluster number by utilizing a contour coefficient and a Clinski-Harabasz score based on the range of the clustering cluster number to obtain a risk assessment result. The method is more suitable for power distribution network equipment risk assessment containing a large number of to-be-detected network points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network risk assessment, and in particular to a distribution network equipment risk assessment method and system based on entropy-weighted t-SNE and Gaussian mixture model. Background Art

[0002] As a key component of the power distribution network, the stability and maintenance of distribution network equipment are fundamental to its reliable operation. In recent years, the number and variety of distribution network equipment has rapidly increased, leading to increasingly complex network structures, more diverse operational modes, and a growing number of risks. Risk assessment of distribution network equipment can proactively identify potential safety hazards, reduce the risk of accidents, and thus safeguard the stability of the power system. It can also help identify weak links within the power system, strengthen protective measures for these links, and improve power system reliability. It can also identify and resolve issues that could impact the sustainable development of the power industry, thereby ensuring its long-term stability. Therefore, developing a scientific and rational distribution network equipment risk assessment method is crucial for ensuring the stability and security of power supply.

[0003] Numerous studies have investigated risk assessment for distribution network equipment, focusing primarily on model characteristics and weight selection. Different models and weight selections can affect the final evaluation results, thereby reducing their credibility. For example, the TOPSIS and MARCOS methods are both prone to order reversal during calculation. CRITIC weights have been used by many scholars for weight selection. However, CRITIC weights use standard deviations to measure the contrast strength of features. Using normalization to eliminate the influence of dimension on standard deviations can alter the size relationship of feature standard deviations, thereby reducing the feasibility of feature weights. Summary of the Invention

[0004] In order to solve the technical problems existing in the above-mentioned prior art, the present invention proposes a distribution network equipment risk assessment method and system based on entropy weight t-SNE and Gaussian mixture model, which can effectively avoid the problem that the evaluation results of other multi-criteria comprehensive evaluation methods are affected by weight selection.

[0005] On the one hand, to achieve the above-mentioned purpose, the present invention provides a distribution network equipment risk assessment method based on entropy-weighted t-SNE and Gaussian mixture model, comprising:

[0006] Obtain a number of distribution network devices to be evaluated, where each distribution network device to be evaluated has n features, and calculate the weight of each feature;

[0007] The weight of each feature is added to the data of each feature to obtain a weighted data set, and the dimension of the weighted data set is reduced to obtain reduced-dimensional data;

[0008] Visualizing the dimensionality-reduced data to determine the range of the number of clusters;

[0009] A Gaussian mixture model is used to perform cluster analysis on the dimensionality reduction data, and based on the range of the number of clusters, the silhouette coefficient and the Calinski-Harabasz score are used to determine the optimal number of clusters to obtain a risk assessment result.

[0010] On the other hand, to achieve the above-mentioned purpose, the present invention also provides a distribution network equipment risk assessment system based on entropy-weighted t-SNE and Gaussian mixture model, comprising:

[0011] Data acquisition unit: used to acquire a number of distribution network devices to be evaluated, wherein each distribution network device to be evaluated has n features and calculate the weight of each feature;

[0012] Dimensionality reduction processing unit: used for adding the weight of each feature to the data of each feature to obtain a weighted data set, and performing dimensionality reduction on the weighted data set to obtain reduced-dimensionality data;

[0013] Visualization unit: used for visualizing the dimensionality reduction data and determining the range of the number of clusters;

[0014] The risk assessment unit is used to perform cluster analysis on the dimensionality reduction data using a Gaussian mixture model, determine the optimal number of clusters based on the range of the number of clusters, and obtain a risk assessment result using the silhouette coefficient and the Calinski-Harabasz score.

[0015] Compared with the prior art, the present invention has the following advantages and technical effects:

[0016] The present invention proposes a distribution network equipment risk assessment model that combines entropy-weighted t-SNE with a Gaussian mixture model. This model can effectively avoid the problem of the evaluation results of other multi-criteria comprehensive evaluation methods being affected by the selection of weights, and does not have the order reversal problem that is prone to occur in methods such as TOPSIS and MARCOS. The present invention is more suitable for the risk assessment of distribution network equipment with a large number of network points to be tested. If more data can be collected in future research, the model performance will be further optimized and improved. Therefore, the distribution network equipment risk assessment model that combines entropy-weighted t-SNE with a Gaussian mixture model has good development prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0018] Figure 1Flowchart of a method for risk assessment of distribution network equipment based on entropy-weighted t-SNE and Gaussian mixture model according to an embodiment of the present invention;

[0019] Figure 2 Schematic diagram of clustering results of a Gaussian mixture model on a noise circle dataset according to an embodiment of the present invention;

[0020] Figure 3 Schematic diagram of clustering results of the Gaussian mixture model on a bimonthly dataset according to an embodiment of the present invention;

[0021] Figure 4 Schematic diagram of the comparison results of t-SNE combined with entropy weight and t-SNE on low-dimensional datasets according to an embodiment of the present invention, wherein (a) is a schematic diagram of the dimensionality reduction and KL divergence of the Iris dataset, (b) is a schematic diagram of the dimensionality reduction and KL divergence of the Raisin dataset, (c) is a schematic diagram of the dimensionality reduction and KL divergence of the Wine dataset, and (d) is a schematic diagram of the dimensionality reduction and KL divergence of the Statlog (Heart) dataset;

[0022] Figure 5 Schematic diagrams showing the comparison results of t-SNE combined with entropy weights and t-SNE on medium-dimensional datasets according to an embodiment of the present invention, wherein (a) is a schematic diagram showing the dimensionality reduction and KL divergence of the Breast cancer dataset, (b) is a schematic diagram showing the dimensionality reduction and KL divergence of the Dermatology dataset, (c) is a schematic diagram showing the dimensionality reduction and KL divergence of the Infrared Thermography Temperature dataset, and (d) is a schematic diagram showing the dimensionality reduction and KL divergence of the Horse Colic dataset;

[0023] Figure 6 Schematic diagrams showing the comparison results of entropy-weighted t-SNE and t-SNE on high-dimensional datasets according to an embodiment of the present invention, wherein (a) is a schematic diagram showing the dimensionality reduction and KL divergence of the Spambase dataset, (b) is a schematic diagram showing the dimensionality reduction and KL divergence of the Lung cancer dataset, (c) is a schematic diagram showing the dimensionality reduction and KL divergence of the Mice Protein Expression dataset, and (d) is a schematic diagram showing the dimensionality reduction and KL divergence of the Connectionist Bench dataset;

[0024] Figure 7 This is a visualization diagram of data dimensionality reduction according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0026] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0027] The Gaussian mixture model is to "mix Gaussian distributions together" and achieve the effect of fitting more complex distributions by mixing multiple Gaussian distributions with different parameters. Specifically, suppose there are K different Gaussian distributions with parameters μ1, μ2, ..., μ K ,σ1,σ2,…,σ K .

[0028] Let these parameters be abbreviated as θ, then a new distribution can be described by a mixture of K Gaussian distributions, whose density function is:

[0029]

[0030] in, They are the weights corresponding to K Gaussian distributions, satisfying This is the Gaussian mixture model.

[0031] In the Gaussian mixture model, the latent variable can be defined as the correspondence between each observation and the Gaussian component, that is, the relationship between the i-th observation and the k-th Gaussian component. The latent variable is denoted as γ ik , if the i-th observation follows the distribution of the k-th Gaussian component, then z ik =1, otherwise z ik = 0. For example, for the problem of height distribution of boys and girls, we can define the latent variable as the true gender of each person. For the i-th observation, z i1 =1 means it is a boy, z i2 =1 means it is a girl.

[0032] After adding latent variables, the complete data likelihood function of the Gaussian mixture model can be written as:

[0033]

[0034] The corresponding log-likelihood is:

[0035]

[0036] For the i-th observation, we want to know the hidden variable z in this data. ik The distribution of , using the conditional probability formula:

[0037]

[0038] Among them, P(x i |z ik ) means that under condition z ik The probability of the next i-th observation value, that is, f(x i |μ k ,Σ k ), P(z ik ) represents the probability that the i-th observation value belongs to the k-th Gaussian distribution, that is, P(x i ) represents the probability of the i-th observation in the Gaussian mixture model, that is:

[0039]

[0040] Therefore, the above formula is written as:

[0041]

[0042] So, for all samples, we can get:

[0043]

[0044] Calculate weight parameters Expectations:

[0045]

[0046] In actual situations, we do not know which data comes from which Gaussian distribution, nor do we know the parameters (mean and variance) of each Gaussian distribution. We can first initialize a set of parameters, so we can calculate γ ik The value of (i.e., the probability that the i-th data belongs to the k-th Gaussian distribution). Substitute the calculated weight into the log-likelihood function and calculate the parameter μ k ,Σ k Taking the partial derivative and setting it equal to zero gives:

[0047]

[0048] Finally, perform iterative processing.

[0049] The Gaussian mixture model is a very flexible model that can fit data distributions of various shapes because it is composed of multiple Gaussian distributions, each of which can capture a cluster in the data. The Gaussian mixture model provides the ability of soft clustering, that is, each data point is assigned to the probability of each cluster, rather than hard clustering to give a clear assignment. Because multiple Gaussian distributions can be used to model data, the Gaussian mixture model performs well when processing multimodal data and is suitable for distributions containing multiple peaks.

[0050] The Gaussian mixture model requires the number of clusters K to be specified in advance. At the same time, the computational complexity of the Gaussian mixture model is high in high-dimensional spaces and large data sets. The estimation of the covariance matrix involves matrix inversion, and for large-scale data sets, calculating the inverse of the covariance matrix may become very expensive. The Gaussian mixture model performs poorly on some irregular data sets, such as artificially synthesized noise circle data sets and bimonthly data sets. Figure 2-Figure 3 .

[0051] Based on the above content, this embodiment proposes a distribution network equipment risk assessment method based on entropy weight t-SNE and Gaussian mixture model, such as Figure 1 ,include:

[0052] Obtain a number of distribution network devices to be evaluated, where each distribution network device to be evaluated has n features, and calculate the weight of each feature;

[0053] The weight of each feature is added to the data of each feature to obtain a weighted data set, and the dimension of the weighted data set is reduced to obtain reduced-dimensional data;

[0054] Visualizing the dimensionality-reduced data to determine the range of the number of clusters;

[0055] A Gaussian mixture model is used to perform cluster analysis on the dimensionality reduction data, and based on the range of the number of clusters, the silhouette coefficient and the Calinski-Harabasz score are used to determine the optimal number of clusters to obtain a risk assessment result.

[0056] Specifically, assuming that the distribution network equipment dataset X to be evaluated has m samples and each sample has n features, then:

[0057]

[0058] Furthermore, the entropy weight method is used to calculate the weight of each feature, specifically:

[0059] W=[W1,W2,…,W n ] T ;

[0060] Where W is the weight of each feature and T is the transposed symbol.

[0061] The weighted data set is obtained as:

[0062] Z = X × diag(W);

[0063] Where Z is the weighted dataset, X is the distribution network equipment dataset to be evaluated, and W is the weight of each feature.

[0064] Furthermore, the dimensionality reduction of the weighted dataset is performed, including:

[0065] The dimensionality of the weighted dataset is reduced by using the t-distributed random nearest neighbor embedding dimensionality reduction method, wherein the t-distributed random nearest neighbor embedding dimensionality reduction method maps data points to a probability distribution through an affine transformation, converts the Euclidean distance into a conditional probability expression of the similarity between points, and uses the KL divergence to optimize the distance between probability distributions.

[0066] Traditional t-SNE dimensionality reduction methods don't account for differences in feature weights. Therefore, t-SNE may not be effective in datasets with large variations in feature weights. To address this issue, a t-SNE algorithm combining entropy weighting is proposed. This method first calculates and assigns feature weights using entropy weighting, incorporating this weight information into the data. t-SNE is then used to reduce the dimensionality of the weighted data.

[0067] The t-distributed random neighbor embedding dimensionality reduction method includes:

[0068] Construct a probability distribution between high-dimensional objects so that similar objects have a higher probability of being selected, while dissimilar objects have a lower probability of being selected;

[0069] Random neighbor embedding constructs the probability distribution of these points in a low-dimensional space, making the two probability distributions as similar as possible.

[0070] Specifically, linear dimensionality reduction methods are very powerful, but they often overlook important nonlinear structures in the data. In machine learning, a manifold refers to a low-dimensional subspace embedded in a high-dimensional data space. Manifold learning achieves dimensionality reduction by mining the intrinsic structure of the data, thereby finding a low-dimensional embedding manifold corresponding to the high-dimensional original data. Compared with principal component analysis, manifolds can be linear, but are more often nonlinear. For this reason, manifold learning is often regarded as a representative of nonlinear dimensionality reduction methods. It not only alleviates the effects of the curse of dimensionality, but also has stronger feature expression capabilities than linear dimensionality reduction methods. In addition to being nonlinear, manifold learning methods are generally nonparametric, which allows the manifold to more freely represent the inherent dimensionality and clustering characteristics of the data, but also makes it more sensitive to noise.

[0071] The t-Distributed Stochastic Neighbor Embedding (t-SNE) dimensionality reduction algorithm is developed from the t-SNE. SNE maps data points to a probability distribution through affine transformation. It mainly includes two steps: (1) SNE constructs a probability distribution between high-dimensional objects, so that similar objects have a higher probability of being selected, while dissimilar objects have a lower probability of being selected; (2) SNE constructs the probability distribution of these points in the low-dimensional space, so that the two probability distributions are as similar as possible.

[0072] SNE converts Euclidean distance into conditional probability to express the similarity between points. Suppose there is a data point x i and x j , the similarity between them is expressed as:

[0073]

[0074] Among them, σ i is a parameter, and its value is different for different data points. Because this embodiment focuses on the similarity between two data points, let p i|i =0.

[0075] For a data point y in the low-dimensional space i and y j , the variance of the Gaussian distribution can be specified as So the similarity between them is as follows (same setting q i|i =0):

[0076]

[0077] If the dimensionality reduction effect is good and the local features are preserved intact, then q j|i =p j|i , so we optimize the distance between the two probability distributions, namely KL divergence (Kullback-Leibler divergences), so the objective function (cost function) is as follows:

[0078]

[0079] Note that different points have different σ i , data point x i The entropy of the similarity distribution will increase with σ i SNE uses the concept of perplexity and then uses binary search to find the best σ i, where perplexity is:

[0080]

[0081] The perplexity can be interpreted as the number of valid neighboring points near a point. SNE is relatively robust to the adjustment of the perplexity. In this embodiment, it is usually selected between 5 and 50.

[0082] Although SNE provides a good visualization method, it is difficult to optimize and suffers from congestion problems. Therefore, Hinton et al. proposed the t-SNE method, which differs from SNE in the following ways: (1) It uses a symmetric version of SNE to simplify the gradient formula; (2) It uses the t-distribution instead of the Gaussian distribution to express the similarity between two points in low-dimensional space.

[0083] Improved p j|i and q j|i One way to replace the KL divergence is to use the joint probability distribution to replace the conditional probability distribution, that is, P is the joint probability distribution of each point in the high-dimensional space, Q is the joint probability distribution in the low-dimensional space, and the objective function is:

[0084]

[0085] in:

[0086]

[0087] This expression makes the overall expression much simpler, but it will introduce the problem of outliers. To solve this problem, the definition of the joint probability distribution in high-dimensional space is modified to:

[0088]

[0089] Avoid the excessive influence of outliers.

[0090] The crowding problem is that clusters are clumped together and cannot be distinguished. To alleviate the crowding problem, in low-dimensional space, a method that emphasizes long-tail distribution is used to convert distances into probability distributions, so that low and medium distances in high dimensions can be mapped to larger distances.

[0091] Comparing the probability density functions of Gaussian distribution and t-distribution, we can see that t-distribution has a longer tail, so the joint distribution in low-dimensional space is modified as follows:

[0092]

[0093] In this embodiment, four low-dimensional datasets, four medium-dimensional datasets, and four high-dimensional datasets were selected for comparative experiments. Datasets can be classified according to the number of features. Generally speaking, datasets with 0 to 19 features are considered low-dimensional datasets, 20 to 49 are considered medium-dimensional datasets, and more than 50 are considered high-dimensional datasets. To ensure the representativeness of the dataset, 12 datasets (four of each type) were selected from UCI as research objects. Table 1 shows the names, number of samples, and number of features of these datasets.

[0094] Table 1

[0095]

[0096]

[0097] We use t-SNE combined with entropy weight and traditional t-SNE to perform dimensionality reduction on different data sets, and use KL divergence as a measurement standard to obtain the KL divergence comparison values of the two methods in different dimensions. For easy observation, we use the dimensionality reduction dimension as the x-axis and the KL divergence value as the y-axis for visualization. First, the performance of the two methods on low-dimensional data sets is as follows: Figure 4 As shown in (a)-(d), it is observed that the two curves intersect at a point in the figure. In order to more objectively compare the advantages and disadvantages of the two methods, the area under the two curves (AUC) is calculated.

[0098] Depend on Figure 4 As shown in (a)-(d), the KL divergence values of the two methods at different dimensions show that t-SNE with entropy weighting outperforms traditional t-SNE on the Iris and Raisin datasets. For the Wine dataset, when the dimensionality reduction is 10, the KL divergence value of t-SNE with entropy weighting is greater than that of traditional t-SNE. Further comparison of the AUCs of the two methods shows that the AUC of t-SNE with entropy weighting is 5.7848, which is significantly lower than the AUC of 11.2656 for traditional t-SNE. For the Statlog (Heart) dataset, the KL divergence values of t-SNE with entropy weighting are greater than those of traditional t-SNE when the dimensionality reduction is 12 and 13. Further comparison of the AUCs of the two methods shows that the AUC of t-SNE with entropy weighting is 8.2397, which is lower than the AUC of 11.9927 for traditional t-SNE. Therefore, t-SNE with entropy weighting outperforms traditional t-SNE on low-dimensional datasets.

[0099] Medium-dimensional and high-dimensional datasets have a large number of features. When the number of features does not exceed 30, the KL divergence values of the two methods at each dimensionality reduction dimension are described in the figure. When the number of features exceeds 30, considering the computational overhead, the maximum dimensionality reduction dimension of the two methods is set to 30. Figure 5 (a)-(d) show the performance of t-SNE combined with entropy weight and traditional t-SNE on medium-dimensional datasets.

[0100] Depend on Figure 5 As shown in (a)-(d), for the Breast cancer and Infrard Thermography Temperatures datasets, the KL divergence values of t-SNE with entropy weighting are smaller than those of traditional t-SNE in all dimensions. For the Dermatology dataset, the KL divergence value of t-SNE with entropy weighting is larger than that of traditional t-SNE only when the dimensionality reduction dimension is 30. Further comparison of the AUCs of the two methods shows that the AUC of t-SNE with entropy weighting is 25.2991, which is lower than the AUC of 32.7028 for traditional t-SNE. For the Horse Colic dataset, the curves of the two methods intersect when the dimensionality reduction dimensions are 14, 23, and 27. Further comparison of the AUCs of the two methods shows that the AUC of t-SNE with entropy weighting is 239.5601, which is lower than the AUC of 253.4995 for traditional t-SNE. Therefore, t-SNE with entropy weighting outperforms traditional t-SNE on medium-dimensional datasets.

[0101] For high-dimensional data sets, the number of features is greater than 49. Considering the time cost of the experiment, we only compare the KL divergence values of the two methods when the dimension reduction dimension does not exceed 30. For example, Figure 6 As shown in (a)-(d).

[0102] Depend on Figure 6As shown in (a)-(d), for the Lung Cancer, Mice Protein Expression, and Spambase datasets, the KL divergence values of t-SNE with entropy weighting are lower than those of traditional t-SNE in all dimensions. For the Lung Cancer and Mice Protein Expression datasets, the difference in KL divergence between the two methods is small when the dimensionality reduction dimension does not exceed 5. However, as the dimensionality reduction dimension increases, the advantage of t-SNE with entropy weighting gradually increases. For the ConnectionistBench dataset, the KL divergence values of t-SNE with entropy weighting are higher than those of traditional t-SNE when the dimensionality reduction dimension is 8, 9, and 10. Further comparison of the AUCs of the two methods shows that the AUC of t-SNE with entropy weighting is 461.2166, which is lower than the AUC of 494.2296 of traditional t-SNE. Therefore, t-SNE with entropy weighting outperforms traditional t-SNE on high-dimensional datasets.

[0103] Furthermore, determining the range of the number of clusters includes:

[0104] The dimensionality-reduced data is visualized using Matplotlib, and the range of the number of clusters is determined based on the distribution of data points in the visualization graph.

[0105] The Gaussian mixture model is trained, comprising:

[0106] Initialize the dimensionality reduction data training set and obtain the probability that the i-th data belongs to the k-th Gaussian distribution;

[0107] Substitute the calculated weights into the log-likelihood function, find the partial derivative of the parameters and set them equal to zero, and finally perform iterative processing to obtain the trained Gaussian mixture model;

[0108] The parameters include:

[0109]

[0110] Where μ k ,Σ k are the parameters of the model, γ ik For the condition x i The hidden variable z ik The conditional probability of x i is the value of the i-th sample, and m is the number of samples.

[0111] This embodiment also proposes a distribution network equipment risk assessment system based on entropy-weighted t-SNE and Gaussian mixture model, including:

[0112] Data acquisition unit: used to acquire a number of distribution network devices to be evaluated, wherein each distribution network device to be evaluated has n features and calculate the weight of each feature;

[0113] Dimensionality reduction processing unit: used for adding the weight of each feature to the data of each feature to obtain a weighted data set, and performing dimensionality reduction on the weighted data set to obtain reduced-dimensionality data;

[0114] Visualization unit: used for visualizing the dimensionality reduction data and determining the range of the number of clusters;

[0115] The risk assessment unit is used to perform cluster analysis on the dimensionality reduction data using a Gaussian mixture model, determine the optimal number of clusters based on the range of the number of clusters, and obtain a risk assessment result using the silhouette coefficient and the Calinski-Harabasz score.

[0116] The proposed model effectively avoids the problem of weight selection affecting the evaluation results of other multi-criteria comprehensive evaluation methods, and does not suffer from the order reversal problem that is common in methods such as TOPSIS and MARCOS. This method is more suitable for risk assessment of distribution network equipment with a large number of network points to be tested. If more data can be collected in future research, the model performance will be further optimized and improved. Therefore, the distribution network equipment risk assessment model combining entropy-weighted t-SNE and Gaussian mixture model has good development prospects.

[0117] In order to more clearly express the technical solution of the present invention, the following specific embodiments are provided to introduce the solution:

[0118] The selection of risk assessment indicators for distribution network equipment varies depending on the research object. For risk assessment of rural distribution network equipment, after consulting relevant literature, we selected six characteristics to form the risk assessment indicator system: the degree of new load (I1), equipment load rate (I2), equipment obsolescence (I3), power supply load importance (I4), power supply radius exceeding the limit (I5), and voltage drop exceeding the limit (I6). Using a tiered combined scoring method, the severity of equipment risk is represented by 0, 2, 4, 6, 8, and 10, respectively, where higher scores indicate more severe equipment risk.

[0119] A case study was conducted on 15 66kV substation main transformers and 36 10kV lines in a rural power grid in a certain area of Northeast China that were to be upgraded and renovated. The risk data of each line are shown in Table 2.

[0120] Table 2

[0121]

[0122]

[0123] The entropy weight method is used to calculate the weight of each feature:

[0124] W=[0.2078,0.1774,0.0115,0.0395,0.2837,0.2800];

[0125] The entropy weighted t-SNE is used to reduce the dimension of the evaluation data and visualize it. Figure 7 It can be seen that the routes to be evaluated can be clustered into 4, 5, 6, or 7 categories. The Gaussian mixture model is further used to cluster the data after dimensionality reduction, and the silhouette coefficient and Calinski-Harabasz score are used to determine that the optimal number of clusters is 5. The clustering results are shown in Table 3.

[0126] Table 3

[0127]

[0128]

[0129] It can be seen from Table 3 that the distribution network equipment risk assessment model based on entropy weight t-SNE and Gaussian mixture model can effectively and reasonably assess the risk of distribution network equipment.

[0130] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A distribution network equipment risk assessment method based on entropy weighted t-SNE and Gaussian mixture model, characterized by: include: Obtain a number of distribution network devices to be evaluated, where each distribution network device to be evaluated has n features, and calculate the weight of each feature; The weight of each feature is added to the data of each feature to obtain a weighted data set, and the dimension of the weighted data set is reduced to obtain reduced-dimensional data; Visualizing the dimensionality-reduced data to determine the range of the number of clusters; A Gaussian mixture model is used to perform cluster analysis on the dimensionality reduction data, and based on the range of the number of clusters, the silhouette coefficient and the Calinski-Harabasz score are used to determine the optimal number of clusters to obtain a risk assessment result.

2. The risk assessment method according to claim 1, characterized in that: The weight of each feature is calculated by the entropy weight method, specifically: W=[W1,W2,…,W n ] T ; Where W is the weight of each feature and T is the transposed symbol.

3. The risk assessment method according to claim 2, characterized in that: The weighted data set is obtained as: Z = X × diag(W); Where Z is the weighted dataset, X is the distribution network equipment dataset to be evaluated, and W is the weight of each feature.

4. The risk assessment method according to claim 1, characterized in that: Performing dimensionality reduction on the weighted data set, comprising: The dimensionality of the weighted dataset is reduced by using the t-distributed random nearest neighbor embedding dimensionality reduction method, wherein the t-distributed random nearest neighbor embedding dimensionality reduction method maps data points to a probability distribution through an affine transformation, converts the Euclidean distance into a conditional probability expression of the similarity between points, and uses the KL divergence to optimize the distance between probability distributions.

5. The risk assessment method according to claim 4, characterized in that: Performing dimensionality reduction on the weighted data set includes: Calculate the KL divergence between the high-dimensional data distribution and the low-dimensional data distribution, and use the gradient descent method to iteratively solve the reduced-dimensional data.

6. The risk assessment method according to claim 1, characterized in that: Determining the range of the number of clusters includes: The dimensionality-reduced data is visualized using Matplotlib, and the range of the number of clusters is determined based on the distribution of data points in the visualization graph.

7. The risk assessment method according to claim 1, characterized in that: The Gaussian mixture model is trained, comprising: Initialize the dimensionality reduction data training set and obtain the probability that the i-th data belongs to the k-th Gaussian distribution; Substitute the calculated weights into the log-likelihood function, find the partial derivative of the parameters and set them equal to zero, and finally perform iterative processing to obtain the trained Gaussian mixture model; The parameters include: Where μ k ,Σ k are the parameters of the model, γ ik For the condition x i The hidden variable z ik The conditional probability of x i is the value of the i-th sample, and m is the number of samples.

8. A distribution network equipment risk assessment system based on entropy-weighted t-SNE and Gaussian mixture model, characterized by: include: Data acquisition unit: used to acquire a number of distribution network devices to be evaluated, wherein each distribution network device to be evaluated has n features and calculate the weight of each feature; Dimensionality reduction processing unit: used for adding the weight of each feature to the data of each feature to obtain a weighted data set, and performing dimensionality reduction on the weighted data set to obtain reduced-dimensionality data; Visualization unit: used for visualizing the dimensionality reduction data and determining the range of the number of clusters; The risk assessment unit is used to perform cluster analysis on the dimensionality reduction data using a Gaussian mixture model, determine the optimal number of clusters based on the range of the number of clusters, and obtain a risk assessment result using the silhouette coefficient and the Calinski-Harabasz score.