A method for classifying voltage quality of a transformer area based on nested K-means clustering

By processing distribution transformer voltage data using nested K-means clustering and Pearson correlation matrix, the accuracy problem of voltage quality analysis in distribution areas was solved, enabling refined classification and anomaly identification of voltage quality in distribution areas, and providing a scientific governance solution.

CN117290746BActive Publication Date: 2026-01-02STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311413449.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2026-01-02
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

Existing technologies lack effective clustering methods in voltage quality analysis of distribution areas, which makes it impossible to accurately describe the voltage quality characteristics of the entire distribution area. In particular, research on the construction of power consumption profiles for distribution areas is still incomplete.

Method used

A nested K-means clustering method is adopted to process distribution transformer voltage data through primary and secondary clustering, combined with the Pearson correlation matrix, to gradually refine the voltage quality classification. Newton interpolation is used to handle missing data, Euclidean distance is used to analyze the limit situation, and principal component analysis is used to reduce dimensionality.

Benefits of technology

It improves the accuracy of voltage quality classification in distribution areas, enables rapid identification of distribution transformers with abnormal voltage, provides scientific solutions for voltage management in distribution areas, and overcomes the shortcomings of traditional European distance in terms of highly correlated data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117290746B_ABST
    Figure CN117290746B_ABST
Patent Text Reader

Abstract

The application relates to a transformer area voltage quality classification method based on nested K-means clustering, which comprises the following steps: obtaining voltage data of each transformer and performing pretreatment; based on a K-means clustering method, the pretreated transformer voltage data is initially clustered to obtain an initial classification result; based on the initial classification result, the out-of-limit condition of the transformer voltage is analyzed; if it is out of limit, the K-means clustering method is used for secondary clustering; if it is not out of limit, the initial classification result is output as the final clustering result; the clustering result of the secondary clustering is analyzed for distinctness, whether the distinctness is obvious is judged, if yes, the clustering result of the secondary clustering is output as the final clustering result, if not, the secondary clustering and the distinctness analysis process are repeated until the final clustering result is output, and a transformer area voltage quality classification result is obtained. Compared with the prior art, the application has the advantages of improving the transformer area voltage quality classification accuracy and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid safety, and particularly relates to a transformer area voltage quality classification method based on nested K-means clustering. BACKGROUND

[0002] The advantages and disadvantages of voltage quality are important indicators of power quality, which will directly endanger the safe and economic operation of the power grid. There is a great difference between the voltage distribution of the transformer area and the medium voltage distribution network. The voltage loss of the medium voltage distribution network is small, and the node voltage distribution curve is approximately a smooth curve. However, the voltage deviation of the first and last nodes of the transformer area is often large, up to 20% to 30% of the rated voltage, and the node voltage distribution curve often shows a rapid decline characteristic. There is also a problem of exceeding the upper limit of the first-end voltage and the lower limit of the last-end voltage. At present, the research on electricity portrait analysis focuses on the analysis of user-level electricity consumption characteristics, but the research on low-voltage transformer area electricity consumption characteristics based on clustering is relatively less, especially the research on the construction of transformer area electricity portrait is still in the preliminary stage. Wang et al. in the document Association rule mining based quantitative analysis approach of household characteristics impacts on residential electricity consumption patterns (Association rule mining based quantitative analysis approach of household characteristics impacts on residential electricity consumption patterns) used the density-based clustering algorithm DBSCAN to extract the seasonal typical electricity consumption mode of each user, and used the K-means clustering algorithm to cluster the electricity consumption mode. Finally, the association rule mining algorithm was used to explore the potential relationship between the residential electricity consumption mode and the user portrait factors. Shanshan et al. in the document Method of Low Voltage Transformer Area Electricity Portrait Based on Clustering (Method of Low Voltage Transformer Area Electricity Portrait Based on Clustering) extracted the transformer area feature tags through various methods based on transformer distribution data, formed a transformer area tag system, and constructed the transformer area portrait under the condition that the number of transformer area electricity consumption modes was unknown. Alie et al. in the document Intelligent Transformer Identification Technology Based on Data Spatio-Temporal Correlation (Intelligent Transformer Identification Technology Based on Data Spatio-Temporal Correlation) proposed to cluster the transformer area electricity data and analyze the relationship between the transformer area static information and the electricity information to determine the electricity mode of the unmonitored transformer area. Li et al. in the document Development of Low Voltage Network Templates—Part I: Substation Clustering and Classification (Development of Low Voltage Network Templates—Part I: Substation Clustering and Classification) established the transformer area fingerprint and the transformer area health system, but the tag system used to describe the transformer area electricity characteristics was not perfect.

[0003] The current research on voltage quality analysis mainly focuses on the user side, and the portrait analysis of the user side as a hot direction cannot accurately describe the voltage quality characteristics of the entire transformer area because it does not combine the voltage characteristics of the upper transformer area. SUMMARY

[0004] The purpose of the present application is to provide a transformer area voltage quality classification method based on nested K-means clustering to improve classification accuracy.

[0005] The purpose of the present application can be achieved by the following technical solutions:

[0006] A transformer area voltage quality classification method based on nested K-means clustering, comprising the following steps:

[0007] Obtain the voltage data of each transformer and perform preprocessing;

[0008] Based on the K-means clustering method, the preprocessed transformer voltage data is initially clustered to obtain an initial classification result;

[0009] Based on the initial classification result, analyze the out-of-limit situation of the transformer voltage, if it is out-of-limit, use the K-means clustering method for secondary clustering, if it is not out-of-limit, output the initial classification result as the final clustering result;

[0010] Perform distinctness analysis on the clustering result of the secondary clustering, and determine whether the distinctness is obvious, if yes, output the clustering result of the secondary clustering as the final clustering result, if no, repeat the secondary clustering and distinctness analysis process until the final clustering result is output, and obtain the transformer area voltage quality classification result.

[0011] Further, the preprocessing operation is:

[0012] The data exceeding the set missing ratio in the transformer voltage data is removed, and the data not exceeding the set missing ratio is processed by using an interpolation method.

[0013] Further, the interpolation method is Newton interpolation method, and the interpolation polynomial is:

[0014] f(x i )=f(x1)+f[x2,x1](x i -x1)+…+f[x n ,x n-1 ,…,x1](x i -x1)…(x i -x n-1 )

[0015] +f[x n ,x n-1x1+x i ](x i -x1)…(x i -x n )

[0016] The interpolation approximation function is:

[0017] N n (x i )=f(x1)+f[x2,x1](x i -x1)+…+f[x n ,x n-1 ,…,x1](x i -x1)…(x i -x n-1 )

[0018] The truncation error is:

[0019] R n (x i )=f[x n ,x n-1 ,…,x1+x i ](x i -x1)…(x i -x n )

[0020] In the formula, f(x i ) is the function value obtained by Newton interpolation, [x1, f(x1)], [x2, f(x2)],..., [x n , f(x n )] are n data collected by a data acquisition device, N n (x i ) is the interpolation approximation function, and R n (x i ) is the truncation error.

[0021] Further, the cluster categories of the initial clustering include more upper limit days, normal days, and "other" days.

[0022] Further, the loss function of the K-means clustering method is:

[0023]

[0024] In the formula, J is the loss function, M is the total number of samples, i.e., the number of data of the voltage of the data acquisition device, x i represents the i th sample, c i represents the cluster to which x i belongs, represents the center point corresponding to the c i cluster.

[0025] Further, before secondary clustering, the secondary clustering features are selected based on the out-of-limit condition, and principal component analysis is used for dimension reduction.

[0026] Further, the secondary clustering features are the maximum, minimum, median, average and out-of-limit rate of the voltage of each month.

[0027] Further, the primary classification result uses Euclidean distance analysis to analyze the out-of-limit condition of the distribution transformer voltage.

[0028] Further, before the repeated secondary clustering and distinctness analysis process, the correlation matrix is used to process the clustering results with unclear distinctness.

[0029] Further, the correlation matrix is a Pearson correlation matrix, and the expression is:

[0030]

[0031] In the formula, D r is the Pearson correlation matrix; r 12 is the correlation coefficient; x1=(x 11 , x 12 , …, x 1n ), and x2=(x 21 , x 22 , …, x 2n ) are any two points in an n-dimensional space. , and are the average values of the given points.

[0032] Compared with the prior art, the present application has the following beneficial effects:

[0033] (1) The present application nests the out-of-limit classification result based on the primary clustering result using the K-means clustering algorithm, repeatedly applies the K-means clustering algorithm at different levels, iteratively clusters the data multiple times, obtains a clustering result with high correlation, and thus effectively improves the accuracy of voltage quality classification.

[0034] (2) The present application uses an unsupervised learning clustering algorithm combined with Pearson correlation, fully considers the clustering index characteristics under different clustering feature conditions, uses the Pearson correlation matrix to combine the distribution transformers with similar voltage curves together, overcomes the problem that the K-means clustering algorithm based on traditional Euclidean distance has poor clustering effect when dealing with data with high correlation, and improves the rationality of the clustering scheme.

[0035] (3)The application can quickly analyze voltage quality by processing out-of-limit data through multiple clustering processes, identify abnormal distribution transformers in a transformer area, classify a large number of distribution transformers with common characteristics according to abnormal indexes, and lay a scientific data foundation for providing a comprehensive solution to transformer area voltage management through targeted research on the divided categories. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 A flowchart of the method of the application is shown in the figure.

[0037] Figure 2 A schematic diagram of the clustering process of the embodiment of the application is shown in the figure.

[0038] Figure 3 The maximum monthly voltage of the first clustering 3-class secondary clustering 1-class of the embodiment of the application is shown in the figure.

[0039] Figure 4 The maximum monthly voltage of the first clustering 3-class secondary clustering 2-class of the embodiment of the application is shown in the figure.

[0040] Figure 5 The maximum monthly voltage of the first clustering 3-class secondary clustering 3-class of the embodiment of the application is shown in the figure.

[0041] Figure 6 The maximum monthly voltage of the first clustering 3-class secondary clustering 4-class of the embodiment of the application is shown in the figure.

[0042] Figure 7 The maximum monthly voltage of the first clustering 3-class secondary clustering 3-class tertiary clustering 1-class of the embodiment of the application is shown in the figure.

[0043] Figure 8 The maximum monthly voltage of the first clustering 3-class secondary clustering 3-class tertiary clustering 2-class of the embodiment of the application is shown in the figure.

[0044] Figure 9 The maximum monthly voltage of the first clustering 3-class secondary clustering 3-class tertiary clustering 3-class of the embodiment of the application is shown in the figure.

[0045] Figure 10 The maximum monthly voltage of the first clustering 3-class secondary clustering 3-class tertiary clustering 4-class of the embodiment of the application is shown in the figure. DETAILED DESCRIPTION

[0046] The application will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments are implemented on the basis of the technical solution of the application, and detailed implementation methods and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.

[0047] The embodiment provides a transformer area voltage quality classification method based on nested K-means clustering, as shown in the figure. Figure 1 The method comprises the following steps:

[0048] S1, acquire voltage data of each distribution transformer and pre-process.

[0049] The embodiment adopts voltage data collected by Shanghai Power Supply Company with 15 minutes as a collection cycle. In order to filter out noise in initial data, data containing more than 10% missing data in voltage data of distribution transformers in Shanghai urban area are rejected, and data containing less than 10% missing data are processed by interpolation. Common interpolation methods include linear interpolation, Lagrange interpolation and Newton interpolation, etc. In order to reduce the number of operations, the embodiment selects Newton interpolation method, and the interpolation polynomial of Newton interpolation method can be written as follows:

[0050] f(x i )=f(x1)+f[x2,x1](x i -x1)+…+f[x n ,x n-1 ,…,x1](x i -x1)…(x i -x n-1 )

[0051] +f[x n ,x n-1 ,…,x1+x i ](x i -x1)…(x i -x n )

[0052] In the formula, [x1, f(x1)], [x2, f(x2)],..., [x n , f(x n )] are n data collected by a distribution transformer, and f(x i ) is a function value obtained by Newton interpolation.

[0053] The interpolation approximation function N n (x i ) is:

[0054] N n (x i )=f(x1)+f[x2,x1](x i -x1)+…+f[x n ,x n-1 ,…,x1](x i -x1)…(x i -x n-1 )

[0055] The truncation error R n (x i ) is:

[0056] R n (x i )=f[x n ,x n-1 ,…,x1+x i ](x i -x1)…(x i -x n )

[0057] S2, based on the K-means clustering method, the pre-processed distribution transformer voltage data is initially clustered to obtain an initial classification result.

[0058] K-means algorithm (K-means) is a kind of unsupervised learning algorithm for clustering data in machine learning. The algorithm has the advantages of simple principle, strong interpretability, easy implementation, fast convergence speed, etc., and the effect is remarkable when dealing with huge data sets. Its basic principle can be explained as follows: through a loop, the class center point is iterated constantly, the distance of each object to the new class center point is calculated, and the objects are reclassified according to the nearest distance principle. When the intra-class distance is minimum and the inter-class distance is maximum, the iteration can be stopped. Physically, it belongs to the clustering algorithm based on sample distance division.

[0059] According to the Power Quality Supply Voltage Allowable Deviation GB 12325-2008, the voltage deviation is in the normal range of-10%~+7%, and exceeds 235.4V (voltage deviation exceeds 7%) as the upper limit, and is lower than 198V (voltage deviation exceeds-10%) as the lower limit. In the initial clustering, the days of exceeding the upper limit, the normal days and the "other" days are taken as the K-Means clustering categories. Since the number of distribution transformers exceeding the lower limit is very small, it is not taken as a clustering category, so as to avoid too many categories and the difference between categories is not obvious. As shown in FIG. 1, the K-means clustering method is used for initial clustering, and the initial classification result obtained is a plurality of subsets, i.e. a plurality of initial classification categories. Figure 2

[0060] S3, based on the initial classification result, the over-limit situation of the distribution transformer voltage is analyzed. If it is over-limit, the K-means clustering method is used for secondary clustering. If it is not over-limit, the initial classification result is output as the final clustering result.

[0061] As shown in FIG. 2, the K-means clustering method is used for secondary clustering, and the final clustering result is obtained. Figure 2 ​As shown, distance calculation is first used to analyze the voltage exceedance of distribution transformers, identifying the transformer categories with the most frequent exceedances within a year. Secondly, for these transformers, the maximum, minimum, median, average voltage, and exceedance rate for each month are selected as secondary clustering features. Since there are many features in the secondary clustering, Principal Component Analysis (PCA) is chosen to reduce the dimensionality of the dataset. PCA is a feature extraction method that uses orthogonal transformation to convert linearly correlated variables into a few linearly independent variables, significantly reducing noise and redundancy.

[0062] There are many methods for calculating distances, such as Minkowski distance, Manhattan distance, Euclidean distance, Mahalanobis distance, correlation coefficient, and cosine similarity. The most commonly used is Euclidean distance. Assume the sample X = {x1, x2, ..., x...} n}, Y = {y1, y2, ..., y n}, then the Euclidean distance between n-dimensional sample vectors can be expressed as:

[0063]

[0064] This embodiment starts with secondary clustering. Since the classification results of cases exceeding the upper limit obtained from the initial clustering are based on this, the main focus is on the voltage trend of the distribution transformers. Therefore, the maximum value of each month within a year is selected as the feature for secondary clustering. Voltage data for Shanghai's urban area in 2022 is missing for December. Therefore, the secondary clustering features ultimately selected in this embodiment are max_1, max_2, ..., max_11, which represent the maximum voltage values ​​from January to December, respectively.

[0065] The data after PCA dimensionality reduction is then subjected to secondary clustering using the K-means clustering algorithm to obtain the secondary clustering results.

[0066] S4. Perform discrimination analysis on the clustering results of the secondary clustering to determine whether the discrimination is significant. If yes, output the clustering results of the secondary clustering as the final clustering result. If not, repeat the secondary clustering and discrimination analysis process until the final clustering result is output to obtain the voltage quality classification result of the transformer area.

[0067] like Figure 2 As shown, this step mainly involves analyzing each subclass after obtaining the results of the secondary clustering, comparing the distance or similarity between the cluster centers of different subclasses. If the distance between the cluster centers of different subclasses is large or the similarity is low, it indicates that the discrimination is significant; otherwise, the discrimination is not significant. If there are categories with low discrimination, the Pearson correlation matrix will be used to replace the original data for clustering.

[0068] Pearson correlation coefficient method is used to detect the degree of linear correlation between two continuous random variables, the value range is [-1, 1], positive value represents positive correlation, negative value represents negative correlation, the greater the absolute value represents the higher degree of linear correlation. In mathematics, Pearson correlation coefficient can be expressed as the quotient of the covariance and the standard deviation of two variables. For any given two points x 11 , x 12 , …, x 1n ), x2=(x 21 , x 22 , …, x 2n ), the correlation coefficient can be expressed as:

[0069]

[0070] Wherein:

[0071]

[0072] According to the above expression of correlation coefficient, the correlation matrix formula can be defined as follows:

[0073]

[0074] In the formula, are the average values of the given points.

[0075] The loss function of the K-means algorithm used in the above embodiment can be defined as:

[0076]

[0077] Wherein, M is the data set of the total number of samples, x i represents the i-th sample, c i represents the cluster to which x i belongs, represents the center point corresponding to the ci cluster. That is, the loss function can be defined as the sum of the error squares of each sample to the center point of the cluster to which it belongs. Through iterative calculation, when the loss function J is no longer changed or convergent, the iteration is stopped.

[0078] According to the clustering process as shown in Figure 2 , the voltage out-of-limit statistical data is clustered, and the best clustering category in the initial clustering is determined to be three according to the Silhouette Coefficient (SC). The initial clustering results are as shown in Table 1:

[0079] Table 1 Initial clustering results

[0080]

[0081] It can be seen that in the initial clustering, the average normal days of class 1 is 315 days, which is a normal transformer category, class 2 is a category in which other days account for the majority, and the average over-limit days of class 3 is 230.9 days. In order to further analyze the voltage over-limit situation of the over-limit transformer, it is necessary to perform secondary clustering. The monthly voltage eigenvalues of 982 transformers in class 3 are selected, and after dimension reduction using PCA, secondary clustering is performed. The SC determines that the optimal number of clusters is 4 classes.

[0082] The secondary clustering results are shown in Figures 3 to 6 , where Figure 3 is the result of secondary clustering class 1, the number of transformers in this class is 339, and the transformer voltage over-limit interval mainly concentrates between 240V and 250V, with a large span but uniform distribution, and the overall fluctuation trend is not large, and the overall monthly voltage maximum deviation is large; the number of transformers in secondary clustering class 2 is the largest, reaching 515, as shown in Figure 4 It can be seen that the monthly voltage maximum of this class of transformers is distributed between 235V and 245V, and there are two relatively obvious decreases in April and May, while the overall monthly voltage maximum trend is relatively stable; according to Figure 5 , we can know that the monthly voltage maximum of the transformers belonging to this class has a large span and changes significantly in the first half and the second half of the year, with an average of about 255V in the first half and a decrease to the interval of 235V to 240V in the second half, so this class needs to be further subdivided and studied, and the number of transformers in this class is 48, accounting for a small proportion; Figure 6 is the transformer in secondary clustering class 4, it can be seen from the figure that this class of transformers has the most serious over-limit and little fluctuation, generally maintaining between 250V and 260V, with the largest deviation among all classes, and the number of transformers in this class is 80. Overall, the distance between classes in the secondary clustering results is small, and the clustering effect is obvious.

[0083] In the above analysis, the K-means algorithm based on Euclidean distance is used for secondary clustering of the class with more over-limit transformers in the initial clustering, and the subdivision results of the over-limit transformers are obtained. When analyzing the voltage over-limit situation of the transformers in each class of secondary clustering, we find that the transformers in secondary clustering class 3 generally have a large difference between the first half and the second half of the year, but the decrease in the second half is not sufficiently distinguished among the classes, so we choose to use K-means based on Pearson similarity matrix for third clustering to obtain more accurate clustering results about the voltage curve fluctuation of the transformers. First, the monthly voltage maximum of the transformers in secondary clustering class 3 is extracted, then the Pearson similarity matrix formula is solved, and finally the K-means algorithm is used to cluster the similarity matrix. After clustering, the clustering results are saved to the original table, and the third clustering results are as shown in Figures 7 to 10as shown.

[0084] The results of the third clustering are shown in Table 3. Figures 7 to 10 As shown in Table 3, the distribution transformers are subdivided into four categories by using the K-means algorithm based on the Pearson similarity matrix to cluster the category 3 in the second clustering, which has a larger difference between the maximum voltage in the first half year and the second half year. Figure 7 The results of the first clustering show that the voltage out-of-limit interval of the distribution transformers in this category is between 255 V and 260 V, and the voltage of several distribution transformers decreases temporarily in July and August. Figure 8 The results of the second clustering show that most of the distribution transformers in this category have a voltage drop in June, and there is no other type of distribution transformer mixed in. Figure 9 The results of the third clustering show that the voltage of some distribution transformers in this category is not high in January and February, but jumps to the average level of the third category in March, and the out-of-limit interval remains around 255 V, and the voltage drops to around 240 V in September and October. Figure 10 The results of the fourth clustering show that the voltage of the distribution transformers in this category is high in January and February, but falls in March and April.

[0085] The above clustering algorithm uses an unsupervised learning clustering algorithm combined with Pearson correlation, fully considers the characteristics of the clustering index under different clustering characteristics, and uses the Pearson correlation coefficient to combine the distribution transformers with similar voltage curves, which overcomes the problem of poor clustering effect of the K-means clustering algorithm based on traditional Euclidean distance when dealing with data with strong correlation, and improves the rationality of the clustering scheme. The above example analysis verifies that the nested K-means clustering algorithm based on the Pearson correlation coefficient method proposed in this embodiment has strong practicability in production practice in terms of the ability to find abnormal fluctuation distribution transformers according to the voltage out-of-limit condition of the distribution transformers in the study of the voltage out-of-limit problem of urban distribution networks.

[0086] If the above functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or partially contribute to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0087] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0088] The present application is described with reference to flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device that implements the flow Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows or blocks.

[0089] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction apparatus, which implements the flow Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows or blocks.

[0090] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operations steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide processes for implementing the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or steps of the functions specified in the flow

[0091] Although preferred embodiments of the application have been described, those skilled in the art will recognize that additional modifications and variations are possible in light of the above teachings. It is therefore intended that the appended claims be interpreted as including all such modifications and variations as fall within the scope of the application.

Claims

1. A method for classifying the quality of voltage in a transformer area based on nested K-means clustering, characterized by, The method comprises the following steps: Obtaining and preprocessing voltage data of each distribution transformer; Based on the K-means clustering method, the preprocessed voltage data of the distribution transformer is initially clustered to obtain an initial classification result, and the initial clustering cluster categories include over-limit days, normal days and other days; Based on the initial classification result, the over-limit situation of the voltage of the distribution transformer is analyzed, if it is over-limit, the K-means clustering method is used for secondary clustering, if it is not over-limit, the initial classification result is output as the final clustering result; The secondary clustering result is analyzed for distinctness, if it is obvious, the secondary clustering result is output as the final clustering result, if it is not obvious, the secondary clustering and distinctness analysis process are repeated until the final clustering result is output, and a transformer voltage quality classification result is obtained; Before repeating the secondary clustering and distinctness analysis process, the clustering result with indistinctness is processed by using a correlation matrix, the correlation matrix is a Pearson correlation matrix, and the expression is: wherein is the Pearson correlation matrix; is the correlation coefficient; , is n any two points in the d-dimensional space; , are the mean values of the given points, respectively.

2. The method according to claim 1, wherein, The preprocessing operation is: Data exceeding a set missing proportion in the voltage data of the distribution transformer is removed, and interpolation method is used to process data not exceeding the set missing proportion.

3. The method of claim 2, wherein the method is characterized by, The interpolation method is Newton interpolation method, and the interpolation polynomial is: The interpolation approximation function is: The truncation error is: wherein is the function value obtained by Newton interpolation, is a piecewise collected n data, is the interpolation approximation function, is the truncation error.

4. The method of claim 1, wherein the method is characterized by, The loss function of the K-means clustering method is: In the formula, For loss function, M The total number of samples, i.e., the number of distribution transformer voltage data. Representing the One sample, Represented as The cluster to which it belongs Then it means The center point corresponding to the cluster.

5. The method of claim 1, wherein the method is characterized by, Before secondary clustering, secondary clustering features are selected based on the over-limit situation, and principal component analysis is used for dimension reduction.

6. The method according to claim 5, wherein, The secondary clustering features are the maximum value, the minimum value, the median, the average value and the over-limit rate of the voltage of the distribution transformer in each month.

7. The method of claim 1, wherein the method is based on nested K-means clustering. The initial classification result is analyzed for the over-limit situation of the voltage of the distribution transformer by using the Euclidean distance.