Power distribution network line loss clustering characteristic analysis and anomaly identification method

By introducing periodic oscillation mutation strategies and the best point set principle in the snow ablation algorithm, the initial clustering center is optimized, and dynamically adjusting the clustering center is solved, the existing technology has poor effect when processing non-uniform or sparse data, and more efficient line loss clustering and abnormal recognition are achieved.

CN120011844APending Publication Date: 2025-05-16NANJING INST OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510085232.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art uses poor snow ablation algorithms when processing non-uniform or sparse data, and the calculation efficiency and convergence speed are insufficient in large-scale problems, making it difficult to apply to large-scale power supply and distribution system operation and maintenance scenarios.

Method used

The initial clustering center is optimized by introducing periodic oscillation mutation strategies and the best point set principle, and dynamically adjusting the clustering center with the improved snow ablation algorithm, improving the convergence and accuracy of the clustering process. At the same time, through adaptive adjustment of weights, the influence of anomalies in clustering analysis is enhanced.

Benefits of technology

It improves the linear loss clustering effect and the accuracy of identifying abnormal line loss, and is suitable for dealing with abnormal line loss caused by line topology abnormalities and load heavy loads in the distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011844A_ABST
    Figure CN120011844A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a power distribution network line loss clustering characteristic analysis and anomaly identification method, and relates to the technical field of power supply and distribution system operation and maintenance. The method comprises the following steps: collecting multi-time-period line loss data of a power distribution network line, and constructing a line loss calculation model; zero mean value unit variance standardization processing is adopted to balance each feature so as to avoid the situation that the clustering process is dominated due to the fact that the factor value range of some features (such as current) is large. Initial clustering center selection is carried out by introducing a periodic oscillation mutation strategy and a good point set principle, and the rationality and clustering effect of an initial clustering result are optimized. In the clustering center optimization process, dynamic adjustment is carried out in combination with an improved snow ablation algorithm, so that the convergence and accuracy of the algorithm are improved. And researching the dynamic change and distribution rule of the line loss data in time and space dimensions through space-time analysis, and carrying out anomaly identification on the line loss data based on a clustering result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power supply and distribution system operation and maintenance, and in particular to a method for analyzing the clustering characteristics of distribution network line losses and identifying anomalies. Background Art

[0002] In the power system, line loss is a key issue that directly affects the reliability and economy of power supply. Line loss not only leads to energy waste, but also affects the operating efficiency and stability of the power system. Therefore, in-depth analysis and clustering of historical line loss data to reveal its potential characteristics and changing laws are crucial for formulating effective loss reduction strategies and achieving efficient operation of the power system. At present, line loss anomaly identification is an important task in power system management, which mainly relies on advanced data analysis and clustering methods to identify and distinguish different types of line losses.

[0003] Common clustering methods include hierarchical clustering, density-based clustering and K-means clustering, among which K-means is widely used in various fields of power systems due to its high efficiency and scalability. However, the K-means algorithm has some shortcomings in practical applications, especially its sensitivity to the selection of initial center points, which may cause the clustering results to be affected by the sample set and the selection of initial center points, showing problems such as local optimality, slow convergence and unstable results. In order to solve these problems, a variety of optimization strategies have been proposed in the prior art. For example, the snow ablation algorithm (SAO) is a heuristic optimization algorithm that gradually shrinks the range of candidate center points by simulating the snow ablation process to increase the probability of finding the global optimal initial center. The algorithm uses density information to select the initial center point, effectively dealing with noise and outliers in the data set, and improving the stability and accuracy of the clustering results.

[0004] However, the standard snow melting algorithm does not work well when processing non-uniform or sparse data, and when facing large-scale problems, the algorithm's computational efficiency and convergence speed are insufficient, making it difficult to apply in large-scale power supply and distribution system operation and maintenance scenarios. Therefore, how to further improve the method of clustering characteristics analysis and abnormal identification of distribution network line losses, thereby improving the line loss clustering effect and the accuracy of identifying abnormal line losses, has become a topic that needs to be studied. Summary of the invention

[0005] The embodiments of the present invention provide a method for analyzing the clustering characteristics of distribution network line losses and identifying abnormalities, which can improve the line loss clustering effect and the accuracy of identifying abnormal line losses.

[0006] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0007] A distribution network line loss clustering characteristic analysis and anomaly identification method, comprising:

[0008] Collect the line loss data of the distribution network and build a line loss model. The line loss data includes the line loss value of the corresponding time period; the load data provided by the smart meter is used to infer the power loss of a certain section of the line; the acquisition system can also obtain line loss data; substations and distribution equipment are usually equipped with various electrical measuring instruments, such as current transformers (CT), voltage transformers (VT), active power and reactive power meters, which are used to measure the current, voltage, and power parameters of each link in the power grid. These devices monitor the operating status of the distribution network in real time and generate data related to line loss. The so-called "line loss value of the corresponding time period" can be understood as: a data obtained every 15 minutes through the line loss model calculation, or it can be specified by itself, such as line loss data for a few hours or a few days. Specifically, the line loss data of the distribution network can be collected, including the line loss values ​​of multiple time periods; then the line loss calculation model is constructed using characteristic parameters such as line current, voltage, area power factor, and user load; before clustering the line loss data, the cluster center of the line loss is optimized and adjusted, wherein the cluster center after optimization is used to screen the abnormal points of the line loss clustering and adaptively adjust the weights; cluster analysis is performed on the line loss data, and the line loss features of different categories are clustered for separation; the line loss features are used for spatiotemporal analysis to identify abnormal line loss data. Among them, the loss features include the combined influence of active line loss, reactive line loss, load, voltage, temperature and other factors, and these features are all corresponding to the formulas such as f(V) and g(L). The features are combined together in Ploss.

[0009] The line loss model includes: P loss represents line loss power, I represents current, R(T) represents resistance at temperature T, PF represents power factor, f(V) represents the influence function of voltage V on line loss, g(L) represents the influence function of load L on line loss. The above I, V, L specifically refer to the current, voltage and load flowing through the transformer in the substation area; where R(T) = R0(1+αT), R0 represents the resistance at the reference temperature, α represents the temperature coefficient of resistance, V0 represents the reference voltage, and L0 represents the reference load.

[0010] Before performing accurate cluster analysis on line loss data, feature parameter standardization processing is performed; the feature parameter standardization processing includes: x i ′ represents the i-th eigenvalue after standardization, x i represents the original i-th eigenvalue, i is the label of the line loss data collected from the distribution network line, μ i represents the mean of the i-th feature; N is the number of samples in the data set; σ i represents the standard deviation of the i-th feature, The characteristic parameters after standardization are recorded in the matrix X. n is the number of samples, p is the number of feature types, and the feature types include: current I, voltage V, power factor PF, temperature T, and load L. Each feature is standardized by zero-mean unit variance transformation to keep the modeling balanced and avoid certain features (such as current) from dominating the clustering due to their large numerical range.

[0011] The optimization and adjustment of the clustering center of the line loss is carried out, the initial center of the line loss clustering is selected, and the periodic oscillation mutation strategy and the principle of the best point set are introduced to improve the rationality of the initial clustering center and the overall clustering effect; then the line loss clustering center is optimized, and the initial center is dynamically adjusted using the improved snow melting algorithm to further enhance the convergence and accuracy of the algorithm; including: setting up the best point set Among them, {r s (n) ·k} means taking the decimal part, n is the number of points, k is a positive integer, s is the dimension of the Euclidean space, r is the best point, r∈M s , M s is a unit cube in s-dimensional Euclidean space; r = {2cos(2πk / p), 1≤k≤s}, p is the smallest prime number satisfying (pt)≥s; represents…,,x i represents the original i-th eigenvalue of the current solution, x rand represents other randomly selected solutions, β and ω are the amplitude and frequency parameters of the oscillation, respectively, and t is the current iteration number.

[0012] The optimized and adjusted cluster center includes: for a line loss data set selected by a good point set, calculating the Euclidean distance between two line loss data points in the line loss data set, and then obtaining the average value of the Euclidean distances of all line loss data points in the data set as the average distance; selecting one of the line loss data points, and querying the number of line loss samples covered within the average distance range around the line loss data as the density value; and taking the line loss data point with the largest density value as the initial cluster center.

[0013] The screening of line loss clustering abnormal points and adaptively adjusting weights include: calculating the outlier degree of all line loss data through the isolation forest model and the local anomaly factor algorithm, and reducing the weight of the abnormal points in the line loss clustering according to the outlier degree results, thereby reducing the impact of abnormal data on the clustering results. In the process of clustering different categories of line loss features for separation, it includes: drawing a curve of the SSE values ​​corresponding to different K values, and selecting the cluster number with a significantly slower SSE decline rate as the optimal K value according to the drawn curve, K = 1, 2, ..., K max , K max is the maximum number of clusters allowed.

[0014] Before using the line loss characteristics for spatiotemporal analysis, it includes: building a model for capturing the mean and fluctuation characteristics of line loss time series data: μ t represents the mean line loss at time point t, x i,t represents the line loss value of the ith line at time point t; where the parameter used to measure the fluctuation intensity is Construct line loss space dimension analysis model: L i represents the weighted line loss mean of area i; w ij represents the spatial weight between regions i and j, x j represents the line loss value of area j; m represents the number of areas involved in the calculation.

[0015] The identifying and obtaining abnormal line loss data includes: LOF(x i ) The LOF value output is used to determine the sample point x i Is it an indicator parameter of an abnormal point? If the LOF value is greater than the set threshold, it is judged as abnormal line loss data. i ,x j ) is the sample point x i With x j The reachable distance, N k (x i ) is the sample point x i k-nearest neighbors, lrd() represents the calculation function of local reachable density.

[0016] The embodiment of the present invention proposes a method for line loss clustering characteristic analysis and anomaly identification. First, multi-time period line loss data of distribution network lines are collected, and a line loss calculation model is constructed using characteristic parameters such as current, voltage, area power factor, and user load. Then, zero mean unit variance normalization is used to balance the features to avoid certain features (such as current) dominating the clustering process due to a large numerical range. Next, the initial clustering center is selected by introducing a periodic oscillation mutation strategy and a good point set principle to optimize the rationality and clustering effect of the initial clustering results. In the process of cluster center optimization, dynamic adjustment is performed in combination with an improved snow melting algorithm to improve the convergence and accuracy of the algorithm. For identifying small-scale abnormal line loss data, the method adaptively increases the weight of abnormal points to enhance their influence in cluster analysis. Finally, the dynamic changes and distribution laws of line loss data in the time and space dimensions are studied through spatiotemporal analysis, and abnormal line loss data are identified based on the clustering results. This solution uses cluster analysis methods combined with characteristic data such as line current, voltage, power factor, and user load in the power system to accurately identify and analyze line loss characteristics in the distribution network. It is particularly suitable for handling abnormal line losses caused by abnormal line topology and overload in the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A schematic diagram of a method flow chart provided by an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of a specific logic flow provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention. It can be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "one", "said" and "the" used herein may also include plural forms. It should be further understood that the term "including" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used here may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items. It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as generally understood by those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.

[0021] The embodiments of the present invention relate to a distribution network line loss analysis and anomaly identification technology, and in particular, a distribution network line loss feature analysis method based on an improved snow melting algorithm to optimize the initial cluster center. Specifically, the present invention uses a cluster analysis method combined with characteristic data such as line current, voltage, power factor, and user load in the power system to accurately identify and analyze line loss characteristics in the distribution network, and is particularly suitable for processing abnormal line losses caused by abnormal line topology and overload in the distribution network.

[0022] Traditional methods for analyzing line losses in distribution networks have problems such as unstable clustering results and difficulty in identifying abnormal points. The present invention optimizes the initial clustering center by introducing a periodic oscillation mutation strategy and the principle of good point sets, and further adjusts the center point in the clustering process by improving the snow ablation algorithm, thereby improving the accuracy and convergence of clustering. In addition, the present invention also solves the problem of identifying small-proportion abnormal data by assigning higher weights to abnormal line loss points, and can effectively distinguish normal line losses from abnormal line losses. To this end, this paper proposes a method that combines an improved snow ablation algorithm with a K-means algorithm, aiming to optimize the clustering characteristic analysis and anomaly identification of line losses in distribution networks. By introducing a periodic oscillation mutation strategy and the principle of good point sets, the improved snow ablation algorithm can improve the convergence speed and global search capability, thereby optimizing the initial center selection of the K-means algorithm.

[0023] The general design idea is to collect multi-time period line loss data of distribution network lines, and use characteristic parameters such as current, voltage, power factor of the substation, and user load to build a line loss calculation model. Then, the zero mean unit variance standardization is used to balance the features to avoid some features (such as current) dominating the clustering process due to the large value range. Then, the initial clustering center selection is carried out by introducing the periodic oscillation mutation strategy and the principle of the best point set to optimize the rationality and clustering effect of the initial clustering results. In the process of cluster center optimization, the improved snow melting algorithm is combined for dynamic adjustment to improve the convergence and accuracy of the algorithm. For the identification of small-scale abnormal line loss data, the method increases the weight of the abnormal points by adaptively increasing their influence in cluster analysis. Finally, the dynamic changes and distribution laws of line loss data in the time and space dimensions are studied through spatiotemporal analysis, and the abnormality of line loss data is identified based on the clustering results.

[0024] The embodiment of the present application discloses a line loss clustering characteristic analysis and anomaly identification method for determining the cause of line line loss. The distribution network line loss clustering characteristic analysis and anomaly identification method obtains the line historical data and recent data for analysis to derive the cause of line loss, and further improves the accuracy of the analysis by collecting data on the line on site to ultimately determine the cause of line loss. It is suitable for distribution network energy efficiency management, fault diagnosis and power dispatch optimization, and can provide more accurate data support for power grid operation and load management, and has important practical application value. Further exploring the spatiotemporal laws of distribution network line loss data through spatiotemporal analysis can provide a more scientific decision-making basis for power grid load dispatching, energy saving and consumption reduction, and power grid optimization management.

[0025] For example, Figure 1 , 2 As shown, the line loss clustering characteristic analysis and anomaly identification method includes:

[0026] Step 1: Collect line loss data of distribution network lines, including line loss values ​​in multiple time periods;

[0027] Step 2: Use characteristic parameters such as line current, voltage, area power factor, and user load to build a line loss calculation model;

[0028] Step 3: Each feature is normalized by zero mean unit variance transformation to keep the modeling balanced and avoid certain features (such as current) dominating the clustering due to their large numerical range;

[0029] Step 4: Select the initial center of line loss clustering, introduce the periodic oscillation mutation strategy and the principle of optimal point set, and improve the rationality of the initial clustering center and the overall clustering effect;

[0030] Step 5: Optimize the line loss clustering center and use the improved snow melting algorithm to dynamically adjust the initial center to further enhance the convergence and accuracy of the algorithm;

[0031] Step 6: Optimize the cluster center according to the line loss, filter the line loss cluster abnormal points and adaptively increase the weight to enhance the influence of the abnormal points and solve the identification of small-proportion abnormal line loss data;

[0032] Step 7: Use the improved K-means algorithm combined with snow melting to optimize the selection of cluster centers to perform accurate cluster analysis on the line loss data and effectively separate the line loss features of different categories;

[0033] Step 8: Focus on the dynamic changes and distribution patterns of power grid line loss data in time and space dimensions through spatiotemporal analysis;

[0034] Step 9: Perform anomaly identification based on the line loss data clustering results.

[0035] The line loss model construction in step 2 includes:

[0036] By considering various characteristics, an extended line loss calculation model including multiple physical quantities can be constructed:

[0037]

[0038] Where: P loss represents line loss power (watt, W); I represents current (ampere, A); R(T) represents resistance at temperature T (ohm, Ω); PF represents power factor; f(V) represents the influence function of voltage V on line loss; g(L) represents the influence function of load L on line loss.

[0039] The relationship between resistance and temperature can be expressed as:

[0040] R(T)=R0(1+αT)

[0041] Where: R(T) represents the resistance at temperature T (ohm, Ω); R0 represents the resistance at reference temperature (ohm, Ω); α represents the temperature coefficient of resistance (1 / ℃); T represents temperature (Celsius, ℃).

[0042] The effect of voltage V on line loss is usually reflected by voltage drop. Generally speaking, the greater the voltage drop, the greater the line loss. A simple voltage influence function can be expressed by the following formula:

[0043]

[0044] Where: f(V) represents the influence function of voltage on line loss; V represents voltage (volts, V); V0 represents the reference voltage (volts, V).

[0045] When the voltage is lower than the reference voltage, the impact of the voltage on the line loss will be aggravated.

[0046] The load L usually directly affects the current through the line. An increase in load will lead to an increase in current, thereby increasing line losses. The load impact function can be expressed by the following relationship:

[0047]

[0048] Where: g(L) represents the influence function of load on line loss; L represents load (kilowatt, kW); L0 represents reference load (kilowatt, kW).

[0049] The characteristic parameter standardization process in step 3 includes:

[0050] The zero mean unit variance transformation is introduced to convert the data of all features to the same scale (i.e., mean zero and standard deviation one), thereby avoiding the numerical range of a specific feature being too large (such as current, voltage, etc.) to dominate the analysis or modeling results.

[0051] Make the obtained data comparable. Whether it is current, voltage or other physical quantities with different dimensions, they can be compared under the same standard to prevent certain features from dominating the analysis due to their large dimensions. This method converts the data into a distribution with a mean of 0 and a standard deviation of 1 by subtracting the mean of each feature value from the feature and then dividing it by the standard deviation of the feature. The standardization formula is as follows:

[0052]

[0053] Where: x i ′ represents the standardized eigenvalue of the i-th feature. i represents the original i-th eigenvalue;

[0054] μ i represents the mean of the i-th feature, and the formula is:

[0055]

[0056] Where: N is the number of samples in the data set.

[0057] σ i Represents the standard deviation of the i-th feature, and the formula is:

[0058]

[0059] Save the collected data for each feature in a matrix:

[0060]

[0061] Where: n is the number of samples, p is the number of features (5 features: I, V, PF, T, L).

[0062] The standardization process can be divided into the following steps:

[0063] First calculate the mean and standard deviation: The formula for calculating the mean is as follows:

[0064]

[0065] The formula for calculating the standard deviation is as follows:

[0066]

[0067] Secondly, each feature is standardized by the obtained standard deviation, that is, each feature of each sample is transformed using the Z-score formula:

[0068]

[0069] For the matrix X, the final standardized data set X′ is:

[0070]

[0071] The standardized data will have the following characteristics:

[0072] The mean of each feature is 0, that is,

[0073] The standard deviation of each feature is 1, that is,

[0074] The initial center selection of line loss clustering in step 4 includes:

[0075] By introducing a good point set, the initial “snow point” is selected through evenly distributed search individuals to increase the diversity of the population and avoid concentration in a certain area. sis the unit cube in s-dimensional Euclidean space, r∈M s , then the good point set can be expressed as:

[0076]

[0077] Where: {r s (n) ·k} represents the decimal part, n is the number of points; P n (k) represents the good point set.

[0078] If there is a deviation conform to Where C(r,ε)n -1+ε is a constant that is only related to r and ε, then we can call P n (k) is the set of good points, and r is called a good point.

[0079] Take r = {2cos(2πk / p), 1≤k≤s}, where p is the smallest prime number satisfying (pt)≥s, then r is a good point.

[0080] Where: It represents the deviation function related to the sample size n and describes the characteristics of the sample distribution.

[0081] For individuals that have been uniformly initialized by the good point set principle, the surrounding density of each individual is evaluated. At the same time, a periodic oscillation mutation strategy is introduced to enhance the global search ability of the algorithm. Every certain number of iterations, the current "snow point" is oscillated and mutated so that it can jump out of the local optimum and explore new possible areas:

[0082]

[0083] Where: x i is the current solution, x rand are other randomly selected solutions, β and ω are the amplitude and frequency parameters of the oscillation, respectively, and t is the current iteration number.

[0084] The steps to improve the snow ablation algorithm in step 5 include:

[0085] Step 1: For the line loss data set X = {x1, x2, …, x n}, calculate the line loss sample data x i and x j (i≠j) The Euclidean distance d(x i ,x j ):

[0086]

[0087] Where: x iRepresents the feature vector of the i-th sample point in the data set; x j represents the feature vector of the jth sample point in the data set; d(x i ,x j ) represents the sample x i and x j The Euclidean distance between it Represents the sample point x i The value of the feature at iteration t; x jt Represents the sample point x j The value of the feature at the tth iteration; d represents the feature dimension of the sample; It means to sum the squares of the differences of all features and calculate the sample point x i and x j Differences in all characteristics.

[0088] Compute the average distance of all samples in the dataset:

[0089]

[0090] Where: Represents the number of combinations between every two n data; D represents the average distance between samples.

[0091] Step 2: Find the number of line loss samples within a range of D from each sample around the selected line loss data sample, which is called the density value and is expressed as:

[0092]

[0093] Where: Used to determine the sample point x j Is it at the sample point x? i Within the range of D; density(x i ) represents the sample point x i The density value reflects the number of sample points within the distance D.

[0094] Step 3: Take the line loss data point x corresponding to the maximum density value i is the initial cluster center.

[0095]

[0096] Where: μ j Indicates the point x corresponding to the maximum density value selected i as the initial center.

[0097] Step 4: Determine the initial center of the line loss cluster obtained, simulate the periodic oscillation process through spring vibration, obtain a new candidate initial center, minimize the potential energy within the cluster, calculate the distance from all sample points in the cluster to the center of the cluster, and compare them. If the distance is less than the previously obtained value, it is identified as a better initial center.

[0098] Step 6 improves the ability to identify small-scale abnormal line loss data, including:

[0099] The line loss data is preprocessed to obtain a structured data set, and the outlier degree of all line loss data is calculated using the isolated forest method and local anomaly factors:

[0100] The anomaly score for Isolation Forest is defined as follows:

[0101]

[0102] Where: h(x) represents the isolation depth of data point x (i.e., the number of tree layers); E(h(x)) represents the average isolation depth of data point x (averaged over all trees); n represents the size of the dataset; c(n) represents the regularization coefficient, which is used to standardize the isolation depth. The formula is: Where: H(i) is the harmonic number of the i-th term:

[0103] When Anomaly Score(x) is close to 1, it means that the data point x is more likely to be isolated and may be an outlier.

[0104] When Anomaly Score(x) is close to 0, it means that the data point x is difficult to be isolated and is a normal point.

[0105] When calculating the local anomaly factor, the reachable distance is calculated first:

[0106] The reachable distance from a data point p to its neighbor point o is defined as:

[0107] Reachability Distance(p,o)=max{K-Distance(o),Distance(p,o)}

[0108] Where: K-Distance(o) represents the kth nearest neighbor distance of point o; Distance(p,o) represents the Euclidean distance or other metric distance between points p and o.

[0109] The local reachability density of a point p is defined as the inverse of the average reachability distance of its neighbors:

[0110]

[0111] Where: Nk (p) represents the k-nearest neighbor set of point p; |N k (p)| represents the number of neighbors of point p.

[0112] The local outlier factor of point p is defined as:

[0113]

[0114] Where: LRD(o) represents the local reachability density of point o; LRD(p) represents the local reachability density of point p.

[0115] When LOF(p)≈1: it means that the density of point ppp is close to the density of its neighbors and it is a normal point.

[0116] When LOF(p)>: it means that the density of point ppp is significantly lower than the density of its neighbors and is an outlier.

[0117] When LOF(p)>>1: it means that the density of point ppp is much lower than the density of its neighbors, which may be an extreme outlier.

[0118] Assign weight ω to data points according to their outlier degree i , the weight of abnormal points is higher than that of normal points. In the traditional K-means algorithm, the objective function is to minimize the intra-class squared error (SSE):

[0119]

[0120] After introducing weights, the objective function becomes:

[0121]

[0122] Where: i is the weight of the ith data point, c k is the center of the cluster.

[0123] The weighted cluster center update method assigns a weight ω to each data point in the cluster. i When calculating the cluster center, data points with larger weights contribute more to the cluster center, while outliers with smaller weights have less impact on the cluster center. This method effectively reduces the impact of outliers and improves the stability and accuracy of clustering. The formula is as follows:

[0124]

[0125] Where: c k It is cluster C k The weighted center of i It is cluster C k The i-th data point in ω; iis the weight of the i-th data point, which is usually set according to the outlier degree of the data point or other factors; is the sum of the weights of all data points in the cluster.

[0126] The snow melting optimization in step 7 improves the K-means line loss accurate clustering analysis, including:

[0127] Based on the candidate cluster centers generated by the snow melting algorithm, the distance from each sample point to each initial center is calculated, and the most suitable cluster center is selected. Collect the objective function set X = {x1, x2, ..., x n}, the data set contains the line loss values ​​of the distribution network lines in different time periods. Select an initial range of cluster number K, for example K = 1, 2, ..., K max . Where K max is the maximum number of clusters allowed.

[0128] When selecting the best K value, relying solely on the DBI index may not fully reflect the effect of clustering. Therefore, in order to more comprehensively evaluate the clustering results, the SSE (Sum of Squares for Error) indicator is introduced as an auxiliary tool. SSE is used to assist in measuring the intra-class error, that is, the sum of the squares of the distance from each point to its cluster center. The smaller the SSE, the more tightly concentrated the points in the cluster are, and the higher the consistency within the class.

[0129] Since the SSE index gradually decreases with the increase of K value, by observing the change of the slope of the SSE index and finding the place with large changes, we can help determine the best K value. Combining the two indicators of DBI and SSE, we finally choose a K value that can effectively balance the intra-class compactness and inter-class separation to ensure the rationality and effectiveness of the clustering results:

[0130]

[0131] Where: C ki For data point x i The cluster center.

[0132] The SSE values ​​corresponding to different K values ​​are plotted into a curve, and the position where the curve presents an "elbow" shape is selected, that is, the cluster number K where the SSE decreases significantly slower is selected. optimal .

[0133] The spatiotemporal analysis of line loss data in step 8 specifically includes:

[0134] By comparing the dynamic changes and distribution patterns of power grid line loss data in time and space, and identifying the periodicity, volatility and regional differences of line loss, abnormal points and potential causes can be located;

[0135] The formula for constructing and capturing the mean and fluctuation characteristics of line loss time series data is as follows:

[0136]

[0137] Where: μ t represents the mean line loss at time point t, indicating the overall operation level; x i,t The line loss value of the i-th line at time point t.

[0138]

[0139] Where: t It represents the standard deviation of line loss at time point t, and measures the intensity of fluctuation.

[0140] Through μ t Analysis of and t Can identify:

[0141] σ t Large (high volatility), which may indicate uneven load or equipment failure;

[0142] σ t Small (low volatility), indicating that the system is running stably.

[0143] The spatial distribution characteristics of line loss are analyzed by weighted average, the differences in line loss in different areas of the power grid are found, and high line loss areas are identified. The line loss spatial dimension analysis formula is constructed:

[0144]

[0145] Where: L i represents the weighted line loss mean of area i; w ij represents the spatial weight between regions i and j, which is determined by the geographical distance or the topological relationship of the power grid; x j represents the line loss value of area j; m represents the number of areas involved in the calculation.

[0146] By comparing the L calculated by substituting the data i The following conclusions can be drawn:

[0147] Higher L i This indicates that the line equipment is aging or the load is uneven; a lower L i This indicates that the line operation efficiency is high.

[0148] Considering both time and space dimensions, a space-time matrix is ​​constructed to unify the time and space characteristics and reveal the dynamic spatial distribution characteristics of line loss data. The matrix contains line loss data of M regions at T time points. The matrix is ​​as follows:

[0149]

[0150] Where: X ij Represents the line loss value of area i at time point j.

[0151] Each row represents the line loss data of a certain area at different points in time. By analyzing the changing trend of each row in the matrix (for example, volatility, periodicity, etc.), we can capture the line loss change pattern of the area in different time periods. Each column in the matrix represents the line loss data of different areas at the same time. By analyzing the columns in the matrix, we can reveal the difference in line loss in different areas at the same time. For example, some areas may have higher line losses due to aging equipment or uneven loads. This spatial distribution feature can be identified by analyzing the matrix columns. Through comprehensive analysis of the entire space-time matrix, especially using methods such as clustering or principal component analysis, the coupling relationship in time and space can be mined.

[0152] During certain specific time periods, certain areas may experience high line losses simultaneously (possibly due to load concentration, equipment failure, etc.). Such spatiotemporal anomaly patterns can be identified through comprehensive analysis of the spatiotemporal matrix.

[0153] The abnormal line loss data identification in step nine includes:

[0154] The clustering results of line loss data are used to identify and isolate areas with abnormally high line losses in the distribution network to provide support for optimization decisions.

[0155] Assume that the line loss data of the distribution network contains d time data of line loss of n lines. The line loss data is analyzed in the time dimension, and the data is normalized and used as the line loss data set X input:

[0156]

[0157] Where: x ij represents the line loss value of the i-th branch line loss in the j-th time period; the row vector x i =[x i1 x i1 …x id ] represents all line loss data of the i-th branch line loss.

[0158] Generate a set of "first generation snow points", which are equivalent to the first generation individuals in traditional evolutionary algorithms and are the starting points of the algorithm search process. In order to ensure that a wide range of search spaces are covered and improve the global search capabilities of the algorithm, these initial snow points can be generated by random distribution or by using more advanced methods such as the best point set principle, so as to evenly distribute these initial points throughout the search space.

[0159]

[0160] Where: row vector cj ∈X(j=1,2,…,Q).

[0161] The data set and the selected initial cluster center points are input into the K-means algorithm for cluster analysis. After clustering is completed, DBI is used to evaluate the clustering results. This indicator comprehensively reflects the intra-class compactness and inter-class separation.

[0162]

[0163] Where: d(c i ,c j ) means c i 、c j The distance between two different cluster centers; S i , S j c i 、c j The average intra-class distance between two clusters is calculated as follows:

[0164]

[0165] Where: N(A j ) represents cluster A j The number of line losses included.

[0166] After clustering is completed, we can first calculate the deviation between each sample point and the center of the cluster. If the distance between a sample point and the cluster center is significantly greater than that of other sample points, the sample point may be an outlier.

[0167] Calculate the deviation of the line loss sample data, the distance from each sample point to its cluster center:

[0168]

[0169] For each cluster, calculate the distance from all sample points to the cluster center and find the point with the largest distance. Assume that the maximum distance is d max , a threshold α can be set, if the distance d(x i ,c k ) is greater than d max +α, the sample point is considered to be an outlier.

[0170] In order to further improve the accuracy of outlier identification, the local outlier factor (LOF) method can be used. This method determines whether it is an outlier by comparing the density of a sample point with its neighborhood.

[0171] Calculating the local reachability density, which is calculated by the relationship between the sample point and other points in the neighborhood, is the core part of the LOF method. The formula is as follows:

[0172]

[0173] Where: reachability(x i ,x j ) is the sample point x i With x j The reachable distance, N k (x i ) is the sample point x i k-nearest neighbors.

[0174] The LOF value is a measure of the sample point x i Whether it is an important indicator of an abnormal point, the calculation formula is as follows:

[0175]

[0176] If the LOF value of a line loss sample point is high, it means that it has a lower density than the points in the neighborhood and may be an outlier. Usually, if the LOF value is greater than the set threshold, the point is considered an outlier.

[0177] The identified anomalies are displayed through visualization means, using scatter plots, heat maps and other visualization methods to display the clustering results of each sample point, the location of the anomaly and its indicators; a report containing anomaly information, identification basis, possible causes, etc. is output and provided to relevant personnel for subsequent processing.

[0178] The above are all preferred embodiments of the present application, and are not intended to limit the protection scope of the present application. Any feature disclosed in this specification (including the abstract and drawings), unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes. That is, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0179] Through the technical scheme of the above-mentioned embodiment, the line loss clustering analysis and abnormality identification method can identify the potential causes of abnormal line loss according to the operation data of the distribution network line and the user's power consumption data, and further locate the possible abnormal lines. This method provides the staff with accurate line loss characteristic analysis results, which helps to quickly lock the location of abnormal line loss, improve the troubleshooting efficiency and processing accuracy. In general, the main advantages of this embodiment are: 1. Through the classification analysis of historical line loss data and typical cases, the data to be analyzed can be matched with historical cases, and the line loss causes with higher probability can be quickly obtained, providing the staff with a reliable decision-making basis for subsequent processing; 2. Combined with the measurement data of the on-site distribution equipment, the potential line loss causes that may exist in the distribution equipment can be deeply analyzed, and at the same time, multi-dimensional cross-validation can be performed in combination with the user's power consumption data to improve the accuracy of the line loss cause analysis; 3. Using the measurement data and the user's power consumption information obtained from the distribution network, it is possible to determine the lines where there may be abnormal line losses, providing strong support for the rapid positioning and processing of abnormal lines.

[0180] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The above is only a specific implementation method of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or replacements that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A distribution network line loss clustering characteristic analysis and anomaly identification method, characterized in that: include: Collecting line loss data of the distribution network and building a line loss model, wherein the line loss data includes line loss values ​​in a corresponding time period; Before clustering analysis of line loss data, the cluster center of line loss is optimized and adjusted, wherein the abnormal points of line loss clustering are screened and the weights are adaptively adjusted through the optimized and adjusted cluster center; Perform cluster analysis on line loss data, cluster different types of line loss features and separate them; The line loss characteristics are used to perform spatiotemporal analysis to identify abnormal line loss data.

2. The method according to claim 1, characterized in that The line loss model includes: P loss represents line loss power, I represents current, R(T) represents resistance at temperature T, PF represents power factor, f(V) represents the influence function of voltage V on line loss, and g(L) represents the influence function of load L on line loss; Among them, R(T)=R0(1+αT), R0 represents the resistance at the reference temperature, α represents the temperature coefficient of resistance, V0 represents the reference voltage, and L0 represents the reference load.

3. The method according to claim 1, characterized in that Also includes: Before conducting accurate cluster analysis on line loss data, the characteristic parameters are standardized; The characteristic parameter standardization process includes: x i ′ represents the standardized eigenvalue, x i represents the original i-th eigenvalue, i is the label of the line loss data collected from the distribution network line, μ i represents the mean of the i-th feature; N is the sample data set of the original parameters, N is x i The number of samples in the set data set; σ i represents the standard deviation of the ith feature, The characteristic parameters after standardization are recorded in the matrix X. n is the number of samples, p is the number of feature types, and the feature types include: current I, voltage V, power factor PF, temperature T, and load L.

4. The method according to claim 1, characterized in that: The optimizing and adjusting the clustering center of the line loss includes: Set up a good point collection in, Indicates taking the decimal part, n is the number of points, k is a positive integer, s is the dimension of the Euclidean space, r is the best point, r∈M s , M s is a unit cube in s-dimensional Euclidean space; r = {2cos(2πk / p), 1≤k≤s}, p is the smallest prime number satisfying (pt)≥s; represents…,,x i represents the original i-th eigenvalue of the current solution, x rand represents other randomly selected solutions, β and ω are the amplitude and frequency parameters of the oscillation, respectively, and t is the current iteration number.

5. The method according to claim 4, characterized in that The optimized and adjusted cluster centers include: For the line loss data set selected by the best point set, the Euclidean distance between two line loss data points in the line loss data set is calculated, and then the average value of the Euclidean distances of all line loss data points in the data set is obtained and used as the average distance; Select one of the line loss data points, and query the number of line loss samples covered within the average distance range around this line loss data as the density value; The line loss data point with the largest density value is taken as the initial cluster center.

6. The method according to claim 5, characterized in that The screening of line loss clustering abnormal points and adaptively adjusting weights include: The outlier degree of all line loss data is calculated through the isolation forest model and the local anomaly factor algorithm, and the weight of the outlier points in the line loss clustering is reduced according to the outlier degree result, thereby reducing the impact of abnormal data on the clustering results.

7. The method according to claim 1, characterized in that The process of clustering and separating different types of line loss features includes: The SSE values ​​corresponding to different K values ​​are plotted into a curve, and the number of clusters with a significantly slower SSE decrease rate is selected as the optimal K value according to the plotted curve, K = 1, 2, ..., K max , K max is the maximum number of clusters allowed.

8. The method according to claim 1, characterized in that: Before using the line loss characteristics to perform spatiotemporal analysis, the following steps are included: Build a model to capture the mean and fluctuation characteristics of line loss time series data: μ t represents the mean line loss at time point t, x i,t represents the line loss value of the ith line at time point t; where the parameter used to measure the fluctuation intensity is Construct line loss space dimension analysis model: L i represents the weighted line loss mean of area i; w ij represents the spatial weight between regions i and j, x j represents the line loss value of area j; m represents the number of areas involved in the calculation.

9. The method according to claim 8, characterized in that The identifying and obtaining abnormal line loss data includes: LOF(x i ) The LOF value output is used to determine the sample point x i Is it an indicator parameter of an abnormal point? If the LOF value is greater than the set threshold, it is judged as abnormal line loss data. i ,x j ) is the sample point x i With x j The reachable distance, N k (x i ) is the sample point x i k-nearest neighbors, lrd() represents the calculation function of local reachable density.

Citation Information

Cited By

  • Data processing method of electric power laying device

    CN120387044A

  • Line loss positioning method based on phase feature clustering

    CN121231923A

  • Self-adaptive temperature control method and system for graphene electric heater

    CN121677032A