Distribution transformer load clustering method and system based on improved topological data analysis and medium
Through polar coordinate mapping and continuous co-modulation method combined with variable pixel parameters and weighting functions, the problem of high topological noise and low resolution in distribution load clustering analysis in distribution networks is solved, more accurate clustering analysis is achieved, and refined support for distribution network management and scheduling is improved.
Patent Information
- Application Number
- CN202510979411.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively capture the internal structural characteristics of the data in distribution load cluster analysis in distribution networks, and there is a lot of topological noise and low image resolution, resulting in inaccurate clustering results.
The topological structure information of the point cloud data set is extracted using polar coordinate mapping and continuous co-modulation method, and noise is filtered through variable pixel parameters and weighting functions, and cluster analysis is performed with the K-mean clustering algorithm.
It improves the accuracy and reliability of matching load clustering, breaks through the limitations of traditional Euclidean space, accurately extracts topological features, and optimizes clustering performance.
Smart Images

Figure CN120492961A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to distribution network technology, and in particular to a distribution transformer load clustering method, system and medium based on improved topology data analysis. Background Art
[0002] In the operation and management of distribution networks, cluster analysis of distribution transformer loads is of great significance for optimizing scheduling and improving power supply reliability. Traditional methods often use clustering methods after dimensionality reduction. These methods have difficulty effectively capturing the intrinsic structural characteristics of the data, resulting in inaccurate clustering results. In addition, existing topological data analysis methods suffer from problems such as excessive topological noise and low image resolution when processing continuous images, which limits their application in distribution transformer load clustering. Therefore, there is an urgent need for an improved topological data analysis method that can accurately extract the topological characteristics of distribution transformer loads, improve the accuracy and reliability of clustering analysis, and provide strong support for the refined management and scheduling of distribution networks. Summary of the Invention
[0003] Purpose of the invention: The purpose of the present invention is to provide a distribution transformer load clustering method based on improved topological data analysis that reduces topological noise and effectively captures the intrinsic structural characteristics of data; another purpose of the present invention is to provide a distribution transformer load clustering system and medium based on improved topological data analysis.
[0004] Technical solution: The distribution transformer load clustering method based on improved topological data analysis described in the present invention includes the following steps: The active power time series data of different users in the same time period and station area are mapped to polar coordinates in chronological order to obtain point cloud data sets of different users. The topological structure information of the point cloud data sets is extracted using the persistent coherence method to obtain persistent life and death graphs of different users, where the horizontal axis of the persistent life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears. The persistent life and death image is preprocessed, a variable pixel parameter p is introduced, the preprocessed persistent life and death image is divided into pixel blocks, a weighting function is introduced to convert the persistent life and death image into a scalar function, topological noise is filtered, the scalar function within the pixel block is integrated, and the integral value is used as the value of the corresponding pixel block. All pixel blocks are sequentially combined to obtain a column matrix pixel map, i.e., a persistent image; All pixel blocks are combined in sequence and flattened to obtain n groups of one-dimensional vectors of specific lengths. Cluster analysis is performed to obtain the clustering results of the distribution transformer load.
[0005] Furthermore, before performing polar coordinate mapping in chronological order, it is necessary to perform median filtering and normalization processing on the collected active power time series data of different users in the same time period and the same station area.
[0006] Furthermore, the active power time series data of different users in the same time period and the same substation area, which are mapped to polar coordinates in chronological order, are time series data with a sampling frequency of 15 minutes, and contain a total of 96 points; Polar coordinate mapping takes the origin of the two-dimensional coordinate as the center of the circle, divides the circle evenly into 96 parts, and performs polar coordinate mapping in chronological order. The polar coordinate corresponding to the nth active power is ,in It represents the normalized value of the nth active power data.
[0007] Furthermore, we use the continuous homology method to extract the topological structure information of the point cloud data set and obtain the continuous life and death graphs of different users, as follows: Based on point cloud data sets from different users, a VR complex is constructed. The Euclidean distance between all points is calculated, and the maximum Euclidean distance is used as the upper threshold of the VR complex parameter ε. Continuous homology analysis is performed on the point cloud data within the ε threshold to extract the topological structure information of the point cloud data set and obtain a continuous life and death graph for different users.
[0008] Furthermore, the continuous life and death diagram is preprocessed as follows: the part above the diagonal of the continuous life and death diagram is retained and rotated 45 degrees clockwise. The length c of the rotated image is the difference between the maximum birth time and the minimum birth time; the width d of the image is the difference between the maximum life cycle and the minimum life cycle.
[0009] Furthermore, the weighting function for ; Where b is the upper bound parameter, a is the lower bound parameter, and t is the vertical coordinate of the rotated image. Using a piecewise weighted function of a linear mapping and changing its lower bound parameter can effectively filter out topological noise.
[0010] Furthermore, cluster analysis is performed on n groups of one-dimensional vectors of specific lengths to obtain the clustering results of distribution transformer loads, as follows: (1) Use the elbow rule to determine the optimal number of clusters k for K-means clustering; (2) Take n groups of one-dimensional vectors of a specific length as sample data, and randomly select k initial data from the sample data as the mean vector, which is the cluster center; (3) Calculate the distance between each sample and each mean vector, and divide the clusters according to the distance; (4) Recalculate the mean vector of each cluster. If any mean vector is updated, replace the mean vector of the previous round and obtain a new mean vector. Return to step (3). If the mean vectors are no longer updated, output the cluster division result, that is, the clustering result of the distribution transformer load.
[0011] The distribution transformer load clustering system based on improved topological data analysis of the present invention includes: The time series data feature mining module is used to perform polar coordinate mapping on the active power time series data of different users in the same time period and station area in chronological order to obtain point cloud data sets of different users. The topological structure information of the point cloud data sets is extracted using the continuous coherence method to obtain continuous life and death graphs of different users, where the horizontal axis of the continuous life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears. The topological data analysis module is used to preprocess the persistent life and death image, introduce a variable pixel parameter p, divide the preprocessed persistent life and death image into pixel blocks, introduce a weighting function to convert the persistent life and death image into a scalar function, filter topological noise, integrate the scalar function within the pixel block, and use the integral value as the value of the corresponding pixel block. All pixel blocks are sequentially combined to obtain a column matrix pixel map, i.e., a persistent image. The distribution transformer load clustering analysis module is used to flatten all pixel blocks after sequential combination to obtain n groups of one-dimensional vectors of specific lengths, perform clustering analysis, and obtain the clustering results of the distribution transformer load.
[0012] The computer device of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0013] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program implements the steps of the above method when executed by a processor.
[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: 1. The present invention obtains point cloud data sets of different users through polar coordinate mapping, and uses a continuous homology method to extract the topological structure information of the point cloud data set to obtain continuous life and death graphs of different users. The topological data analysis method is used to extract features of the distribution transformer load, retain the part above the diagonal of the continuous life and death graph, rotate 45 degrees clockwise, and divide the image based on the variable pixel parameter p. The continuous life and death graph is converted into a scalar function on the plane through a weighted function, and the points in each area are integrated as the size of the regional pixel value to obtain a continuous image. This method breaks through the limitations of traditional Euclidean space and focuses on the structural characteristics of the data itself, making feature extraction more accurate and precise; 2. The present invention improves the weighting function of the continuous image, using a linearized piecewise function to filter out the original topological noise, while ensuring that topological features with longer life cycles have greater weights; 3. The present invention improves the pixel size of the continuous image, breaking through the traditional fixed pixel size of the continuous image at 1. By controlling the pixel size, the length of the continuous image vectorization is controlled. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a multi-peak load polar coordinate mapping diagram; Figure 3 It is a polar coordinate mapping diagram of single peak load; Figure 4 For the continuous life and death diagram; Figure 5 is the continuous image when the parameter a is 0; Figure 6 This is the persistent image when the parameter a is 0.01; Figure 7 This is the persistent image when the parameter a is 0.03; Figure 8 This is the persistent image when the parameter a is 0.05; Figure 9 Schematic diagram of clustering results. DETAILED DESCRIPTION
[0016] The present invention will be further described below with reference to the accompanying drawings.
[0017] The distribution transformer load clustering method based on improved topological data analysis of the present invention comprises the following steps: Set the sampling frequency to 15 minutes to obtain 96 points of active power time series data for different users in the same area and time period. Perform median filtering on this data and normalize the filtered data. Polar coordinate mapping is performed in chronological order to obtain a point cloud data set for each user.
[0018] Polar coordinate mapping takes the origin of the two-dimensional coordinate as the center, draws a circle with a radius of 1, divides the circle evenly into 96 parts, and performs polar coordinate mapping in chronological order. The polar coordinate corresponding to the nth active power is ,in It represents the normalized value of the nth active power data.
[0019] The persistent homology method is used to extract the topological structure information of the point cloud data set, and a persistent life and death graph of different users is obtained, where the horizontal axis of the persistent life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears.
[0020] The Vietoris-Rips Complex (VR complex) is used to construct complexes of different "thicknesses" based on given point cloud data by adjusting threshold parameters to capture the topological evolution process of data from discrete points to fully connected points.
[0021] Based on the point cloud data sets of different users, a VR complex is constructed; the Euclidean distance between all points is calculated, and the maximum Euclidean distance is used as the upper threshold of the VR complex parameter ε. The value range of the VR complex parameter ε is determined. Within the threshold of ε, for each ε value, a corresponding VR complex is constructed to form a complex sequence. For each VR complex, the topological features of each dimension are calculated, such as the 0-dimensional topological feature is the number of connected branches, and the one-dimensional topological feature is the number of rings. As ε increases, the generation and extinction process of the topological features of each dimension is observed. The data features under each simple complex are analyzed to determine the topological structure information, and the continuous life and death graph of different users is obtained. In the continuous life and death graph, there is a diagonal line. , the farther away from the diagonal line, the more obvious the topological features are.
[0022] Preprocess the continuous life and death diagram: retain the part above the diagonal of the continuous life and death diagram, rotate it 45 degrees clockwise, and the length c of the rotated image is the difference between the maximum birth time and the minimum birth time; the width d of the image is the difference between the maximum life cycle and the minimum life cycle.
[0023] A variable pixel parameter p is introduced to divide the preprocessed persistent life and death image into pixel blocks. A weighted function is introduced to transform the persistent life and death image into a scalar function. The topological noise is filtered out, and the scalar function within the pixel block is integrated. The integral value is used as the value of the corresponding pixel block. All pixel blocks are combined in sequence to obtain a column matrix pixel map, i.e., a persistent image.
[0024] The variable pixel p controls the vector length of the output persistent image by controlling the size of this parameter. The pixel parameter represents the size of a single pixel when generating a persistent image. The size of a single pixel block is p. 2 .
[0025] Weighting function for ; The weighting function introduces two variables: the upper bound parameter b and the lower bound parameter a. A piecewise function is used to perform a weighted mapping on the rotated points. Here, t is the vertical coordinate of the rotated image. Using a piecewise weighting function with a linear mapping and varying its lower bound parameter can effectively filter out topological noise.
[0026] All pixel blocks are combined in sequence and flattened to obtain n groups of one-dimensional vectors of specific lengths. Cluster analysis is performed to obtain the clustering results of the distribution transformer load.
[0027] Using the K-means clustering algorithm, we perform cluster analysis on n groups of one-dimensional vectors of a specific length. The clustering results of the distribution transformer load are as follows: (1) Use the elbow method to determine the optimal number of clusters k for K-means clustering; (2) Randomly select k initial data from the sample data as the mean vector, which is the cluster center; (3) Calculate the distance between each sample and each mean vector, and divide the clusters according to the distance; (4) Recalculate the mean vector of each cluster. If any mean vector is updated, replace the mean vector of the previous round to obtain a new mean vector. Return to step (3). If the mean vectors are no longer updated, output the cluster division results.
[0028] Consider the time series data of active power from 4,862 users in a certain district of a city on a certain day. This data is median filtered and then normalized using the max-min filter. The processed data is numbered by timestamp, and the angle corresponding to each data point is determined by the number. The normalized value is the corresponding length.
[0029] like Figure 2 As shown in the figure, mapping the distribution transformer load data onto the circle not only shows obvious topological characteristics, but also maintains the original physical characteristics of the distribution transformer load. It can be seen from the circle that the active power of the distribution transformer load is almost 0 from 10 pm to 8 am. During the day, the active power data shows a bimodal pattern, with obvious peaks in electricity consumption in the morning and afternoon, and a trough in electricity consumption at noon, which is consistent with the original physical characteristics of the data.
[0030] On the data point cloud after polar coordinate mapping, a VR complex is constructed, and the Euclidean distance between all points is calculated. The maximum Euclidean distance is used as the upper threshold of the parameter ε, and a set of increasing parameters is constructed from 0 to the upper threshold of the parameter. ,in As the upper threshold of parameter ε, the parameter ε of VR complex is continuously increased, and the life cycle of topological features is calculated.
[0031] The topological characteristics of the calculated distribution transformer load data are visualized in the form of a continuous life-death diagram.
[0032] A persistence birth-and-death diagram is used to illustrate the lifecycle of topological features in data, such as connected components and holes. Each point in the diagram represents a topological feature, with the horizontal axis representing the feature's birth time (Brith) and the vertical axis representing the feature's death time (Death). The distribution and persistence of the points (i.e., their distance from the diagonal) can reveal the stability and importance of the data's topological structure. For example, points far from the diagonal represent features that persist for a long time in the data and may represent important patterns in the load data.
[0033] Based on the continuous life and death diagram, the continuous image method is used to vectorize the image. First, the continuous life and death diagram is divided along the diagonal line, the upper part is retained and rotated 45 degrees clockwise, and the rotated image is divided. L The maximum birth time With the youngest birth time The difference between the width and height of the image W For the maximum life cycle The difference from the minimum life cycle .
[0034] ; ; The number of pixels in the image is p, and the size of a single pixel block is p 2 The improved weighting function is used to transform the continuous life and death graph into a scalar function on a plane. The corresponding scalar function is integrated in each pixel area, and the integral value is used as the value of the pixel area. The improved weighting function is: ; like Figure 5 As shown in the figure, after using the weighting function, the values of a are 0, 0.01, 0.03, and 0.05, respectively. As the value of a increases, H1 with a shorter life cycle, that is, the topological features originally located at the bottom, are filtered out. The filtered part corresponds to the topological noise in the sample, while H1 with a longer life cycle is preserved, which is more conducive to the extraction of the sample's topological features. This shows that the proposed improvement scheme can effectively filter out some topological noise.
[0035] However, as the value of the lower bound parameter a continues to increase, some topological features with shorter life cycles will be shielded, resulting in the incorrect filtering of some important topological feature information. In actual use, the appropriate lower bound parameter a should be selected according to the actual situation of the data set to retain most of the important topological features for clustering while filtering out the topological noise in the data set, making the clustering effect more obvious.
[0036] After combining and flattening these calculated pixel block parameters, we obtain n groups of one-dimensional vectors of a specific length. K-means clustering is used to obtain clustering results for the distribution transformer load. 900 curves were selected for cluster analysis, and the results are as follows: like Figure 6As shown, the red line represents the cluster center. This example explores clustering methods for time-series distribution transformer load data from the perspective of topological data analysis. The life cycle of topological feature contraction and disappearance can reveal users' electricity usage habits in different time periods. This enables refined clustering, providing data support for profiling different users and supporting scheduling decisions for the distribution network.
[0037] The distribution transformer load clustering system based on improved topological data analysis of the present invention includes: The time series data feature mining module is used to perform polar coordinate mapping on the active power time series data of different users in the same time period and station area in chronological order to obtain point cloud data sets of different users. The topological structure information of the point cloud data sets is extracted using the continuous coherence method to obtain continuous life and death graphs of different users, where the horizontal axis of the continuous life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears. The topological data analysis module is used to preprocess the persistent life and death image, introduce a variable pixel parameter p, divide the preprocessed persistent life and death image into pixel blocks, introduce a weighting function to convert the persistent life and death image into a scalar function, filter topological noise, integrate the scalar function within the pixel block, and use the integral value as the value of the corresponding pixel block. All pixel blocks are sequentially combined to obtain a column matrix pixel map, i.e., a persistent image. The distribution transformer load clustering analysis module is used to flatten all pixel blocks after sequential combination to obtain n groups of one-dimensional vectors of specific lengths, perform clustering analysis, and obtain the clustering results of the distribution transformer load.
[0038] The computer device of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0039] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program implements the steps of the above method when executed by a processor.
Claims
1. A distribution transformer load clustering method based on improved topological data analysis, characterized in that: The following steps are involved: The active power time series data of different users in the same time period and station area are mapped to polar coordinates in chronological order to obtain point cloud data sets of different users. The topological structure information of the point cloud data sets is extracted using the persistent coherence method to obtain persistent life and death graphs of different users, where the horizontal axis of the persistent life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears. The persistent life and death image is preprocessed, a variable pixel parameter p is introduced, the preprocessed persistent life and death image is divided into pixel blocks, a weighting function is introduced to convert the persistent life and death image into a scalar function, topological noise is filtered, the scalar function within the pixel block is integrated, and the integral value is used as the value of the corresponding pixel block. All pixel blocks are sequentially combined to obtain a column matrix pixel map, i.e., a persistent image; All pixel blocks are combined in sequence and flattened to obtain n groups of one-dimensional vectors. Cluster analysis is performed to obtain the clustering results of distribution transformer loads.
2. The distribution transformer load clustering method based on improved topology data analysis according to claim 1 is characterized in that: Before performing polar coordinate mapping in chronological order, the collected active power time series data of different users in the same time period and the same station area need to be median filtered and normalized.
3. The distribution transformer load clustering method based on improved topology data analysis according to claim 2 is characterized in that: The active power time series data of different users in the same time period and the same substation are mapped in polar coordinates in chronological order. The sampling frequency is 15 minutes, and it contains 96 points in total. Polar coordinate mapping takes the origin of the two-dimensional coordinate as the center of the circle, divides the circle evenly into 96 parts, and performs polar coordinate mapping in chronological order. The polar coordinate corresponding to the nth active power is ,in It represents the normalized value of the nth active power data.
4. The distribution transformer load clustering method based on improved topology data analysis according to claim 1 is characterized in that: The continuous homology method is used to extract the topological structure information of the point cloud data set, and the continuous life and death graphs of different users are obtained as follows: Based on point cloud data sets from different users, a Vitalis-Lipps VR complex is constructed. The Euclidean distance between all points is calculated, and the maximum Euclidean distance is used as the upper threshold of the VR complex parameter ε. Continuous homology analysis is performed on the point cloud data within the ε threshold to extract the topological structure information of the point cloud data set and obtain continuous life and death graphs for different users.
5. The distribution transformer load clustering method based on improved topology data analysis according to claim 4 is characterized in that: The continuous life and death diagram is preprocessed as follows: the part above the diagonal of the continuous life and death diagram is retained and rotated 45 degrees clockwise. The length c of the rotated image is the difference between the maximum birth time and the minimum birth time; the width d of the image is the difference between the maximum life cycle and the minimum life cycle.
6. The distribution transformer load clustering method based on improved topology data analysis according to claim 5 is characterized in that: Weighting function for ; Among them, b is the upper bound parameter, a is the lower bound parameter, and t is the vertical coordinate of the rotated image.
7. The distribution transformer load clustering method based on improved topology data analysis according to claim 6 is characterized in that: Perform cluster analysis on n groups of one-dimensional vectors and obtain the clustering results of distribution transformer loads, as follows: (1) Use the elbow rule to determine the optimal number of clusters k for the K-means clustering algorithm; (2) Take n groups of one-dimensional vectors as sample data, and randomly select k initial data from the sample data as the mean vector, which is the cluster center; (3) Calculate the distance between each sample and each mean vector, and divide the clusters according to the distance; (4) Recalculate the mean vector of each cluster. If any mean vector is updated, replace the mean vector of the previous round and obtain a new mean vector. Return to step (3). If the mean vectors are no longer updated, output the cluster division result, that is, the clustering result of the distribution transformer load.
8. A distribution transformer load clustering system based on improved topological data analysis, characterized in that: include The time series data feature mining module is used to perform polar coordinate mapping on the active power time series data of different users in the same time period and station area in chronological order to obtain point cloud data sets of different users. The topological structure information of the point cloud data sets is extracted using the continuous coherence method to obtain continuous life and death graphs of different users, where the horizontal axis of the continuous life and death graph is the time when the topological feature appears, and the vertical axis is the time when the topological feature disappears. The topological data analysis module is used to preprocess the persistent life and death image, introduce a variable pixel parameter p, divide the preprocessed persistent life and death image into pixel blocks, introduce a weighting function to convert the persistent life and death image into a scalar function, filter topological noise, integrate the scalar function within the pixel block, and use the integral value as the value of the corresponding pixel block. All pixel blocks are sequentially combined to obtain a column matrix pixel map, i.e., a persistent image. The distribution transformer load clustering analysis module is used to flatten all pixel blocks after sequential combination to obtain n groups of one-dimensional vectors, perform cluster analysis, and obtain the clustering results of the distribution transformer load.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Electroencephalogram signal continuous feature extraction method based on continuous homology
CN112183477A
Self-adaptive clustering method based on continuous coherence
CN116628532A
Feature analysis method, device and equipment for distribution transformer load data and storage medium
CN120317526A