Power load curve clustering method, system and device based on improved density peak clustering algorithm and storage medium
By improving the density peak clustering algorithm and combining the local density estimation methods of K-nearest neighbors and natural nearest neighbors, the problem of uneven distribution of user types in power load curve clustering is solved, the clustering accuracy and stability are improved, and better power system management and planning are supported.
Patent Information
- Application Number
- CN202510681599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-10-31
Smart Images

Figure CN120873650A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart distribution networks, and specifically to a method, system, device, and storage medium for clustering power load curves based on an improved density peak clustering algorithm. Background Technology
[0002] With the rapid development of information technology and intelligentization in active distribution networks, the power industry is undergoing unprecedented transformation. The massive amounts of electricity big data collected by smart meters are a key driving force behind this transformation, profoundly changing the operation and management models of power systems. Analysis of user-side electricity load characteristics is crucial for the planning and operation of smart grid systems and distribution network management. It helps retailers and distribution network operators improve their understanding of consumer electricity consumption behavior, thereby identifying the social attributes of electricity users and better implementing energy policies and infrastructure planning strategies. How to efficiently perform cluster analysis of electricity load curves in a data-driven manner is a critical issue in analyzing electricity load characteristics.
[0003] Load characteristic analysis is mainly based on cluster analysis of load curves, but existing technologies have certain problems. Existing clustering algorithms are mainly divided into supervised and unsupervised methods. Supervised learning mainly includes Bayesian models based on statistical principles, and Extreme Learning Machines, Support Vector Machines (SVM), and decision trees based on artificial neural networks. These methods often require a large amount of labeled data, and their quality is highly dependent on the training labels of the samples, making it difficult to obtain large amounts of labeled data in real-world scenarios. Unsupervised learning mainly includes K-means, DBSCAN, fuzzy C-means clustering, and Gaussian Mixture Model (GMM) clustering. While these methods do not require a large amount of labeled data, they each have their own problems: the K-means algorithm is easily affected by the initial cluster centers, and the method for determining the number of clusters needs improvement; the density-based DBSCAN algorithm is easily affected by parameter values; GMM suffers from sensitivity to cluster number selection and slow convergence speed; and the fuzzy C-means algorithm performs poorly when the sample distribution is imbalanced.
[0004] Density peaks clustering (DPC), proposed by researchers in 2014, has relatively few hyperparameters and is less sensitive to parameter values. It can cluster clusters of arbitrary shapes, offering advantages over other clustering algorithms and has been widely applied in traffic flow prediction, image recognition, and other fields. However, DPC still has shortcomings when applied to electricity load curve clustering. With the deepening construction of new power distribution networks and the refinement of user-side management, the number of users with different electricity consumption patterns varies significantly, resulting in an uneven overall distribution. When clustering daily load curves, the density differences between clusters are large, leading to an uneven overall distribution. The local density obtained from the original cutoff distance is insufficient to characterize the density between sample points, resulting in errors in the allocation of remaining load curves and affecting the allocation of subsequent sample points.
[0005] To address the problems existing in current power load curve clustering methods in novel power distribution network systems, this invention proposes a power load curve clustering method based on an improved density peak clustering algorithm. In this method, a novel local density estimation approach is proposed based on the concepts of K-nearest neighbors and natural nearest neighbors. This approach calculates the topological relationships and distances between the daily load clustering curves to obtain clustering decision values. Based on these decision values, daily load curves are selected as cluster centers, and finally, the remaining daily load curves are allocated to achieve the final clustering. Summary of the Invention
[0006] To address the aforementioned technical issues, a power load curve clustering method based on an improved density peak clustering algorithm is proposed. This method includes sampling and segmenting power load data to construct a daily load curve set, and performing standardization processing on each daily load curve to generate a standardized sample set.
[0007] Based on the standardized sample set, the sample distance relationship of the daily load curve set is calculated, and the corresponding distance matrix is constructed;
[0008] Based on the distance matrix, a local density estimation model is established, and the density estimate of each sample curve is calculated;
[0009] Based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, a set of clustering decision values is constructed, and several sample curves with clustering decision values higher than the threshold are selected as cluster centers.
[0010] Using the cluster centers as a reference, the remaining sample curves are sequentially assigned to the corresponding clusters based on the density estimate and the distance to each cluster center, thus completing the sample clustering.
[0011] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of sampling and segmenting the power load data to construct a daily load curve set, and performing standardization processing on each daily load curve to generate a standardized sample set includes segmenting the power load data according to a set time interval, with each segment containing a sequence of equally spaced sampling points within 24 hours, and converting it into a daily load curve.
[0012] Then, subtract the average value of the current curve from the original load value of all sampling points in each daily load curve, and divide by the standard deviation of the current curve to obtain the standardized load curve set.
[0013] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of calculating the sample distance relationship of the daily load curve set and constructing the corresponding distance matrix based on the standardized sample set includes:
[0014] Calculate the Euclidean distance between the daily load curves in the daily load curve set, generate the distance matrix of the daily load curve set, and dynamically determine the optimal number of neighbors K by analyzing the statistical distribution characteristics of the distance matrix, and calculate the local density estimate of each daily load curve.
[0015] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of establishing a local density estimation model based on the distance matrix and calculating the density estimate value of each sample curve includes:
[0016] Based on the distance matrix, determine the K nearest neighbors for each daily load curve, which are the K nearest neighbors.
[0017] Determine whether each pair of samples is each other's K nearest neighbor. If they satisfy the mutual inclusion relationship, they are determined to be a pair of natural nearest neighbors.
[0018] The local density estimate of the sample point is calculated by weighting the distances between each sample in the K-nearest neighbor set and the target sample, and by combining the number of natural nearest neighbors as a correction factor.
[0019] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of constructing a clustering decision value set based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, and selecting several sample curves with clustering decision values higher than a threshold as cluster centers includes:
[0020] A clustering decision value set is constructed for all samples, and the selection threshold or selection ratio of cluster centers is determined based on the degree of difference between the clustering decision values of each sample in the set.
[0021] When a cluster decision value is higher than that of other samples, the current sample is directly selected as the cluster center;
[0022] When several clustering decision values are close, multiple samples are selected as cluster centers according to preset sorting rules or quantity limits.
[0023] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of constructing a clustering decision value set based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, and selecting several sample curves with clustering decision values higher than a threshold as cluster centers, further includes:
[0024] Based on the local density estimate and relative distance value of each sample curve, the corresponding clustering decision value is calculated;
[0025] The clustering decision values of all samples are compared, and the set of sample curves with the largest clustering decision value is identified as the cluster center, and the initial structure of the corresponding cluster is established.
[0026] As a preferred embodiment of the power load curve clustering method based on the improved density peak clustering algorithm described in this invention, the step of using cluster centers as references and sequentially assigning the remaining sample curves to the corresponding cluster clusters according to the density estimate and the distance to each cluster center to complete the sample clustering partitioning includes:
[0027] For each daily load curve that is not assigned to a cluster, identify all cluster center samples whose local density values are greater than the current curve.
[0028] From the current cluster center samples, select the cluster corresponding to the sample with the smallest Euclidean distance to the current curve;
[0029] The current daily load curve is assigned to the cluster to which the selected cluster center sample belongs;
[0030] Repeat the process for the remaining unclassified daily load curves until all daily load curves are classified.
[0031] Another objective of this invention is to provide a power load curve clustering system based on an improved density peak clustering algorithm. This invention solves the problem in existing power load curve clustering systems that it is difficult to accurately identify cluster centers and effectively divide clusters when dealing with uneven distribution of user types and significant differences in sample density.
[0032] As a preferred embodiment of the power load curve clustering system based on the improved density peak clustering algorithm described in this invention, it is characterized by including a data preprocessing module, a distance calculation and density modeling module, a cluster center identification module, and a sample classification module.
[0033] The data preprocessing module divides the collected power load data into daily load curves according to a set time interval, and performs standardization processing on each daily load curve to generate a standardized daily load curve set.
[0034] The distance calculation and density modeling module is based on the standardized daily load curve set, calculates the Euclidean distance between any samples, constructs a distance matrix, and establishes a local density estimation model to obtain the local density estimate of each sample.
[0035] The cluster center identification module constructs clustering decision values based on the relative distance between local density estimates and samples, and determines cluster center samples based on the clustering decision values.
[0036] The sample classification module assigns unclassified samples to the corresponding clusters based on their distance and density to each cluster center, thus completing the sample clustering.
[0037] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the power load curve clustering method based on an improved density peak clustering algorithm.
[0038] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the power load curve clustering method based on an improved density peak clustering algorithm.
[0039] The beneficial effects of this invention are as follows: This invention proposes an improved local density estimation method by combining KNN and the natural nearest neighbor algorithm, which effectively characterizes the nonlinear relationship of power load curves under large-scale datasets, mines the density relationship of each curve when the dataset is unevenly distributed, solves the misclassification problem of power load curve sets when the density difference is too large, effectively improves the clustering accuracy of load curves, and provides reliable identification of the social attributes of power users for the power supply side, so as to better implement energy planning strategies. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 The following is a flowchart of an overall method for clustering power load curves based on an improved density peak clustering algorithm, provided as an embodiment of the present invention.
[0042] Figure 2 This is a visualization of the clustering results of the DPC algorithm for a power load curve clustering method based on an improved density peak clustering algorithm, provided as an embodiment of the present invention.
[0043] Figure 3 This is a visualization of the clustering results of the DPC algorithm for a power load curve clustering method based on an improved density peak clustering algorithm, provided as an embodiment of the present invention.
[0044] Figure 4 The image shows the clustering results of the iDPC algorithm in a power load curve clustering method based on an improved density peak clustering algorithm, provided in an embodiment of the present invention. In this image, a represents the first type of curve shape, b represents the second type of curve shape, c represents the third type of curve shape, and d represents the fourth type of curve shape.
[0045] Figure 5 The GMM algorithm clustering result is provided in an embodiment of the present invention for a power load curve clustering method based on an improved density peak clustering algorithm, wherein a is the first type of curve shape, b is the second type of curve shape, c is the third type of curve shape, and d is the fourth type of curve shape.
[0046] Figure 6 The K-medoids algorithm clustering result is provided in an embodiment of the present invention for a power load curve clustering method based on an improved density peak clustering algorithm, wherein a is the first type of curve shape, b is the second type of curve shape, and c is the third type of curve shape.
[0047] Figure 7 The image shows the K-DPC algorithm clustering result in a power load curve clustering method based on an improved density peak clustering algorithm, provided in an embodiment of the present invention. Here, a represents the first type of curve shape, b represents the second type of curve shape, and c represents the third type of curve shape.
[0048] Figure 8 The present invention provides a method for clustering power load curves based on an improved density peak clustering algorithm, wherein a is a first type of curve shape, b is a second type of curve shape, and c is a third type of curve shape. Detailed Implementation
[0049] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0050] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a power load curve clustering method based on an improved density peak clustering algorithm, including:
[0051] S1. Sample and segment the power load data to construct a set of daily load curves, and perform standardization processing on each daily load curve to generate a standardized sample set.
[0052] The input daily power load data is normalized to eliminate dimensional differences. The Z-score standardization method is used to ensure that the load data at each point in the load curve follows a distribution with a mean of 0 and a standard deviation of 1. Based on the data sampling frequency, it is converted into a daily load curve set, with each row of data representing the standardized power load information.
[0053] First, the power load data is converted into a daily power load curve based on the sampling interval. Specifically, based on the sampling interval T, every 24 / T sampling points are truncated, and each segment of power load data is represented as a daily power load curve, that is, 24-hour power load data.
[0054] After processing the daily power load curve, the power load data will have a stronger periodic trend, which will help improve the clustering effect.
[0055] Preprocessing steps include outlier handling and data standardization. Outlier removal methods are used to handle outliers; if a load curve has missing or outlier data accounting for 10% or more of the total, that load curve is removed to prevent outliers from affecting the clustering results. To avoid the impact of amplitude differences on similarity metrics and improve data comparability, the Z-score standardization method is used to transform the daily load curves into a distribution with a mean of 0 and a standard deviation of 1.
[0056] The Z-score standardization method can be expressed as a formula as follows:
[0057]
[0058] Where x' is the z-score standardized number, μ is the mean of sequence x, and σ is the standard deviation of sequence x.
[0059] S2. Based on the standardized sample set, calculate the sample distance relationship of the daily load curve set and construct the corresponding distance matrix.
[0060] Calculate the Euclidean distance between the daily load curves in the daily load curve set, generate the distance matrix of the daily load curve set, and dynamically determine the optimal number of neighbors K by analyzing the statistical distribution characteristics of the distance matrix, and calculate the local density estimate of each daily load curve.
[0061] To visually represent the distances between daily load curves, the pairwise Euclidean distances between daily load curves in the daily load curve set are calculated based on the sampling interval of the daily load curves. This leads to the construction of the daily load curve sample Euclidean distance matrix D, the calculation formula of which is expressed as:
[0062]
[0063] Where, x a (i) and x b (i) represent the load data of the two daily load curves at time i, and T represents the sampling interval in hours. Assuming that there are a total of k daily load curves in the set, the Euclidean distance matrix D of the daily load curve samples is a k×k matrix.
[0064] S3. Based on the distance matrix, establish a local density estimation model and calculate the density estimate of each sample curve.
[0065] Based on the distance matrix, determine the K nearest neighbors for each daily load curve, which are the K nearest neighbors.
[0066] Determine whether each pair of samples is each other's K nearest neighbor. If they satisfy the mutual inclusion relationship, they are determined to be a pair of natural nearest neighbors.
[0067] The local density estimate of the sample point is calculated by weighting the distances between each sample in the K-nearest neighbor set and the target sample, and by combining the number of natural nearest neighbors as a correction factor.
[0068] In a preferred embodiment of this invention, the local density estimation model is calculated based on the Euclidean distance matrix D on the daily power load curves to obtain a local density estimate for each daily load curve. This characterizes the sparse density representation and nonlinear characteristics of each daily load curve within the set of daily load curves. The calculation formula is as follows:
[0069]
[0070] Wherein, KNN(x) i ) represents the relationship with sample point x i The set of K nearest neighbors with the smallest Euclidean distance, where the number of K nearest neighbors is 1%-2% of the total number of sample points, based on the DPC cutoff distance; d ij For sample point xi With sample point x j The Euclidean distance between them, d ik For sample point x i The distance between a point and its K nearest neighbors is given by NNN(i,j). NNN(i,j) represents the natural nearest neighbors (NNN) of a point, calculated as follows:
[0071]
[0072] When sample point x i With x j When sample points are within each other's K nearest neighbor neighborhood, these two sample points are a pair of natural nearest neighbors. By performing a weighted summation of the K nearest neighbor samples of a sample point, and taking into account the Euclidean clustering between sample points and the density of the cluster to which the sample belongs, the local density estimate of the sample point is adjusted according to the number of its natural nearest neighbor sample points.
[0073] In one optional embodiment of the local density estimation model in this invention, a fixed-radius neighborhood modeling method based on the kernel density function is used. This method can be used as a lightweight alternative or auxiliary initialization mechanism in some modules.
[0074] For each sample point x in the constructed Euclidean distance matrix D i Define a fixed distance threshold ε;
[0075] The set N of sample points falling within the neighborhood of this threshold is statistically analyzed. ε (x i The density is estimated using a kernel function, such as the Epanechnikov kernel or the Gaussian kernel, which takes the following form:
[0076]
[0077] The obtained ρ i The density estimate of sample points in the neighborhood of radius ε is used to replace the density index in the traditional DPC method. It is input into the γ value calculation stage and participates in the selection of cluster center candidates.
[0078] In this invention, the fixed-radius kernel density estimation method can be used as a sub-strategy in the distance calculation and density modeling module to quickly estimate local density when the sample size is small or in scenarios with high tolerance for accuracy (such as coarse-grained clustering of power users and pre-screening for anomaly detection).
[0079] Its output density value can also be directly involved in the construction process of clustering decision value γ, but it is generally not used in conjunction with the natural nearest neighbor mechanism, but rather as an independent strategy path.
[0080] For example, in the deployment of data gateways at the edge nodes of the distribution network, when the system performs clustering and classification operations on the initial load profiles of newly sampled users each day, this method can be used to make a low-latency preliminary judgment on their clustering tendency.
[0081] This result can be used in subsequent steps to quickly build a preliminary set of cluster centers, improving the loading efficiency of subsequent calculation modules.
[0082] It should be further explained that:
[0083] The density estimation mechanism based on K-nearest neighbors and natural nearest neighbors used in this invention has the following technical advantages compared to the fixed radius method:
[0084] Introducing a two-layer neighborhood structure (K-neighbors and mutually inclusive NNNs) enables finer-grained modeling of local sample relationships and cluster structure changes; it maintains the discriminative power of density indices and effectively supports γ value ranking even in scenarios with significant density gradient changes or blurred cluster boundaries; by using the maximum distance within the neighborhood as a scale normalization method, it avoids fixed threshold selection errors and improves the stability of cluster center selection; ultimately, it significantly improves the stability of density assessment and the overall robustness of clustering, especially suitable for distribution scenarios with highly varied daily load behaviors of power users.
[0085] S4. Based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, construct a set of clustering decision values, and select several sample curves with clustering decision values higher than the threshold as cluster centers.
[0086] A clustering decision value set is constructed for all samples, and the selection threshold or selection ratio of cluster centers is determined based on the degree of difference between the clustering decision values of each sample in the set.
[0087] When a cluster decision value is higher than that of other samples, the current sample is directly selected as the cluster center;
[0088] When several clustering decision values are close, multiple samples are selected as cluster centers according to preset sorting rules or quantity limits.
[0089] Based on the local density estimate and relative distance value of each sample curve, the corresponding clustering decision value is calculated;
[0090] The clustering decision values of all samples are compared, and the set of sample curves with the largest clustering decision value is identified as the cluster center, and the initial structure of the corresponding cluster is established.
[0091] In this invention, the step of constructing cluster decision values and selecting cluster centers is one of the core mechanisms of density peak clustering. Its role is to establish decision indicators based on the local density estimates and relative distances between samples obtained from previous calculations, screen out representative daily load curves as cluster centers, and thus guide the initial construction of clusters.
[0092] In this invention, a preferred embodiment of constructing clustering decision values and selecting several curves as cluster centers is as follows:
[0093] For each daily load curve sample x i Based on the local density estimate ρ calculated in the previous step i The density of the sample is compared with that of the remaining samples. The "relative distance" δ is defined. i The construction logic is as follows:
[0094] If x i For the point in the sample set with the largest local density estimate, its relative distance δ i Set to the maximum Euclidean distance between this sample and all other sample points, that is:
[0095]
[0096] Otherwise, its relative distance is defined as the nearest distance to all samples in the set whose local density is higher than its own, i.e.:
[0097]
[0098] It indicates that the local density is greater than x. i The distance between the point and the nearest point to the sample point.
[0099] This relative distance is used to measure the marginality and independence of the current sample within a high-density sample structure.
[0100] This invention uses the local density estimate ρ i With relative distance δ i Multiplying these values yields the clustering decision value γ used for determining cluster centers. The formula for calculating the clustering decision value γ is as follows:
[0101] γ i =ρ i δ i
[0102] The γ value of the cluster center should be significantly greater than that of other daily load curves.
[0103] This clustering decision value is used in the actual clustering process to simultaneously measure the significance of a sample in terms of density and its relative separation in the sample distribution space. Samples with higher values tend to be located in local high-density centers and maintain a certain distance from other centers, thus being preferentially selected as cluster centers.
[0104] Complete all samples γ i After the values are calculated, the cluster center selection process begins. This process includes the following sub-steps:
[0105] Sort the set of clustering decision values {γ1,γ2,...,γn} of all samples.
[0106] Set selection criteria: When a certain sample point x i γ i If the value is significantly higher than other sample points (e.g., higher than the mean plus standard deviation of the set of all sample γ values), the sample is directly designated as the cluster center.
[0107] If there are multiple samples with γ i If the values are close to each other (i.e., fail to form a significant boundary), then the sample density value ρ is used as the basis for classification. i A second sorting process is then performed, prioritizing those with higher densities.
[0108] The number of cluster centers can be set to a dynamic upper limit, such as 5% of all samples, to ensure that the center points are not too densely distributed while meeting the representativeness requirements.
[0109] Once the cluster centers are selected, the selected samples constitute the initial center points of the clusters and serve as the reference basis for allocating the remaining samples in the sample classification module. This will guide the subsequent determination of the affiliation of the remaining unclassified daily load curve samples. The selected cluster center samples will be assigned fixed numbers and used as identifiers for each cluster.
[0110] Constructing clustering decision values and selecting several curves as cluster centers is an optional embodiment of the present invention, which is a cluster center selection method based on density ranking index.
[0111] For each sample curve, calculate its corresponding local density estimate. Then, sort all the sample density estimates from highest to lowest and assign each sample a density ranking number. The sample with the highest density is ranked first, and so on, with higher numbers for lower densities.
[0112] Next, calculate the relative distance value for each sample, which is the minimum distance between that sample and all samples with a higher density (if its density is the highest, then the maximum distance with all other samples is used as a reference). Sort all samples by relative distance values from largest to smallest to generate a relative distance ranking number. Similarly, samples with larger relative distances are ranked higher.
[0113] Subsequently, the density ranking number and the relative distance ranking number of each sample are summed to obtain a ranking index used to determine the priority of cluster centers. The smaller the value of this ranking index, the more outstanding the sample is in both local density and relative distance dimensions, that is, it has both high density and strong spatial separation, and is therefore more representative.
[0114] In the actual process of determining cluster centers, an upper limit can be set for the maximum number of cluster centers. Following the ascending order of the aforementioned ranking indices, a number of samples, not exceeding the set limit, are selected as the final cluster centers. These samples will serve as the core representatives of the initial clusters, guiding the subsequent assignment of other samples.
[0115] In an optional embodiment of the present invention, the determination of cluster centers no longer involves constructing a clustering decision value by multiplying the density estimate and the relative distance value. Instead, a cluster center selection method based on a ranking index is employed. This method ranks each sample separately in both the density and relative distance dimensions, and then sums the two ranking numbers to obtain a ranking index. Samples with smaller ranking indices are considered to have greater potential to become cluster centers.
[0116] The optional embodiments do not need to consider the scaling or normalization of density and distance, and have a certain ability to resist outliers under the sorting mechanism.
[0117] In a preferred embodiment of the invention, the selection of cluster centers is based on a clustering decision value calculated by multiplying the local density estimate of each sample by its relative distance value. The density value of a sample represents its concentration in a local region, while the distance value represents its relative separation from other high-density samples. The clustering decision value constructed by multiplying these two values simultaneously captures the dual characteristics of "representativeness" and "independence."
[0118] It should be further explained that:
[0119] The preferred embodiment of the present invention accurately captures the dual characteristics of each sample point in density distribution and spatial topology by constructing a product model of local density and relative distance, thereby enabling more stable and accurate identification of representative cluster center samples.
[0120] Compared with the sorting index strategy in the optional implementation, the technical solution of the present invention has stronger structural separation ability and discrimination efficiency when dealing with problems such as uneven density, fuzzy cluster boundaries and complex power load curve structure, and further improves the accuracy and stability of clustering effect.
[0121] S5. Using the cluster centers as a reference, the remaining sample curves are sequentially assigned to the corresponding cluster centers based on the density estimate and the distance to each cluster center, thus completing the sample clustering.
[0122] For each daily load curve that is not assigned to a cluster, identify all cluster center samples whose local density values are greater than the current curve.
[0123] From the current cluster center samples, select the cluster corresponding to the sample with the smallest Euclidean distance to the current curve. Assign the current daily load curve to the cluster of the selected cluster center sample. Repeat this process for the remaining daily load curves that have not yet been assigned to clusters until all daily load curves have been assigned.
[0124] Example 2, refer to Figure 2 Figure * shows the second embodiment of the present invention, which provides a power load curve clustering method based on an improved density peak clustering algorithm. To verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0125] Two-dimensional visualization of over a thousand daily load curves selected from traditional DPC algorithms (e.g.) Figure 2 As shown in the figure, the clustering distribution characteristics of different user types can be observed. In the figure, the blue, green, and yellow clusters represent over 800 load curves from factories, office buildings, and supermarkets, respectively, while the red cluster corresponds to over 200 primary and secondary school user data. Because the red cluster has a significantly smaller sample size and is closer to the centroids of the yellow and green clusters, aliasing occurs in the feature space. It is worth noting that in the two-dimensional projection, the coordinate axes are dimensionless, and the Euclidean distance between points only represents the relative similarity of the original high-dimensional space—the closer the distance, the higher the consistency of the load curve shape. Simulation experiments show that when the data scale expands to 100,000 records, the proportion of misassigned samples in the aliasing region will surge.
[0126] The improved local density estimation method employs a weighted aggregation strategy on the K nearest neighbors of a sample point, while simultaneously considering both the Euclidean distance between sample points and the density characteristics of the cluster to which the sample belongs. This allows for dynamic adjustment of the local density estimate based on the number of natural neighbors of each sample point. This improvement effectively avoids the problem of inaccurate local density estimation in accurately depicting the relationships between samples when the sample point distribution is sparse, thus preventing misassignment. Figure 3 Presented with Figure 2 The distribution results in two-dimensional space after clustering the same daily load curve data using the iDPC (Improved Density Peak Clustering) algorithm incorporating improved local density estimation are compared with... Figure 2 As can be seen from the comparison, the iDPC algorithm successfully solved the problem of sample points overlapping between the red cluster and the yellow and green clusters.
[0127] This study uses commercial user load data publicly available from the U.S. Department of Energy's Open Energy Information (OpenEI) website as an application case. The OpenEI dataset covers commercial user load information from 16 industries across multiple regions in the United States. Data for each user is sampled at one-hour intervals, with 24 time points collected daily, totaling 4800 daily load curves. Commonly used clustering evaluation metrics, such as the sum of squares of errors (SSE), the Davies Bouldin Index (DBI), and the Silhouette Coefficient (SC), are selected to measure the validity of the clustering results. If the load curve set is divided into K clusters, the above metrics are defined as follows:
[0128] This study uses commercial user load data publicly available from the U.S. Department of Energy's Open Energy Information (OpenEI) website as an application case. The OpenEI dataset covers commercial user load information from 16 industries across multiple regions in the United States. Data for each user is sampled at one-hour intervals, with 24 time points collected daily, totaling 4800 daily load curves. Commonly used clustering evaluation metrics, such as the sum of squares of errors (SSE), the Davies Bouldin Index (DBI), and the Silhouette Coefficient (SC), are selected to measure the validity of the clustering results. If the load curve set is divided into K clusters, the above metrics are defined as follows:
[0129] 1) SSE:
[0130]
[0131] In the formula: x i * C i Cluster center; x ic C i Non-central sample points within a cluster. The SSE index reflects the degree of cohesion of each cluster; the smaller the value, the more compact the clusters and the better the clustering effect.
[0132] 2) DBI:
[0133]
[0134] In the formula: ||x represents the average distance from each sample point in the i-th cluster to the cluster center; i * -x j* ||2 represents the distance from the i-th cluster center to the j-th cluster center. DBI comprehensively considers the similarity between samples within a cluster and the difference between samples between clusters; the smaller the value, the better the clustering effect.
[0135] 3)SC:
[0136]
[0137] In the formula: b(x) i ) represents the sample point x i Minimum average distance to sample points within other clusters; a(x i ) represents x i The average distance to other sample points within the same cluster. The value of SC ranges from -1 to 1. A larger value indicates that the clusters are more distant from each other, the differences within the cluster are smaller, and the clustering effect is better.
[0138] Gaussian Mixture Model (GMM), K-medoids clustering, K-DPC (Kernel Density Estimation Density Peak Clustering), and Density Sort Index Density Peak Clustering (D-DPC) based on density ranking index cluster centers were selected as comparison algorithms. Each algorithm was applied to the OpenEI public dataset to obtain the clustering results for reference. Figures 4-8 .
[0139] The clustering results of the iDPC algorithm are as follows: Figure 4 As shown in the figure, the clustering results are generally accurate, but there are significant differences in morphology between different clusters, and the four curve shapes show obvious differences. Figure 4 (a) Class and Figure 4 (b) The load curve exhibits a bimodal shape, in which Figure 4 (a) is the hotel category, whose peak load curves for morning and evening are similar; Figure 4 (b) is the load curve for apartments, where the peak load in the evening is significantly higher than that in the morning. Figure 4 (c) class and Figure 4 (d) The load curves all exhibit a single-peak shape, among which Figure 4 (c) The peak duration of class (c) is slightly longer than Figure 4 (d) The peak period is roughly distributed between 9:00 and 17:00. This category mainly includes places where people work during the day and for a long time, such as hospitals, middle schools, and office buildings. Figure 4(d) The load curve only covers primary schools, and its peak period is concentrated between 9:00 and 16:00. Clustering results using GMM, K-medoids, K-DPC, and D-DPC algorithms are as follows: Figure 5 , Figure 6 , Figure 7 and Figure 5 As shown, all four algorithms exhibit a mixture of two curve shapes, resulting in poor overall performance. Table 1 presents the performance of three clustering evaluation metrics. Comparing the metrics for clustering effectiveness, iDPC has the lowest SSE and DBI metrics, indicating the highest degree of cluster cohesion and the most compact clusters; its SC is the highest, indicating the highest degree of sparsity between clusters. Therefore, iDPC can effectively divide the clusters correctly among the four clustering methods, with the largest inter-cluster distance and the smallest intra-cluster distance, resulting in the best clustering effect.
[0140] Table 1 Comparison of clustering evaluation metrics for different algorithms
[0141] Clustering evaluation metrics Optimal number of clusters SSE DBI SC iDPC 4 8440.9 0.7463 0.4495 K-DPC 3 9240.9 0.8327 0.3941 K-Medoids 3 10028.6 1.1048 0.4107 GMM 4 9802.7 1.0527 0.3887 D-DPC 3 1.5515 1.4440 0.3279
[0142] The classic DPC algorithm requires constructing a fully connected distance matrix and performing a global sort, resulting in a time complexity of O(n log n). 2 In the proposed iDPC algorithm, the improved local density estimation requires calculating the density weighted sum within the K nearest neighbors of each sample point, with a time complexity of O(n^2). 2 +n 2 Therefore, the time complexity of the algorithm is O(n^2). 2 +n 2 )~O(n 2 Compared to the original DPC algorithm, by introducing a controllable increase in computational complexity, the clustering effect is significantly improved, and no significant waste of computational resources is caused in the case of larger-scale data.
[0143] Practical applications show that this method, by employing an improved density peak clustering algorithm, takes into account the problem that when the uneven distribution of data leads to excessive density differences between clusters, it cannot adequately reflect the density and nonlinear characteristics between sample points. It avoids the problem that when the local density estimation cannot accurately represent the relationship between sample points when the sparsity of data points varies greatly, resulting in incorrect allocation of other sample points, thereby improving the effectiveness of clustering.
[0144] Example 3 is the third embodiment of the present invention, which differs from the previous two embodiments in that:
[0145] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0147] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0148] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0149] Example 4 is the fourth embodiment of the present invention. This embodiment provides a power load curve clustering system based on an improved density peak clustering algorithm, including a data preprocessing module, a distance calculation and density modeling module, a cluster center identification module, and a sample classification module.
[0150] The data preprocessing module divides the collected power load data into daily load curves according to a set time interval, and performs standardization processing on each daily load curve to generate a standardized daily load curve set.
[0151] The distance calculation and density modeling module is based on the standardized daily load curve set, calculates the Euclidean distance between any samples, constructs a distance matrix, and establishes a local density estimation model to obtain the local density estimate of each sample.
[0152] The cluster center identification module constructs clustering decision values based on the relative distance between local density estimates and samples, and determines cluster center samples based on the clustering decision values.
[0153] The sample classification module assigns unclassified samples to the corresponding clusters based on their distance and density to each cluster center, thus completing the sample clustering.
[0154] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for clustering power load curves based on an improved density peak clustering algorithm, characterized in that: include, The power load data is sampled and segmented to construct a set of daily load curves, and each daily load curve is standardized to generate a standardized sample set. Based on the standardized sample set, the sample distance relationship of the daily load curve set is calculated, and the corresponding distance matrix is constructed; Based on the distance matrix, a local density estimation model is established, and the density estimate of each sample curve is calculated; Based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, a set of clustering decision values is constructed, and several sample curves with clustering decision values higher than the threshold are selected as cluster centers. Using the cluster centers as a reference, the remaining sample curves are sequentially assigned to the corresponding clusters based on the density estimate and the distance to each cluster center, thus completing the sample clustering.
2. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 1, characterized in that: The process of sampling and segmenting power load data to construct a set of daily load curves, and performing standardization processing on each daily load curve to generate a standardized sample set includes segmenting the power load data according to a set time interval, with each segment containing a sequence of equally spaced sampling points within 24 hours, and converting it into a daily load curve. Then, subtract the average value of the current curve from the original load value of all sampling points in each daily load curve, and divide by the standard deviation of the current curve to obtain the standardized load curve set.
3. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 2, characterized in that: The process of calculating the sample distance relationships of the daily load curve set based on the standardized sample set and constructing the corresponding distance matrix includes: Calculate the Euclidean distance between the daily load curves in the daily load curve set, generate the distance matrix of the daily load curve set, and dynamically determine the optimal number of neighbors K by analyzing the statistical distribution characteristics of the distance matrix, and calculate the local density estimate of each daily load curve.
4. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 3, characterized in that: The process of establishing a local density estimation model based on the distance matrix and calculating the density estimate for each sample curve includes: Based on the distance matrix, determine the K nearest neighbors for each daily load curve, which are the K nearest neighbors. Determine whether each pair of samples is each other's K nearest neighbor. If they satisfy the mutual inclusion relationship, they are determined to be a pair of natural nearest neighbors. The local density estimate of the sample point is calculated by weighting the distances between each sample in the K-nearest neighbor set and the target sample, and by combining the number of natural nearest neighbors as a correction factor.
5. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 4, characterized in that: The step of constructing a clustering decision value set based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, and selecting several sample curves with clustering decision values higher than a threshold as cluster centers, includes... A clustering decision value set is constructed for all samples, and the selection threshold or selection ratio of cluster centers is determined based on the degree of difference between the clustering decision values of each sample in the set. When a cluster decision value is higher than that of other samples, the current sample is directly selected as the cluster center; When several clustering decision values are close, multiple samples are selected as cluster centers according to preset sorting rules or quantity limits.
6. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 5, characterized in that: The step of constructing a clustering decision value set based on the local density estimate of each sample curve and its relative distance to other samples in the distance matrix, and selecting several sample curves with clustering decision values higher than a threshold as cluster centers, further includes: Based on the local density estimate and relative distance value of each sample curve, the corresponding clustering decision value is calculated; The clustering decision values of all samples are compared, and the set of sample curves with the largest clustering decision value is identified as the cluster center, and the initial structure of the corresponding cluster is established.
7. The power load curve clustering method based on an improved density peak clustering algorithm as described in claim 6, characterized in that: The step of using cluster centers as references, and then sequentially assigning the remaining sample curves to the corresponding clusters based on the density estimates and their distances to each cluster center, to complete the sample clustering process includes: For each daily load curve that is not assigned to a cluster, identify all cluster center samples whose local density values are greater than the current curve; From the current cluster center samples, select the cluster corresponding to the sample with the smallest Euclidean distance to the current curve; The current daily load curve is assigned to the cluster of the selected cluster center sample. This process is repeated for the remaining daily load curves that have not yet been assigned to clusters until all daily load curves have been assigned.
8. A power load curve clustering system based on an improved density peak clustering algorithm, employing the power load curve clustering method based on an improved density peak clustering algorithm as described in any one of claims 1 to 7, characterized in that, It includes: a data preprocessing module, a distance calculation and density modeling module, a cluster center identification module, and a sample classification module; The data preprocessing module divides the collected power load data into daily load curves according to a set time interval, and performs standardization processing on each daily load curve to generate a standardized daily load curve set. The distance calculation and density modeling module is based on the standardized daily load curve set, calculates the Euclidean distance between any samples, constructs a distance matrix, and establishes a local density estimation model to obtain the local density estimate of each sample. The cluster center identification module constructs clustering decision values based on the relative distance between local density estimates and samples, and determines cluster center samples based on the clustering decision values. The sample classification module assigns unclassified samples to the corresponding clusters based on their distance and density to each cluster center, thus completing the sample clustering.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the power load curve clustering method based on the improved density peak clustering algorithm according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the power load curve clustering method based on the improved density peak clustering algorithm according to any one of claims 1 to 7.
Citation Information
Cited By
Power cable insulation state evaluation method and system based on multi-dimensional feature fusion
CN121299386A
Track clustering method and device based on space-time mahalanobis distance and density peak value
CN121524664A
Power load curve coding and clustering method and system, storage medium and terminal
CN122020228A