Typical load curve extraction method and system based on improved DPC algorithm

By adaptively determining the initial cluster centers and introducing cluster cross density and boundary density parameters to correct the initial clusters, the DPC algorithm is improved, which solves the uncertainty and allocation errors of the traditional DPC algorithm when manually selecting cluster centers, improves the accuracy and stability of the DPC algorithm in extracting typical load curves, and is suitable for multi-dimensional data sets and actual data sets.

CN120670880APending Publication Date: 2025-09-19STATE GRID JIBEI ELECTRIC POWER CO LTD TANGSHAN POWER SUPPLY CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510738869.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional DPC algorithms require manual selection of cluster centers when extracting typical load curves, which leads to subjective uncertainty and allocation errors, resulting in poor accuracy and stability of clustering results, making it difficult to meet the requirements of flexible resource identification.

Method used

The method of adaptively determining the initial cluster center is adopted. The data features are identified by calculating the local density and relative distance, and the initial cluster center is automatically selected. The cluster cross density and boundary density parameters are introduced to correct the initial cluster and improve the allocation error.

Benefits of technology

The automation level and clustering accuracy of the DPC algorithm are improved, subjective errors are reduced, and typical load curves can be extracted more accurately. It is applicable to multi-dimensional data sets and actual data sets, verifying the effectiveness of the improved strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670880A_ABST
    Figure CN120670880A_ABST
Patent Text Reader

Abstract

A typical load curve extraction method based on an improved DPC algorithm is characterized by comprising the following steps that a daily load curve of a typical load is collected, and the n-dimensional daily load curve is described as a data point containing n data; calculating local density rho and relative distance delta of all data points, and determining an initial clustering center; and defining a clustering cross density Ji, j and a clustering boundary density Bi, j of the two clusters Ci and Cj, so as to correct the initial cluster and output a clustering result. According to the method, the DPC algorithm is improved, so that the accuracy and the stability during typical load curve extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a field, and more specifically, to a typical load curve extraction method and system based on an improved DPC algorithm. Background Art

[0002] Load curves are essential data for power system operation analysis. They are massive and complex in form. To accurately describe load fluctuations and support the optimal regulation of flexible resources, clustering algorithms are needed to extract the most representative load curves from these massive load curves. Research on load clustering focuses on the following three areas:

[0003] First, clustering algorithms are modified based on the objectives and datasets to enhance their applicability and accuracy. This is currently the most common approach to load clustering research. This paper studies the optimal K value selection for the K-means clustering algorithm (K-means), taking into account the uncertainty of demand response load alignment. The deep differencing method is used to determine the optimal number of clusters for the dataset. The K-means algorithm is then combined with the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm and the Agglomerative Hierarchical Clustering (AHC) algorithm. Experimental comparisons are conducted with traditional K-means, DBSCAN, and AHC algorithms, concluding that the K-means-AHC algorithm outperforms the other algorithms. Based on the improved K-means clustering method for distribution network load forecasting, this paper proposes an adaptive k-means++ load characteristic clustering algorithm. This algorithm automatically determines the optimal number of clusters through iterative graph segmentation, avoiding bias caused by manual setting and improving classification accuracy. A case study demonstrates the algorithm's high accuracy and robustness. The paper proposes a K-medios clustering algorithm based on the analysis of power load characteristics using the adaptive k-means++ algorithm. The number of clusters in this algorithm is determined by the cluster radius, and simulation results verify that this clustering algorithm is effective in clustering load model parameters. The paper Clustering of Uncertain Load Model Parameters with K-medoids Algorithm improves the K-means algorithm by combining the parallel computing characteristics of the cloud platform to improve its clustering effect on massive electricity consumption data and achieve efficient analysis of user behavior. The paper Research on Residential Electricity Consumption Behavior Analysis Model Based on Cloud Computing proposes a new user clustering method. Considering the massiveness, instability, and uncertainty of smart metering data, simulation results show that this method is superior to the traditional K-means algorithm in load analysis. The paper Multi-resolution load profile clustering for smart metering data proposes a clustering method based on dynamic characteristics, which is applied to the online rapid and refined assessment of frequency safety after anticipated faults. The literature proposes a low-order simulation frequency safety analysis method based on dynamic characteristic clustering, which uses a Gaussian mixture model to fit the user's daily load data and adopts a hierarchical clustering algorithm for cluster analysis. The results show that this method can effectively realize user clustering.The paper "Aversatile clustering method for electricity consumption pattern analysis in households" proposes a wind-solar power correlation probability interval prediction model based on similar day clustering. This model extracts similar day scenarios for combined wind and solar power output using correlation coefficients. Simulation results demonstrate that the model effectively extracts wind-solar power correlation features and outperforms existing models in prediction accuracy. The paper also proposes a short-term wind-solar power correlation probability interval prediction method based on similar day clustering and the WOA-BiLSTM-Copula algorithm. This method uses symbolic aggregation to approximate individual user load curves to reduce data dimensionality, and employs a Markov model to model the dynamic characteristics of electricity consumption behavior. Subsequently, a novel density peak point fast search clustering algorithm is used to rapidly cluster massive load data. The paper "Clustering of electricity consumption behavior dynamics toward big data applications" proposes a combined prediction method using ensemble clustering and an improved Markov chain. First, ensemble clustering is used to divide photovoltaic output into trend sequences and random sequences. Second, an improved Markov chain model is used to predict the trend portion. The prediction accuracy is improved by an average of 7.76% compared to the traditional Markov chain model, demonstrating the applicability of this method to different output types and geographic locations.

[0004] The second is to use different methods to extract features from the load curve to reduce data dimensionality and improve clustering effects. The paper proposes an SVD-KICIC clustering algorithm based on ensemble clustering and an improved Markov chain model for photovoltaic power short-term prediction, which combines singular value decomposition (SVD) with a weighted K-means clustering approach by integrating intra-cluster and inter-cluster distances (KICIC). This method uses SVD technology to identify effective load curve features, reduce data dimensionality, and obtain singular values ​​of the data to determine the weights of the KICIC algorithm. Simulation results show that this method improves the computational efficiency and clustering effect of the KICIC algorithm. The paper Method for Clustering Daily Load Curve Based on SVD-KICIC proposes a load curve clustering method based on discrete wavelet transform (DWT). This method first transforms the load curve through multi-level DWT to reduce the dimensionality, and then uses a clustering fusion algorithm to optimize the clustering results obtained in the previous stage. Experimental results show that this algorithm performs better than other algorithms. The paper A Used Load Curve Clustering Algorithm Based on Wavelet Transform extracts multiple daily load characteristic indicators from the user load curve and uses the Delphi method to configure the weights of each characteristic indicator to achieve the purpose of dimensionality reduction and shorten the algorithm running time. The paper Clustering Analysis of Daily Load Curve Based on Characteristic Indicator Dimensionality Reduction proposes a typical operation mode extraction method based on convolutional neural network and self-attention mechanism. Clustering is performed by introducing a feature clustering layer for joint optimization. The new energy-load-traditional energy combination mode indicator and clustering effect evaluation index are used to characterize the operation mode clustering results. The paper examined various dimensionality reduction methods for extracting typical operating modes of high-proportion renewable energy power systems based on a convolutional self-attention clustering algorithm, finding that an algorithm that integrates principal component analysis (PCA) dimensionality reduction achieved the best performance. The paper also proposed a method based on information entropy segmented aggregation approximation to re-express load fluctuation characteristics and cluster them using a spectral clustering algorithm. This method effectively preserves load characteristics while reducing dimensionality.The paper proposes a load classification method based on information entropy segmented aggregation approximation and spectral clustering, and proposes a line selection method based on parameter-optimized variational mode decomposition (VMD) and improved K-clustering criterion fusion. This method, combined with an improved K-clustering algorithm, achieves multi-criterion fusion, overcoming the limitations of a single criterion and demonstrating excellent resistance to harmonic and noise interference. The paper also proposes a distribution network fault line selection method based on parameter-optimized VMD and improved K-clustering criterion fusion. The method uses a self-organizing map neural network to study user load curves, performs low-dimensional mapping through the neural network, and combines the K-means method to achieve user classification.

[0005] The third approach is to improve load similarity measurement methods to make clustering results more accurate. A paper proposes a spatiotemporal joint clustering method for clustering power user load curves based on a self-organizing map neural network. This method divides time series into multiple subsequences through time window shifting, enabling clustering analysis of spatiotemporal joint data. A paper proposes a feature selection strategy based on user load characteristics for power transmission and transformation equipment status anomaly detection based on a spatiotemporal joint clustering method. The effectiveness of this strategy in improving load clustering accuracy is verified using real-world data. A paper proposes a feature selection strategy based on load curve morphological characteristics using adaptive segmented aggregation approximation dimensionality reduction. This method accurately describes curve morphological characteristics while improving algorithm execution speed. A paper proposes a typical load curve morphology clustering algorithm using adaptive segmented aggregation approximation, combining cosine distance and Gaussian kernel function to measure load curve similarity, achieving stable and effective load clustering. A paper proposes a load curve ensemble spectral clustering algorithm that considers dual-scale similarity. In substation clustering analysis, a Euclidean distance between load curves and substation user composition is used to measure inter-substation similarity. The literature compares existing load curve morphological similarity measurement methods based on the substation characteristics analysis of the multivariate clustering model and the two-stage clustering correction algorithm, and finds that the dynamic time normalization method has the most outstanding performance. The literature Experimental comparison of representation methods and distance measures for time series data [J]. Data Mining and Knowledge Discovery combines distance similarity and morphological feature similarity to measure the similarity of load curves, and realizes reasonable classification of users through clustering. The literature Clustering load profiles for demand response applications proposes the use of dynamic time normalization method to measure the sequence composed of feature points. This method accurately reflects the characteristics of time series while improving the efficiency of the algorithm. The literature proposes a clustering method based on time series variable density processing based on time series mining based on segmented time bending distance. This method can effectively improve the load clustering effect and truly reflect the electricity consumption characteristics of residential users.

[0006] To address the above problems, a typical load curve extraction method and system based on an improved DPC algorithm is urgently needed. Summary of the Invention

[0007] In order to solve the deficiencies in the prior art, the present invention provides a typical load curve extraction method and system based on an improved DPC algorithm.

[0008] The present invention adopts the following technical solutions.

[0009] The first aspect of the present invention relates to a typical load curve extraction method based on an improved DPC algorithm, the method comprising the following steps: collecting a daily load curve of a typical load, describing an n-dimensional daily load curve as a data point containing n data; calculating the local density ρ and relative distance δ of all data points, and determining the initial cluster center; defining two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results.

[0010] Collect daily load curves of typical loads and describe the n-dimensional daily load curve as a data point containing n data, including: collecting 340 days of load data of 20 household users with a sampling interval of 8s; expanding the sampling interval to 30min, and adding the loads of 20 household users at the same time as large load data; selecting 200 daily load curves with relatively similar loads in the data set as the load data set for testing.

[0011] 200 daily load curves with relatively similar loads in the data set are selected as the load data set for testing, including: selecting 20% ​​of the load data to increase the load at certain moments and marking them with the "Cluster 2" label, selecting 10% of the load data to reduce the load at certain moments and marking them with the "Cluster 3" label, and marking the remaining 70% of the load data that has not been modified with the "Cluster 1" label, thereby simulating the actual load data set with cluster labels.

[0012] Calculate the local density ρ and relative distance δ of all data points and determine the initial cluster center, including: the initial cluster center is:

[0013] Ω0=Ω ρ ∪Ω δ

[0014] Where,

[0015] is the average value of the local density of all data points, is the average value of the relative distances of all data points,

[0016] x i is the i-th data point in the load data set.

[0017] Define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results, including: calculating the cluster set The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p , p is the number of clusters; J is taken in descending order of local density of cluster centers p×p and B p×p If the elements in J i,j >B i,j , then delete the original two cluster centers, reselect the data point with the largest local density in the two clusters as the cluster center, and cluster C i and C j Merge into one cluster; calculate the new cluster cross density matrix and cluster boundary density matrix, and continue looping until the clusters can no longer be merged, stop the loop, and output the clustering results.

[0018] Calculate cluster sets The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p ,include:

[0019] Define two clusters C i and C j The cluster cross density is J i,j for

[0020] J i,j =max{ρ x |x∈Y i ∩Y j}

[0021] Where Y i It's Y i ∩Y j The xth data point and cluster C in i The distance of the data point is less than the cutoff distance d c A collection of data points,

[0022] Y j It's Y i ∩Y jThe xth data point and cluster C in j The distance of the data point is less than the cutoff distance d c A collection of data points,

[0023] Y i ∩Y j For two clusters C i and C j The data points at the critical position.

[0024] Calculate cluster sets The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p ,include:

[0025] Define cluster C i The boundary point set is B i ,

[0026] B i ={x j |x j =sort(ρ j )[:m]}

[0027] m=ceil[length(C i )×20%]

[0028] In the formula, sort(ρ j ) indicates that C i The local density of each data point in the ascending order, [:m] means taking the first m elements of the set, ceil means rounding up, length(C i ) means C i The number of data points in the cluster C i The boundary point is C i Data points where the local density is in the lower 20% range;

[0029] Define two clusters C i and C j Cluster boundary density B i,j for

[0030]

[0031] Where, is cluster C i The boundary point set B i The local density average of each data point in .

[0032] The second aspect of the present invention relates to a typical load curve extraction system based on an improved DPC algorithm, the system utilizing the implementation of the method described in the first aspect of the present invention; the system comprises an acquisition module, a calculation module and an output module; the acquisition module is used to acquire the daily load curve of the typical load and describe the n-dimensional daily load curve as a data point containing n data; the calculation module is used to calculate the local density ρ and relative distance δ of all data points to determine the initial cluster center; the output module is used to define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results. The third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.

[0033] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.

[0034] The beneficial effect of the present invention is that, compared with the prior art, a typical load curve extraction method and system based on an improved DPC algorithm in the present invention proposes a method for adaptively determining the initial clustering center to eliminate the subjective uncertainty of the traditional DPC algorithm that requires manual selection of clustering centers. This method can automatically select the initial clustering center by identifying the characteristics of the data, thereby improving the degree of automation and accuracy of the DPC algorithm; secondly, an initial clustering improvement strategy is proposed to correct the initial clustering. The proposed strategy can improve the allocation errors when allocating non-clustering centers; finally, multiple algorithms are used to cluster data sets of multiple dimensions and actual data sets, and the clustering results are analyzed and compared according to clustering evaluation indicators. The improved DPC algorithm and other algorithms are used to extract typical load curves of the actual data set to verify the effectiveness of the proposed DPC improvement strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is the overall architecture diagram of the present invention;

[0036] Figure 2 It is the cluster decision diagram of the present invention;

[0037] Figure 3 Schematic diagram of the strategy for improving the initial clustering;

[0038] Figure 4 Flowchart for improving the DPC algorithm;

[0039] Figure 5To improve the clustering effect diagram of DPC algorithm and other algorithms;

[0040] Figure 6 This is the clustering effect diagram for clustering the Spiral dataset;

[0041] Figure 7 This is the clustering effect diagram for clustering the Flame dataset;

[0042] Figure 8 This is the clustering effect diagram for clustering the Jain dataset;

[0043] Figure 9 This is the clustering effect diagram for clustering the Aggregation data set;

[0044] Figure 10 Comparison of clustering effects of four methods on different data sets;

[0045] Figure 11 This is the modified load data curve;

[0046] Figure 12 The graph shows the results of clustering the modified load curve using four algorithms;

[0047] Figure 13 The cluster center graph selected by the common DPC algorithm and the improved DPC algorithm in the decision graph in this clustering;

[0048] Figure 14-21 The following are the typical load curve sets and typical load curve diagrams obtained by the four algorithms. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present invention clearer and more accurate, the technical solutions of the present invention are described in detail below through multiple specific embodiments. The embodiments used in the present invention are only used to explain the present invention and are not intended to limit the content of the present invention.

[0050] In distribution networks, the typical load curve of a power system is an important description of load fluctuation patterns, reflecting the typical electricity consumption behavior in the area covered by the power system. Accurately extracting the most representative typical load curves for each power system is crucial for accurately identifying and optimizing flexibility resources.

[0051] The typical load curve is usually extracted using a clustering algorithm. Common clustering algorithms include DBSCAN, K-means, and DPC. Among them, the DPC algorithm is a density-based clustering algorithm. Compared with other algorithms, it has many advantages, such as the DPC algorithm can identify non-convex clusters, does not require the number of clusters to be specified in advance, simple parameter settings, and better performance when processing high-dimensional data. However, the traditional DPC algorithm has some problems in the application process, such as the need to manually select cluster centers, subjective uncertainty, and the possibility of allocation errors when allocating non-cluster center points. Therefore, in order to address the problem that the typical load curves extracted by common clustering algorithms are less representative and difficult to meet the requirements of flexible resource identification, the present invention aims to propose an improved DPC algorithm to improve its accuracy and stability in extracting typical load curves.

[0052] The present invention proposes a method for adaptively determining the initial cluster centers to eliminate the subjective uncertainty of the traditional DPC algorithm which requires manual selection of cluster centers. This method can automatically select the initial cluster centers by identifying the characteristics of the data, thereby improving the automation and accuracy of the DPC algorithm. Secondly, an initial clustering improvement strategy is proposed to correct the initial clustering. The proposed strategy can improve the allocation errors when allocating non-cluster centers, and obtain locally convex classes for judging intra-class features. Finally, multiple algorithms are used to cluster data sets of multiple dimensions and actual data sets. The clustering results are analyzed and compared according to clustering evaluation indicators. The improved DPC algorithm and other algorithms are used to extract typical load curves of the actual data set to verify the effectiveness of the proposed DPC improvement strategy.

[0053] A first aspect of the present invention relates to a typical load curve extraction method based on an improved DPC algorithm, the method comprising the following steps:

[0054] Step 1: Collect the daily load curve of the typical load and describe the n-dimensional daily load curve as a data point containing n data.

[0055] Overall architecture of the present invention Figure 1 The typical load curve and atypical load curve for the same region differ significantly, and the frequency of occurrence of typical loads is significantly greater than that of atypical loads. This is reflected in high-dimensional vector space as a significant difference in sample density. A convex cluster is one in which the line connecting each pair of points lies within the cluster. A non-convex cluster is one in which the line connecting each pair of points passes outside the cluster. Most clustering algorithms cannot effectively identify non-convex clusters. The DPC algorithm is a density-based clustering algorithm that can discover non-convex clusters and is sensitive to sample density. It can accurately cluster high-density clusters, making it suitable for load curve clustering.

[0056] The clustering performance of the four clustering algorithms was tested using the REFIT electrical load measurement dataset, which contains load data from 20 households over the past two years, with a sampling interval of 8 seconds. After data processing, 340 days of load data for the 20 households were obtained. The sampling interval was expanded to 30 minutes, resulting in 340 x 48 load records. The load data for the 20 households over the same period was summed to form a single large load record.

[0057] The REFIT dataset is a real-world dataset without cluster labels. To quantitatively analyze the clustering effect of the improved DPC algorithm on this real-world dataset, we first selected 200 daily load curves with relatively similar loads from the dataset as a test load dataset. Then, we randomly selected a certain percentage of load data from this dataset and artificially modified them to create significant differences from the original data. These data were then labeled with clusters. 20% of the load data were then increased at certain times and labeled "Cluster 2." 10% of the load data were decreased at certain times and labeled "Cluster 3." The remaining 70% of the unmodified load data were labeled "Cluster 1." This simulated load dataset with cluster labels was obtained. In the modified load curve data, the blue curve represents "Cluster 1," the red curve represents "Cluster 2," and the green curve represents "Cluster 3."

[0058] Step 2: Calculate the local density ρ and relative distance δ of all data points to determine the initial cluster center.

[0059] The DPC algorithm requires the following conditions to achieve clustering: (1) the local density of the data point serving as the cluster center is higher than that of the surrounding data points; and (2) the relative distances between the data points serving as cluster centers are relatively large. When applied to load curve clustering, an n-dimensional daily load curve is a single data point containing n data points. The dimension of a daily load curve refers to the number of sampling periods within the curve. For example, a daily load curve with a sampling interval of 1 hour has 24 sampling periods, resulting in a dimension of 24. This leads to two important variables in the DPC algorithm: local density and relative distance.

[0060] Where, the local density (cutoff kernel) of the x-th load curve is ρ x The definition is as follows:

[0061]

[0062] Where, d x,y Represents the Euclidean distance between the xth data point and the yth data point, d c is the cutoff distance, and x(m) is the judgment function. For larger data sets, the Gaussian kernel can be used to calculate the local density, which is defined as follows:

[0063]

[0064] Euclidean distance d x,y The calculation formula is:

[0065]

[0066] Where n is the data dimension, x i and y i are the i-th dimension data of x data points and y data points respectively. When applied to load curve clustering, x i and y i are the load values ​​of the ith period of the xth load curve and the yth load curve respectively. c The definition of is as follows:

[0067] d c =sort(d) round(Np%)

[0068] Where N is the total number of data points in the dataset, p is a percentage parameter, usually ranging from 1 to 10, sort(d) indicates that the Euclidean distances between all data points in the dataset are sorted in ascending order, and round(N p %) means rounding off Np%, that is, after sorting all Euclidean distances in ascending order, take the round(N p The distance value of % is the cutoff distance. The expression of χ(m) is as follows:

[0069]

[0070] The relative distance δ of the xth data point x The definition is as follows:

[0071]

[0072] When the local density of the xth data point is the largest, it may be the cluster center, and its relative distance is the maximum distance from other data points to the xth data point; when the local density of the xth data point is not the largest, its relative distance is the shortest distance from all data points with a local density greater than its to it.

[0073] After calculating the local density ρ and relative distance δ of all data points, a decision diagram is constructed with ρ as the horizontal axis and δ as the vertical axis. For example, Figure 2 shown.

[0074] The DPC algorithm requires manual selection of data points with larger local density and relative distance as cluster centers, and assigning the remaining data points to the cluster centers with the nearest local density greater than their own local density to obtain clustering results. Figure 2In the figure, the solid red line and dashed blue line represent two of the possible selection methods. Different selection methods will produce different numbers of clusters and clustering results. Therefore, the accuracy of the DPC algorithm's clustering results is heavily influenced by the user's subjective experience and is subject to uncertainty. Furthermore, in the DPC algorithm, once a data point is incorrectly assigned, all points with lower local density within its cutoff distance will also be assigned to the incorrect cluster.

[0075] Commonly used clustering evaluation indicators include accuracy (ACC), precision, recall, Rand Index, Adjusted Rand Index (ARI) and FMI (Fowlkes-Mallows Index). In order to compare the performance of different clustering algorithms, this paper introduces three clustering evaluation indicators: ACC, ARI and FMI. Among them, ACC reflects the correctness of sample clustering, ARI quantifies the relationship between clustering results and random allocation, and FMI comprehensively considers clustering accuracy and recall. Using these three clustering evaluation indicators, the clustering effect of the clustering algorithm can be evaluated more comprehensively.

[0076] The calculation formula for ACC is:

[0077]

[0078] In the formula, m is the number of clusters, r j To be correctly assigned to cluster C j where n is the total number of data points. ACC represents the ratio of the number of correctly identified data points to the total number of data points. ACC measures how accurately a clustering algorithm assigns data points to clusters. ACC only considers the accuracy of sample clustering and does not take into account the precision and recall of clustering.

[0079] The calculation formulas for ARI and FMI are:

[0080]

[0081] Where a is the number of data point pairs that belong to the same cluster in both the actual and experimental cases, b is the number of data point pairs that belong to the same cluster in the actual case but not in the experimental case, c is the number of data point pairs that do not belong to the same cluster in the actual case but do belong to the same cluster in the experimental case, and d is the number of data point pairs that do not belong to the same cluster in both the actual and experimental cases. ARI accounts for the effects of random assignment, and FMI is the geometric mean of clustering precision and recall, both of which are used to measure the similarity between clustering results and the true categories.

[0082] The maximum values ​​of the above three clustering evaluation indicators are all 1. The larger the values ​​of the three indicators, the better the clustering accuracy and clustering precision of the algorithm, and the stronger the clustering performance.

[0083] Through the analysis of the DPC algorithm principle, it can be seen that the traditional DPC algorithm has the problems of uncertainty in the subjective selection of cluster centers and the existence of allocation errors. To solve the above problems, an adaptive selection method for the initial cluster centers is proposed, and the cluster cross density and cluster boundary density parameters are introduced. Based on this, an improved strategy for the DPC algorithm is proposed.

[0084] The data points used as cluster centers generally have both large local density and large relative distance. Therefore, in order to avoid noise points with small local density and large relative distance and boundary points with large local density and small relative distance from being selected as cluster centers, two thresholds are defined to improve the accuracy of the initial cluster center selection.

[0085] Let DATA={x1,x2,...,x n} is the load curve data of a certain n days, the local density and relative distance of each load curve are calculated, and the initial cluster center is determined.

[0086]

[0087] Ω0=Ω ρ ∪Ω δ

[0088] Where, Ω ρ is the set of load curve data points whose local density is greater than the average local density, Ω δ is the set of load curve data points whose relative distance is greater than the average relative distance, x i is the data point in DATA, ρ i is the corresponding local density, δ i is the corresponding relative distance, is the average of the local densities of all data points, is the average of the relative distances of all data points, and Ω0 is the set of initial cluster centers.

[0089] Step 3: Define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results.

[0090] The initial cluster center can be clustered to obtain the core point set Ω and the initial cluster set In actual situations, the local density and relative distance of some non-core points are also greater than the average value. Therefore, the initial cluster center set contains not only the actual core points but also some non-core points. The initial clustering is not accurate and needs to be corrected.

[0091] Define two clusters C i and C j The cluster cross density is J i,j , which is defined as follows:

[0092] J i,j =max{ρ x |x∈Y i ∩Y j}

[0093] Where Y i It's Y i ∩Y j The xth data point and cluster C in i The distance of the data point is less than the cutoff distance d c A collection of data points, Y j It's Y i ∩Y j The xth data point and cluster C in j The distance of the data point is less than the cutoff distance d c A collection of data points.

[0094] Define cluster C i The boundary point set is B i , which is defined as follows:

[0095] B i ={x j |x j =sort(ρ j )[:m]}

[0096] m=ceil[length(C i )×20%]

[0097] In the formula, sort(ρ j ) indicates that C i The local density of each data point in the set is sorted in ascending order, [:m] means taking the first m elements of the set. ceil means rounding up, length(C i ) means C i The number of data points in the cluster C i The boundary point is C i Data points where the local density is in the lower 20% range.

[0098] Define two clusters C i and C j The cluster boundary density is B i,j , which is defined as follows:

[0099]

[0100] Where, is cluster C i The boundary point set B i The local density average of each data point in .

[0101] For the two clusters C in the initial clustering i and C j , if J is satisfied i,j >B i,j , then cluster C i and C j Merge into one cluster. The principle of the initial clustering improvement strategy is as follows:

[0102] If the data points that should belong to the same cluster are mistakenly divided into two clusters C i and C j In, such as Figure 3 (a) As shown in the left figure, the red circle is centered at a data point and has a cutoff distance d c The number of data points contained in a circle with a radius of C is the local density of the point. i and C j The cluster intersection data point set Y i ∩Y j That is Figure 3 (a) The set of red circle data points in the right figure, cluster cross density J i,j is the local density value of the data point with the largest local density. Obviously, the set of data points with lower local density in the two clusters is the boundary point set B. i and B j is the blue circle data point set in the right figure, B i and B j The average local density is the cluster boundary density B i,j Obviously less than the cluster cross density J i,j , so we can judge that C i and C j In fact, they should belong to the same cluster.

[0103] If two clusters C i and C j is correctly divided, such as Figure 3 As shown in (b), the cluster intersection data point set Y of these two clusters is i ∩Y j It should contain only a few data points at the edge of the two clusters, or an empty set, then the local density of the most edge data points is J i,j Should be smaller than B i,j , you can judge C i and C j do not belong to the same cluster.

[0104] The specific steps to improve the strategy are as follows:

[0105] Step 1: Calculate the cluster set according to the above formula The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p , where p is The number of clusters included.

[0106] Step 2: Take J in descending order of local density of cluster centers p×p and B p×p For example, if the 4th and 7th cluster centers in Ω are the two data points with the highest local density in Ω, then first compare J 4,7 and B 4,7 Compare, if it does not meet J 4,7 >B 4,7 , then take the parameters of the first and third points of local density for comparison, and so on, until J is satisfied i,j >B i,j Or all clusters are compared pairwise.

[0107] Step 3: If J is satisfied i,j >B i,j , delete the original two cluster centers from Ω, and reselect the data points with the largest local density in the two clusters as cluster centers and add them to Ω, Merge the two clusters in . Return to step 1, calculate the new cluster cross density matrix and cluster boundary density matrix, and continue the cycle. If J is still not satisfied in the end i,j >B i,j , that is, the existing The clusters in can no longer be merged, so the loop stops and the clustering results are output.

[0108] After improving the DPC algorithm, adaptive selection of initial cluster centers and correction of initial clusters were achieved.

[0109] The algorithm complexity of the improved DPC algorithm mainly consists of the following five parts:

[0110] Calculate the Euclidean distance between each data point, the complexity is O(n 2 );

[0111] Calculate the local density of each data point, the complexity is O(n 2 );

[0112] Calculate the relative distance of each data point, the complexity is O(n 2 );

[0113] The initial cluster center is selected by calculation, with a complexity of O(n);

[0114] Assign all data points to different clusters, the complexity is O(n);

[0115] The complexity of calculating cluster cross density, cluster boundary density and merging clusters is O(n 2 );

[0116] The algorithm complexity of the improved DPC algorithm is O(n 2 ), which is the same as the traditional DPC algorithm. The improvement of the traditional DPC algorithm does not increase the complexity of the algorithm. At the same time, three other common clustering algorithms are selected to compare with the improved DPC algorithm. Among them: the complexity of the K-means algorithm is O(n), the complexity of the DPC algorithm is O(n 2 ), the complexity of the DBSCAN algorithm is O(n 2 ). Compared with common clustering algorithms, the complexity of the improved DPC algorithm is at an average level.

[0117] All algorithms in this paper were implemented using MATLAB. Six two-dimensional datasets, four multidimensional datasets, and one real-world dataset were used to test and compare the improved DPC algorithm with three common advanced algorithms: DPC, K-means, and DBSCAN. The improved DPC algorithm was used to cluster the real-world datasets and effectively extract their typical load curves.

[0118] The data set parameters are shown in Table 1:

[0119] Table 1 Dataset parameters

[0120]

[0121] The algorithm parameter settings are shown in Table 2:

[0122] Table 2 Algorithm parameter settings

[0123]

[0124]

[0125] For different data sets, the parameter settings of each algorithm are shown in Table 2. The parameter setting standard of the algorithm is to make the clustering effect of the algorithm as optimal as possible.

[0126] Among them, the clustering effect diagrams of the six two-dimensional data sets are as follows: Figure 5 Each point in the figure represents a two-dimensional data point, and the horizontal and vertical coordinate values ​​represent the position coordinates of the point in the two-dimensional coordinate system, which has no physical meaning.

[0127] Depend on Figure 5It can be seen that the improved DPC algorithm and DBSCAN algorithm are more accurate in clustering the Moons dataset; the DPC algorithm has some clustering errors, the reason being that the algorithm has a joint allocation error when allocating non-core data points; the reason why the K-means algorithm has errors is that the algorithm cannot effectively identify non-convex clusters.

[0128] Depend on Figure 6 It can be seen that the first three density-based algorithms can accurately cluster the Spiral dataset, while the K-Means algorithm cannot identify non-convex clusters, resulting in poor clustering results.

[0129] Depend on Figure 7 It can be seen that the improved DPC and traditional DPC algorithms have better clustering effects on the Flame dataset, the DBSCAN algorithm mistakenly identifies some data points as noise, and the K-Means algorithm has the worst clustering effect.

[0130] Depend on Figure 8 It can be seen that when allocating non-core points, the improved DPC algorithm mistakenly allocates some data points from another cluster that is relatively closer to the current cluster, but the clustering effect is still better than the other three algorithms.

[0131] Depend on Figure 9 It can be seen that the improved DPC algorithm has a better clustering effect on the Aggregation dataset; the traditional DPC algorithm only clusters 6 clusters because only 6 cluster centers with large local density and relative distance are selected in the decision graph, and the clustering effect is poor due to subjective uncertainty; the clustering effect of the DBSCAN algorithm is relatively worse because the parameter setting of the DBSCAN algorithm is difficult and it is difficult to find parameters suitable for this dataset; the reason why the K-means algorithm has a poor clustering effect is that it cannot identify non-convex clusters.

[0132] Depend on Figure 10 It can be seen that the improved DPC algorithm and the traditional DPC algorithm have better clustering effects; the DBSCAN algorithm mistakenly identifies some data points as noise.

[0133] Table 3 shows a comparison of the clustering performance of the four methods for different datasets. The improved DPC algorithm performs better than the other three clustering algorithms for all ten datasets. For two-dimensional datasets, the improved DPC algorithm maintains a high accuracy rate. For the Jain dataset, the improved DPC algorithm's ARI value does not exceed 0.9 because the density of the two classes of elements in certain areas of the Jain dataset is too close, making it difficult for the algorithm to effectively classify the different classes. For high-dimensional datasets, the improved DPC algorithm significantly outperforms the other three algorithms.

[0134] Table 3 Comparison of clustering effects

[0135]

[0136]

[0137] The modified load data curve is as follows: Figure 11 , we select 20% of the load data and increase its load at certain times and label it "Cluster 2." We select 10% of the load data and reduce its load at certain times and label it "Cluster 3." We label the remaining 70% of the load data "Cluster 1." This simulates an actual load data set with cluster labels. The blue curve represents the "Cluster 1" load, the red curve represents the "Cluster 2" load, and the green curve represents the "Cluster 3" load.

[0138] The algorithm parameter settings are shown in Table 4:

[0139] Table 4 Algorithm parameter settings

[0140]

[0141] Four algorithms are used to cluster the modified load curves. The parameter settings of the four algorithms are shown in Table 4. The results are shown in Table 4. Figure 12 As shown in Figure 2, the red, blue, and green curves are respectively clustered into the same class, and the black curve is the curve identified as noise by the DBSCAN algorithm.

[0142] It can be seen that the improved DPC algorithm has the best clustering effect, and its clustering results are very close to the labels; the DBSCAN algorithm can cluster the "Cluster 2" curves more accurately, but it mistakenly assigns some "Cluster 3" curves to "Cluster 1" and mistakenly identifies some curves as noise; similarly, the K-means algorithm can also cluster the "Cluster 2" curves well, but it mixes the "Cluster 1" curves with the "Cluster 3" curves; the DPC algorithm mistakenly assigns all "Cluster 3" curves to "Cluster 1", mainly because the subjective selection of cluster centers in the decision diagram is very limited. The cluster centers selected in the decision diagram by the ordinary DPC algorithm and the improved DPC algorithm in this clustering are as follows: Figure 13 As shown, in actual situations, the improved DPC algorithm does not produce a decision diagram. A decision diagram is drawn here for further explanation.

[0143] Because the decision graph contains two points with large local density ρ and relative distance δ, the DPC algorithm typically selects only these two points as cluster centers when the actual number of clusters in the dataset is unknown, causing a dataset containing three categories to be incorrectly clustered into two. The improved DPC algorithm does not require the subjective selection of cluster centers on the decision graph, thus avoiding the uncertainty of subjective selection.

[0144] The clustering results of the four algorithms are shown in Table 5:

[0145] Table 5 Clustering results of 4 algorithms

[0146]

[0147] It can be seen that for the REFIT electrical load measurement dataset, the ACC, ARI, and FMI clustering evaluation index values ​​of the improved DPC algorithm are all the largest, especially the ARI value is much higher than that of other algorithms. That is, the clustering results of the improved DPC algorithm are most similar to the label values, and its clustering effect is better than the other three algorithms. Since the K-means algorithm cannot identify non-convex clusters and performs poorly on high-dimensional data sets, it simply identifies data points with relatively close Euclidean distances as the same class without considering the specific layout of each data point in the high-dimensional space. Therefore, its clustering effect is the worst, and its ACC, ARI and FMI are the lowest among the four algorithms; the DBSCAN algorithm can identify non-convex clusters, but its parameter setting is difficult, and some normal curves are mistakenly identified as abnormal curves, resulting in poor clustering effect; the traditional DPC algorithm has subjective uncertainty and thus only divides the data set into two categories, resulting in poor clustering effect. Compared with the DBSCAN algorithm, the traditional DPC algorithm has higher ACC and FMI, but lower ARI. The reason is that the DBSCAN algorithm divides the data set into three categories, and the clustering results are more similar to the labels than the traditional DPC; the improved DPC algorithm adaptively selects the correct clustering center, and its clustering effect is the best.

[0148] In summary, the proposed improved DPC algorithm is superior to other algorithms in clustering high-dimensional data, and its biggest advantage over the traditional DPC algorithm is that it can automatically select cluster centers, effectively avoiding the uncertainty of subjective selection of cluster centers.

[0149] Four algorithms are used to cluster the REFIT electrical load measurement data set. The parameter settings and clustering results of the four algorithms are shown in Tables 5 and 6. The maximum cluster set obtained by each algorithm is the typical load curve set. For the improved DPC algorithm and the traditional DPC algorithm, the cluster center of the largest cluster is taken as the typical load curve; for the DBSCAN algorithm, the data point with the highest density in the largest cluster is taken as the typical load curve; for the K-means algorithm, the data point closest to the centroid of the largest cluster is taken as the typical load curve. The typical load curve sets and typical load curves obtained by the four algorithms are shown in Tables 5 and 6. Figure 14-21 shown.

[0150] Table 6 Algorithm parameters and clustering results

[0151]

[0152] As can be seen, the number of typical load curves in the improved DPC and DBSCAN algorithms is relatively reasonable, accounting for about 80% of the total number of curves. This number of curves can represent the load's power usage during most of the time. However, the DPC algorithm classifies over 98% of the curves into the typical load curve set, which is too high. The K-means algorithm only classifies 40% of the curves as typical load curves, which is too low. Neither algorithm can effectively represent typical power usage.

[0153] Comparing the typical load curve sets clustered by the four algorithms shows that the typical load curve set clustered by the improved DPC algorithm contains few load curves containing extreme loads, while the typical load curve sets clustered by the traditional DPC algorithm and the DBSCAN algorithm contain a large number of prominent extreme load curves, resulting in poor clustering effect. The typical load curve set clustered by the K-means algorithm contains only 137 curves, and many curves that should belong to the typical load curve set are assigned to other sets, resulting in missing features of the typical load curve set and inaccurate typical load curves. In summary, the improved DPC algorithm proposed in this invention extracts more accurate and representative typical load curves than the other clustering algorithms.

[0154] The present invention studies a typical load curve extraction method based on an improved DPC algorithm. First, the principle of the DPC algorithm and its existing problems are analyzed, including the subjective uncertainty of selecting cluster centers and the associated errors in the allocation of data points. In response to the above problems, a method for adaptively determining the initial cluster centers is proposed. This method can automatically select all data points that may be correct cluster centers as initial cluster centers based on the local density and relative distance characteristics of data points in the data set, and then cluster them using the DPC algorithm according to the initial cluster centers to obtain a set of initial clusters. Secondly, an initial cluster center improvement strategy is proposed. The strategy defines two parameters, cluster cross density and cluster boundary density, and corrects the initial clusters by comparing the sizes of the two parameters. Then, the algorithm complexity of the improved DPC algorithm is analyzed, and it is concluded that the complexity of the proposed improved DPC algorithm is the same as that of the traditional DPC algorithm. Among the common clustering algorithms, the algorithm complexity of the improved DPC algorithm is at the average level. Finally, the improved DPC algorithm and other common clustering algorithms were used to cluster data sets of various dimensions and actual data sets, and the results were compared and verified. The improvement effect was as follows: the proposed improved DPC algorithm can adaptively determine the initial cluster center and determine the final cluster center through the initial cluster improvement strategy. It is significantly better than the traditional DPC algorithm in the case where the cluster center cannot be effectively determined manually. In the verification of data sets of various dimensions, the clustering evaluation index values ​​of the proposed improved DPC algorithm are better than those of other algorithms. Among them, ACC, ARI and FMI are 25.40%, 46.92% and 21.83% higher than those of other algorithms, respectively, verifying the excellent performance and effect of the improved DPC algorithm. Using the improved DPC algorithm to cluster actual data sets can effectively extract typical load curves from the actual load curve data set, and the extracted typical load curves are more representative than those of other algorithms.

[0155] The second aspect of the present invention relates to a typical load curve extraction system based on an improved DPC algorithm using the method described in the first aspect of the present invention; the system is implemented using the method described in the first aspect of the present invention; the system includes an acquisition module, a calculation module and an output module; the acquisition module is used to acquire the daily load curve of the typical load and describe the n-dimensional daily load curve as a data point containing n data; the calculation module is used to calculate the local density ρ and relative distance δ of all data points and determine the initial cluster center; the output module is used to define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results.

[0156] A third aspect of the present invention relates to a terminal, comprising a processor and a storage medium; the storage medium is used to store instructions; and the processor is used to operate according to the instructions to execute the steps of the method described in the first aspect of the present invention.

[0157] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect of the present invention.

[0158] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art will appreciate that the technical solutions of the present invention still include modifications or equivalent substitutions that may be made to the specific embodiments of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention are intended to be covered by the claims of the present invention.

Claims

1. A typical load curve extraction method based on an improved DPC algorithm, characterized in that: The method comprises the following steps: Collect the daily load curve of typical load and describe the n-dimensional daily load curve as a data point containing n data; Calculate the local density ρ and relative distance δ of all data points to determine the initial cluster center; Define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results.

2. The typical load curve extraction method based on the improved DPC algorithm according to claim 1 is characterized in that: The daily load curve of the typical load is collected, and the n-dimensional daily load curve is described as a data point containing n data, including: Collect 340 days of load data from 20 household users with a sampling interval of 8 seconds; The sampling interval is extended to 30 minutes, and the loads of 20 household users at the same time are added together to form the large load data; 200 daily load curves with relatively similar loads in the data set are selected as the load data set for testing.

3. The typical load curve extraction method based on the improved DPC algorithm according to claim 2 is characterized in that: The 200 daily load curves with relatively similar loads in the selected data set are used as the load data set for testing, including: 20% of the load data are selected to increase the load at certain moments and are labeled "Cluster 2". 10% of the load data are selected to reduce the load at certain moments and are labeled "Cluster 3". The remaining 70% of the load data that has not been modified are labeled "Cluster 1". This simulates the actual load data set with cluster labels.

4. The typical load curve extraction method based on the improved DPC algorithm according to claim 3 is characterized in that: The calculation of the local density ρ and relative distance δ of all data points to determine the initial cluster center includes: The initial cluster centers are: Ω0=Ω ρ ∪Ω δ Where, is the average value of the local density of all data points, is the average value of the relative distances of all data points, x i is the i-th data point in the load data set.

5. The typical load curve extraction method based on the improved DPC algorithm according to claim 4 is characterized in that: The definition of two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results, including: Calculate cluster sets The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p , p is the number of clusters; Take J in descending order of local density of cluster centers p×p and B p×p If the elements in J i,j >B i,j , then delete the original two cluster centers, reselect the data point with the largest local density in the two clusters as the cluster center, and cluster C i and C j Merge into one cluster; Calculate the new cluster cross density matrix and cluster boundary density matrix, and continue looping until the clusters can no longer be merged. Stop the loop and output the clustering results.

6. The typical load curve extraction method based on the improved DPC algorithm according to claim 5 is characterized in that: The calculated clustering set The cluster cross density and cluster boundary density between clusters in the cluster are used to form the cluster cross density matrix J p×p and the cluster boundary density matrix B p×p , include: Define two clusters C i and C j The cluster cross density is J i,j for J i,j =max{ρ x |x∈Y i ∩Y j } Where Y i It's Y i ∩Y j The xth data point and cluster C in i The distance of the data point is less than the cutoff distance d c A collection of data points, Y j It's Y i ∩Y j The xth data point and cluster C in j The distance of the data point is less than the cutoff distance d c A collection of data points, Y i ∩Y j For two clusters C i and C j The data points at the critical position.

7. The typical load curve extraction method based on the improved DPC algorithm according to claim 6 is characterized in that: The cluster cross density and cluster boundary density between clusters in the cluster set C are calculated to form a cluster cross density matrix J p×p and the cluster boundary density matrix B p×p , include: Define cluster C i The boundary point set is B i , B i ={x j |x j =sort(ρ j )[:m]} m=ceil[length(C i )×20%] In the formula, sort(ρ j ) indicates that C i The local density of each data point in the set is sorted in ascending order, [:m] means taking the first m elements of the set, ceil means rounding up, and length(Ci) means C i The number of data points in the cluster C i The boundary point is C i Data points where the local density is in the lower 20% range; Define two clusters C i and C j Cluster boundary density B i,j for Where, is cluster C i The boundary point set B i The local density average of each data point in .

8. A typical load curve extraction system based on an improved DPC algorithm, characterized by: The system is implemented using the method according to any one of claims 1 to 7; The system includes an acquisition module, a calculation module and an output module; The acquisition module is used to acquire a daily load curve of a typical load and describe the n-dimensional daily load curve as a data point containing n data; The calculation module is used to calculate the local density ρ and relative distance δ of all data points and determine the initial cluster center; The output module is used to define two clusters C i and C j The cluster cross density J i,j , cluster boundary density B i,j , to correct the initial clustering and output the clustering results.

9. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.