Power consumer industry classification analysis method based on full-automatic density clustering algorithm

Through the power user classification method based on the fully automatic density clustering algorithm, the problem of inaccurate classification of power users in the existing technology is solved, accurate analysis of the power user industry and identification of changes in power usage properties is realized, and the management efficiency and decision-making support capabilities of the power system are improved.

CN120217025APending Publication Date: 2025-06-27GUANGDONG POWER GRID CO LTD DONGGUAN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510289082.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing power user classification methods cannot effectively deal with the diverse and changing power user groups, especially when the nature of users' electricity consumption changes, it is difficult to accurately reflect the type of industry and specific situations of the users' industry, resulting in low decision-making accuracy and efficiency in power management, load forecasting, demand response and other work.

Method used

The industry classification method for power users belonging to is adopted based on the fully automatic density clustering algorithm. By collecting real-time data of power users, normalizing and smoothing processing is performed, and the data is pre-processed in combination with the timing feature layering method. Then, based on the density clustering algorithm, a classification model of the industry to which the power user belongs is established, dynamically adjust the Eps and MinPts parameters, identify high-density areas and allocate cluster tags, and ultimately realize the identification of industry classification and changes in electricity consumption properties.

Benefits of technology

It realizes accurate analysis of the power user industry, can identify the load characteristics of power users at different stages, improves the accuracy and adaptability of industry classification, and enhances the management efficiency and decision-making support capabilities of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217025A_ABST
    Figure CN120217025A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of electric power, and provides a power consumer industry classification analysis method based on a full-automatic density clustering algorithm, which comprises the following steps: firstly, collecting power consumer data, and carrying out segmentation processing and feature mapping on a time sequence of power load; combining the polynomial parameters with the statistical features to generate low-dimensional vectors; carrying out clustering analysis on the mapped spatial features by adopting a density-based clustering method; and finally, analyzing a clustering result by applying a Rand index, homogeneity, integrity and a harmonic average. According to the invention, through multi-level processing of the power consumer load data and combination of time sequence feature extraction and a density clustering algorithm, accurate analysis of the power consumer industry is realized, and the accuracy and adaptability of power consumer industry classification are improved. Power supply enterprises can be powerfully helped to quickly discover power consumers with unknown industry classification and changed industry classification, and the management efficiency and decision support capability of a power system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric power, and particularly relates to a method for classifying the industries to which power users belong based on a fully automatic density clustering algorithm. Background Art

[0002] With the rapid development of the power system, especially driven by emerging industries and new energy, the diversity and complexity of power users are increasing day by day. Existing marketing management systems have problems such as inconsistent input of power user information and decentralized management, and there are also situations such as missing filling, incorrect filling, and failure to go to the power supply bureau in time after the change of the power consumption nature of users, resulting in the lack of industry identification, incorrect industry identification, or untimely update of power user data information, and it is impossible to efficiently carry out the industry analysis work of power users. Traditional power user classification methods mainly classify according to on-site conditions such as power consumption load and region. However, this method has certain limitations, especially in the case where the power consumption nature at the customer site changes and is not updated in time. Facing a large-scale, diverse, and highly variable power user group, these traditional methods cannot fully consider the multi-dimensional index time series characteristics of users, and it is difficult to accurately reflect the industry type and specific situation to which users belong, resulting in low decision-making accuracy and efficiency in power management, load forecasting, demand response, and other work. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a method for classifying the industries to which power users belong based on a fully automatic density clustering algorithm to solve the problems in the prior art. The technical solution adopted by the present invention is as follows:

[0004] A method for classifying the industries to which power users belong based on a fully automatic density clustering algorithm includes the following steps:

[0005] Step 1: Collect real-time data of power users, use multiple indicators such as maximum power, peak-to-total ratio, flat-to-valley ratio, valley-to-total ratio, and daily average load as clustering analysis indicators, and perform normalization processing:

[0006]

[0007] where x i is the initial data, x max is the maximum daily load, and x' i is the value after normalization;

[0008] Perform smoothing processing to generate a smoothed load sequence:

[0009]

[0010] Step 2: Preprocess the data based on the time series feature stratification method;

[0011] Step 3: Establish a classification model for the industries to which power users belong based on the fully automatic density method;

[0012] Step 4: Industry mapping and matching correction;

[0013] Step 5: Cluster performance analysis.

[0014] Furthermore, Step 2 includes:

[0015] Step 2.1, segment the time series of power load:

[0016] Segment the time series X = {x1, x2, … x i} of length I, divide the time series into m segments, and find m - 1 cut points

[0017] Introduce the time series data into the dynamically growing window point by point in an iterative manner, and calculate the sum of squared errors SSE within the window q , the maximum standard error is SEP max , when SSE q > SEP max , the current point is marked as the cut point t i , then end the current segment. Each time an error exceeding the threshold is detected, record the cut point t i . The initial segment starts from t0 = 1. When t1 = n is detected, it means that the first n data points belong to the first segment. Use the cut point set {t1, t2, … t m-1} to define the segment range: segment q: [t q-1 + 1, t q ;

[0018] Step 2.2, perform segment mapping and grouping on the segmented segments:

[0019] Map each segment to an array including polynomial fitting parameters and statistical features. Each segment is projected into an l - dimensional space, where l is the length of the mapped segment. The coefficients p of the polynomial are directly obtained from the least - squares fitting process of the segmentation q , and at the same time calculate the following statistical features:

[0020] Variance

[0021] Skewness γ 1q :

[0022] Autocorrelation coefficient AC q :

[0023] By combining the polynomial fitting parameters and statistical features:

[0024]

[0025] where p q is the polynomial approximation parameter for approximately segment q;

[0026] The segments are divided into two categories, and the number of groups k = 2;

[0027] Step 2.3, perform time series mapping on the grouped segments:

[0028] For each time series Xi, extract the corresponding cluster center where:

[0029] i ∈ {1, …, T} represents the time series index, and T is the number of time series;

[0030] j ∈ {1, …, k} represents the index of the cluster in each time series, and k is the number of clusters;

[0031] For each cluster, C ij Extract the following: the cluster center the average value of all segments in the cluster; the extreme segment the segment with the largest variance in the cluster;

[0032] Map the length of the cluster to w = l × 2, where l is the length of the mapped segment, for each cluster Map it as:

[0033]

[0034] Select a time series as a reference, calculate the cluster centers of all other time series, and match them with the cluster center of the reference time series, and select the cluster closest to the reference center for correspondence;

[0035] When the matching is completed, each time series will be converted into a mapped time series X i ′, and this time series includes the extracted cluster features. The length of the mapped time series is: (w × k) + v, where w × k is the total feature dimension of all clusters, and v = 2 is two additional information, including the error difference and the number of segments.

[0036] Furthermore, step 3 includes:

[0037] Step 3.1, calculate the radius Eps of the circular area around the specified point:

[0038] For each data point P(x, y), calculate the distance between this point and its nearest neighbor. Given the coordinates (x, y) of the data point and the coordinates (x c , yc ), the distance D between two nearest neighbors min (P(x, y)) is calculated by the following formula:

[0039]

[0040] For each point in the mapped time series X i ′, apply formula (3-1) to calculate the shortest distance from each point to its nearest neighbor, store all the shortest distances in a list ListD, and sort them in ascending order:

[0041] List D = Sorted(List D )(3-2)

[0042] Apply a statistical histogram to display the data distribution of the extracted distances, and calculate the normalized distance list through formula (3-3):

[0043]

[0044] In the formula, R is the sub-range value where the normalized distances become similar, and it is calculated by the following formula:

[0045] R = int(N × C)(3-4)

[0046] In the formula, N is the number of sampling points in the data set, and C is the division factor;

[0047] Divide the number of non-zero distances in the normalized histogram by the distance range R Av , thereby obtaining the frequency value F, which represents the degree of noise and long distances in the data set:

[0048]

[0049]

[0050] In the formula, Hstg A represents the number of frequency points that are non-zero within the histogram range, and R Avg is the average value of the data points within each histogram range;

[0051] By calculating the distance range between the maximum value P max and the minimum value P min in the data points, and then adjusting the radius Eps according to the characteristics of noise and outliers, use the following formula to calculate the final radius Eps value:

[0052] Eps = P min + ((P min - P max ) × F)(3-7)

[0053] Step 3.2, calculate the minimum number of points Min according to the final radius Eps value Pts :

[0054] Scan all points in the dataset using the radius Eps and calculate the number of points within the radius Eps:

[0055]

[0056] where D Eps (p(x,y)) represents the number of points within the radius Eps around the given point P(x,y). Repeat this process for each point in the dataset and finally store the results in a list:

[0057]

[0058] Calculate the data graph of List Eps and determine the most frequent value;

[0059] Step 3.3, assign cluster labels to the sample points in the dataset:

[0060] For each point P(x,y) in the dataset, construct a circular region with a radius of Eps and calculate the number of points within this region:

[0061]

[0062] If the number of points within this region is greater than or equal to Min Pnts , then consider this point as a clustering core and label all points within its neighborhood as the same cluster;

[0063] If there are no already labeled data points within the neighborhood, assign the next value of the maximum label L max previously used to the current point:

[0064] L[P] = L max +1(3 - 12)

[0065] If at least one adjacent data point has already been assigned a label, assign the current point to the label of the point with the smallest label in the neighborhood:

[0066] L[P] = min(L)(3 - 13)

[0067] where L is the label;

[0068] If the number of points N within the radius Eps region is less than the value of Min Pnts , then consider the density of this region to be too low to form a cluster; repeat the step of assigning cluster labels for all points in the dataset to identify high-density regions as potential cluster cores;

[0069] Process all the unclustered points P in the dataset, check whether there are labeled points within the neighborhood radius Eps. If there are, assign the current point to the nearest labeled point in the neighborhood, and iterate this process until all points are labeled as members of a cluster;

[0070] Merge clusters: Use the Union-Find method of disjoint-set data structure to identify clusters that are connected to each other and merge them into one cluster. Reassign labels to the merged clusters, and assign the smallest label value to the new merged cluster to obtain the final clustering result.

[0071] Furthermore, step 4 includes:

[0072] Step 4.1, Industry mapping and matching correction:

[0073] Based on the clustering result of step 3, assign corresponding industry identifiers to each cluster label, and compare with the registered industry information of the user to ensure the accuracy of classification;

[0074] Match according to the characteristics of the assigned cluster label L and industry characteristics, select the closest industry classification, and assign the industry identifier Z to the cluster label j :

[0075]

[0076] where Z j is the industry identifier of the cluster label L i ; Z k is the industry feature template; Distane is the distance function for feature matching;

[0077] Traverse all users, extract the original registered industry information Z u , and compare it with the assigned industry identifier Z j ;

[0078] If Z u = Z j , it is considered that the matching is successful, and record the industry classification of the user;

[0079] If Z u ≠ Z j , it indicates that the electricity consumption nature of the user has changed. Record the changed industry classification of the user and mark the user as "pending manual verification" status;

[0080] Step 4.2, Classification result statistics and user information registration:

[0081] Count the number of users in each industry category and their main load characteristics, and generate a classification statistical table, including industry category, number of users, and average load characteristics. Then, use bar charts and heat maps to display the classification results;

[0082] Record all Z u ≠Z j user information, including user ID, original industry category, current load characteristics, and recommended industry category, and send a notice to the management staff to verify the industry classification of this electricity user.

[0083] Furthermore, step 5 includes:

[0084] Step 5.1, Rand Index:

[0085]

[0086] where a: the number of identical point samples found in the same group of A and the same group of B;

[0087] b: the number of identical point samples found in different groups of A and different groups of B;

[0088] c: the number of identical point samples found in the same set of A and different sets of B;

[0089] d: the number of identical point samples found in different groups of A and the same group of B;

[0090] In the formula, A is the true value of the label, and B is the clustering result;

[0091] Step 5.2, Homogeneity:

[0092]

[0093] In the formula, H(C|K) represents the distribution degree of points belonging to different categories in each cluster, and H(C) represents the entropy of category C;

[0094]

[0095] In the formula, n c,k is the number of samples belonging to category c and belonging to cluster k, n k is the total number of samples in cluster k, and N is the total number of samples in the dataset;

[0096] Step 5.3, Completeness:

[0097]

[0098] In the formula, H(K|C) represents the distribution of samples assigned to different clusters in each category, and H(K) represents the entropy of cluster K;

[0099]

[0100] Wherein, n c,k is the number of samples belonging to category c and cluster k, and n c is the total number of samples in category c, and N is the total number of samples in the dataset;

[0101] Step 5.4, harmonic mean:

[0102] The harmonic mean is the weighted average of homogeneity and completeness, and is calculated using the normalized mutual information:

[0103]

[0104] Wherein, H represents homogeneity and C represents completeness;

[0105] Step 5.5, calculate the Rand index, homogeneity, completeness, and harmonic mean based on the original data and features provided in Steps 1 and 2, and summarize the calculation results of each index; verify the rationality of the clustering result obtained in Step 3 and the effect of the industry classification in Step 4 through performance evaluation.

[0106] The present invention has the following beneficial effects:

[0107] (1) The present invention performs segmented processing and feature mapping on time series data, reducing the computational overhead and enabling the method to efficiently process large-scale user data. Then, by combining polynomial fitting parameters with statistical features, low-dimensional vectors are generated, reducing the computational complexity while retaining the important information of the data. And by dynamically adjusting Eps and MinPts, the method can adapt to data distributions with complex shapes and different densities, avoiding misclassification problems caused by improper parameter settings in traditional clustering methods. Finally, high-density regions are identified by detecting core points in the data, and cluster labels are assigned to the sample points in the dataset;

[0108] (2) The present invention realizes the accurate analysis of the power user industry through multi-level processing of power user load data, combined with the extraction of time series features and density clustering algorithms. This method can effectively identify the load characteristics of power users at different times and perform clustering based on these characteristics, which can effectively help power supply enterprises quickly discover power users with unclear industry classification or changed industry classification, improve the accuracy and adaptability of power user industry analysis, and thus provide support for load forecasting, demand response, etc. of the power system, and further improve the management efficiency and decision-making support ability of the power system. Description of the Drawings

[0109] Figure 1 is the method flow chart of the present invention. Detailed Embodiments

[0110] The following will combine with the Figure 1 in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. If not specifically specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0111] The purpose of this application of the present invention is to provide a method for classifying and analyzing the power user industry based on the fully automatic density clustering algorithm: First, collect power user data, and preprocess the data based on the time series feature stratification method, segment the time series of power load and perform feature mapping. Then combine polynomial parameters with statistical features to generate low-dimensional vectors. Then use the density-based clustering method to perform clustering analysis on the mapped spatial features, so as to classify the time series into different groups. And assign industry identifiers to the clustering results to identify users with changed electricity consumption properties. Finally, use four clustering performance indicators, namely the Rand index, homogeneity, completeness, and harmonic mean, to analyze the clustering results. The specific steps are as follows:

[0112] Step 1: Power user data collection:

[0113] Collect real-time power user data and clean these data. The present invention selects the daily load curve of users as the main data, and at the same time selects some influencing factors that have a greater impact on load fluctuations: five load fluctuation characteristics, namely the maximum power, peak-to-total ratio, flat-to-valley ratio, valley-to-total ratio, and daily average load, as clustering analysis indicators.

[0114] In the power consumption data of users, due to power outages, equipment failures, and unexpected situations during transmission, it is easy to have data loss or abnormal data collection. For missing values, mean filling is used, and this processing reduces the impact of missing data on the data set to a certain extent.

[0115] To reduce the differentiation between the value ranges of different features of the data, normalize the data supplemented in the previous step.

[0116]

[0117] In the formula, x i is the initial data, x max is the maximum daily load, and x' i is the value after normalization.

[0118] During the acquisition and transmission of power load data, it is subject to noise interference, resulting in data fluctuations and instability. Therefore, in order to reduce the impact of noise, the data needs to be smoothed. Specifically, through Equation (1-2), the continuous n data values around each load data point are averaged to generate a smoothed load sequence.

[0119]

[0120] Step 2: Preprocess the data based on the time series feature stratification method:

[0121] Step 2.1, segment the time series of the power load:

[0122] Segment the time series X = {x1, x2, …, x i} with length I, divide the time series into m segments, and determine m - 1 cut points The set of segments Q = {q1, q2, …, q m} includes

[0123] Adopt the dynamic growth window method, iteratively introduce the data points of the time series into the growth window, and simultaneously update the corresponding least squares polynomial approximation of the segment and its error. The window continuously expands to find the cut point t such that the approximation error of each segment q k is minimized, while ensuring that the error does not exceed the set threshold SEP max . Repeat this process until the end of the time series is reached;

[0124] Iteratively introduce the time series data point by point into the dynamic growth window, calculate the sum of squared errors SSE q within the window, and artificially define the maximum standard error as SEP max . When SSE q > SEP max , the current point is marked as the cut point t i , then end the current segment. Each time an error exceeding the threshold is detected, record the cut point t i . The initial segment starts from t0 = 1. When t1 = n is detected, it means that the first n data points belong to the first segment. Use the set of cut points {t1, t2, …, t m-1} to define the segment range: segment q: [t q-1 + 1, t q ;

[0125] Step 2.2, perform segment mapping and grouping on the segmented segments:

[0126] After the time series is segmented, each segment is mapped to an array that includes polynomial fitting parameters and statistical features. Each segment is projected into an l-dimensional space, where l is the length of the mapped segment. The purpose of this step is to represent the original time series segment with low-dimensional features, thus facilitating subsequent classification tasks.

[0127] The coefficients p of the polynomial are directly obtained from the least squares fitting process of the segmentation. q 。

[0128] Meanwhile, the following statistical features are calculated:

[0129] (1) Variance Measures the amplitude of change of the segment values:

[0130]

[0131] (2) Skewness γ 1q : Measures the symmetry of the segment value distribution relative to the mean:

[0132]

[0133] (3) Autocorrelation coefficient AC q : Describes the correlation between adjacent points in the segment:

[0134]

[0135] By combining the polynomial fitting parameters and statistical features, each segment is mapped to an l-dimensional vector l = c + f, where c is the order of the polynomial and f is the number of statistical features.

[0136]

[0137] where p q is the polynomial approximation parameter for approximating segment q, and this process can simplify the original length of the section from t q -t q-1 +l to the low-dimensional vector c + f.

[0138] The above mapping step successfully transforms the time series data into a structured feature space. Next, all segments are grouped by the hierarchical classification method, and the segments in the time series are represented as feature vectors with the same length, thus significantly reducing the dimension of the representation and the computational complexity.

[0139] The ward distance is used as the similarity metric to measure the similarity between segments, ensuring that each merging operation can minimize the variance within the cluster and improve the grouping quality. For all datasets and time series, we uniformly choose to divide the segments into two categories, that is, the number of groups k = 2.

[0140] Step 2.3, perform time series mapping on the grouped segments:

[0141] In the time series mapping stage, each time series is converted into a set of clustered segments, and all time series are represented in a unified dimensional space.

[0142] For each time series X i , extract the corresponding cluster center from the aforementioned segmentation process Where:

[0143] i ∈ {1, …, T} represents the time series index, and T is the number of time series;

[0144] j ∈ {1, …, k} represents the index of the cluster in each time series, and k is the number of clusters.

[0145] For each cluster, C ij Extract the following:

[0146] 1) The cluster center The average value of all segments in this cluster.

[0147] 2) The extreme segment X Cij : The segment with the largest variance in this cluster.

[0148] The length of the mapped cluster is w = l × 2, where l is the length of the mapped segment, and multiplying by 2 is because each cluster contains two main features: the cluster center and the extreme segment. Therefore, for each cluster C ij , we represent its mapping as:

[0149]

[0150] At the same time, consider the remaining features of the time series in this step.

[0151] 1) The error difference between the most dissimilar (farthest) and most similar (nearest) segments in the time series, measuring their similarity to the cluster center, and this error is evaluated by the mean square error (MSE).

[0152] 2) N Ci : The number of segments of the time series.

[0153] Select a time series as a reference, calculate the cluster centers of all other time series, and match them with the cluster center of the reference time series. Select the cluster that is closest to the reference center for correspondence.

[0154] Once the matching is completed, each time series will be converted into a mapped time series X i′, the time series includes the extracted clustering features. The length of the mapped time series is: (w × k) + v.

[0155] Among them, w × k is the total sum of the feature dimensions of all clusters, and v = 2 are two additional pieces of information, including the error difference and the number of segments.

[0156] Step 3: Establish a classification model for the industries to which power users belong based on the fully automatic density method:

[0157] Use the density-based clustering method to perform clustering analysis on the mapped spatial features, so as to classify the time series into different groups.

[0158] The density-based clustering method includes Eps and Min Pts Two important parameters. Eps is the radius of the circular area around a specified point, and the points covered are used to check the density. And Min Pts is the threshold of the number of points in the Eps area, which is used to consider that these points belong to the same density.

[0159] Step 3.1, calculate the radius Eps of the circular area around the specified point:

[0160] The size of Eps is closely related to the characteristics of the data distribution and the distance between points.

[0161] For each data point P(x, y), calculate the distance between this point and its nearest neighbor. Given the coordinates (x, y) of the data point and the coordinates (x c , y c ) of the neighbor point, the distance D min (P(x, y)) of the two nearest neighbors is calculated by the following formula:

[0162]

[0163] Apply the above formula (3-1) to each point in the mapped time series X i ′ to calculate the shortest distance from each point to its nearest neighbor. Store all these shortest distances in a list ListD and sort them in ascending order:

[0164] List D = Sorted(List D ) (3-2)

[0165] Next, apply a statistical histogram to display the data distribution of the extracted distances. There are close-range values in the dataset space. To handle the problem of small differences between these distances, apply a normalization process to the list to balance the close distances into equal-range intervals. Calculate the normalized distance list by formula (3-3):

[0166]

[0167] Where R is the sub-range value where the distances become similar after normalization, and is calculated by the following formula:

[0168] R = int(N × C) (3 - 4)

[0169] Where N is the number of sampling points in the data set and C is the division factor.

[0170] The data distribution is controlled by the interval range of the distance values and the frequency of the distances. When determining the optimal Eps value in data sets with different densities, these two factors interact, resulting in a more complex distribution in the data set. Therefore, the number of non-zero distances in the normalized histogram is divided by the distance range R Av , thus obtaining the frequency value F, which represents the degree of noise and long distances in the data set:

[0171]

[0172] Where Hstg A represents the number of non-zero frequency points within the histogram range, and R Avg is the average value of the data points within each histogram range.

[0173] By calculating the distance range between the maximum value P max and the minimum value P min in the data points P, and then adjusting Eps according to the characteristics of noise and outliers. Finally, the following formula is used to calculate the final Eps value:

[0174] Eps = P min + ((P min - P max ) × F) (3 - 7)

[0175] This method can dynamically adjust the Eps value according to the distribution characteristics of the data, making it adapt to data sets with different densities and avoiding over-segmentation or merging of clusters.

[0176] Step 3.2, calculate the minimum number of points Min Pts :

[0177] Min Pts is used to define how many points at least are required to be considered a density-connected cluster within a specified radius Eps.

[0178] First, use Eps to scan all the points in the data set and calculate the number of points within Eps:

[0179]

[0180] Where DEps (p(x, y)) represents the number of points within the Eps radius around the given point P(x, y). This process is repeated for each point in the dataset, and the results are finally stored in a list, as shown in Equation (3-10).

[0181]

[0182] Store the results in List Eps After that, calculate the data graph of List Eps and determine the most frequent value. In an organized dataset, this value corresponds to the central point and is suitable to be set as Min Pts ; in some datasets with highly diverse distributions, the statistical standard deviation of the number of points in Eps provides a more accurate measure and is suitable to be set as Min Pts . Therefore, set the value of Min Pts to the maximum value between the two parameters in the dataset.

[0183] Step 3.3, assign cluster labels to the sample points in the dataset:

[0184] This step mainly identifies high-density regions (i.e., potential clustering cores) by detecting core points in the data. A core point is defined as a point whose number of neighboring points is greater than or equal to the preset minimum number of points Min Pts .

[0185] For each point P(x, y) in the dataset, construct a circular region with a radius of Eps and calculate the number of points within this region:

[0186]

[0187] where L is the label;

[0188] If the number of points within this region is greater than or equal to Min Pnts , then this point is considered a clustering core, and all points within its neighborhood are labeled with the same cluster.

[0189] Cluster label assignment:

[0190] (1) If there are no labeled data points in the neighborhood, assign the next value of the maximum label L max previously used to the current point. This can ensure that a unique label is assigned to each cluster and avoid any potential conflicts with existing labels:

[0191] L[P] = L max + 1 (3-12)

[0192] (2) If at least one adjacent data point has been assigned a label in the previous core clustering detection step, then assign the current point to the label of the point with the smallest label in the neighborhood.

[0193] L[P] = min(L)(3 - 13)

[0194] If the number of points N within the Eps region is less than the value of Min Pnts , it is considered that the density of this region is too low to form a cluster. Repeat the step of assigning cluster labels for all points in the dataset to identify high-density regions as potential cluster cores.

[0195] Next, process all unclustered points P in the dataset to check if there are any labeled points within their neighborhood Eps. If there are, then assign the current point to the nearest labeled point in the neighborhood. Iteratively apply this process until all points are labeled as members of a cluster. However, in some datasets with highly distributed data, some outlier points may still not be assigned to potential clusters, so these outlier points can be connected to the nearest cluster points to assign labels.

[0196] For the dataset with complex shapes or low density in the previous step, the cluster labels may still be incomplete and it is still necessary to merge connected clusters.

[0197] Identify connected clusters: If some clusters are "connected", that is, they share some common points or are close enough in the dataset, then they should be merged into one cluster. This process tracks the labels of neighboring points through a label association list.

[0198] Merge clusters: Use the Union-Find method of disjoint-set data structure to identify those clusters that are connected to each other and merge them into one cluster, that is, merge clusters with shared labels. The merged clusters will be re-assigned labels, and the new merged cluster will be given the smallest label value, ensuring label consistency and avoiding problems of duplicate or inconsistent merged clusters. At the same time, the final clustering result after merging can better reflect the density and spatial structure of the data. The output of this step is the final clustering result of the proposed algorithm.

[0199] Step 4: Industry mapping and matching correction, including:

[0200] Step 4.1, Industry mapping and matching correction:

[0201] According to the clustering result obtained in Step 3, assign a corresponding industry category Z to each cluster label j , and this mapping relationship is constructed based on the actual application requirements of the power industry by matching the characteristics of the industry feature template with the characteristics of the cluster labels.

[0202] Match the characteristics of the assigned clustering label L with the industry characteristics, and select the closest industry classification. The industry characteristics include the duration of the peak period, load volatility, peak and valley values, etc. in the load curve. Calculate the similarity distance between each clustering label and all industry characteristic templates, and select the industry category corresponding to the industry characteristic template with the smallest distance as the industry classification result of this clustering.

[0203]

[0204] Where Z j is the industry category of the clustering label L i ; Z k is the industry characteristic template; Distane is the distance function for feature matching.

[0205] Traverse all users, extract the original registered industry information Z u , and compare it with the assigned industry identifier Z j ;

[0206] If Z u =Z j , it is considered a successful match, and record the industry classification of this user;

[0207] If Z u ≠Z j , it indicates that the electricity consumption nature of the user has changed. Record the industry classification after the change of this user, and mark this user as "pending manual verification" status.

[0208] Step 4.2, Classification result statistics and user information registration:

[0209] Statistically count the number of users in each industry category and their main load characteristics, and generate a classification statistical table, including information such as industry category, number of users, average load characteristics, etc. Then use visualization methods such as bar charts and heat maps to display the classification results.

[0210] Record all user information where Z u ≠Z j , including user ID, original industry category, current load characteristics, recommended industry category, etc., and send a notice to relevant management personnel to further verify the industry classification of this electricity user.

[0211] Step 5: Clustering performance analysis:

[0212] Step 5.1, Rand index: The Rand index is used to measure the similarity between two labeled data sets. It measures the consistency between the clustering result and the true label by ignoring the permutations of combinations.

[0213]

[0214] Where a is the number of identical point samples found in the same group of A and the same group of B;

[0215] b is the number of identical point samples found in different groups of A and different groups of B;

[0216] c is the number of identical point samples found in the same set of A and different sets of B;

[0217] d is the number of identical point samples found in different groups of A and the same group of B.

[0218] In the formula, A is the true value of the label and B is the aggregation result.

[0219] Step 5.2, Homogeneity: Homogeneity measures whether each cluster mainly contains samples of the same category.

[0220]

[0221] In the formula, H(C|K) represents the distribution degree of points belonging to different categories in each cluster, and H(C) represents the entropy of category C.

[0222]

[0223] In the formula, n c,k is the number of samples belonging to category c and belonging to cluster k, n k is the total number of samples in cluster k, and N is the total number of samples in the dataset.

[0224] Step 5.3, Completeness: Completeness evaluates whether similar samples are assigned to the same cluster.

[0225]

[0226] In the formula, H(K|C) represents the distribution of samples assigned to different clusters in each category, and H(K) represents the entropy of cluster K.

[0227]

[0228] In the formula, n c,k is the number of samples belonging to category c and belonging to cluster k, n c is the total number of samples in category c, and N is the total number of samples in the dataset.

[0229] Step 5.4, Harmonic Mean: The harmonic mean is the weighted average of Homogeneity and Completeness, usually calculated using the Normalized Mutual Information (NMI).

[0230]

[0231] In the formula, H represents homogeneity and C represents integrity.

[0232] Step 5.5: Calculate the above Rand index, homogeneity, integrity, and harmonic mean based on the original data and features provided in Step 1 and Step 2, and summarize the calculation results of each index as shown in the following table:

[0233] Index Value Rand Index 0.85 Homogeneity (H) 0.90 Completeness (C) 0.87 Harmonic Mean (F) 0.88

[0234] Verify the rationality of the clustering result obtained in Step 3 and the effect of the industry classification in Step 4 through performance evaluation.

[0235] The embodiments described above are only descriptions of the preferred modes of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations, variations, modifications, and substitutions made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. The power user industry classification analysis method based on the fully automatic density clustering algorithm is characterized by: The following steps are involved: Step 1: Collect real-time power user data, and use multiple indicators such as daily maximum power, peak-to-total ratio, flat-to-valley ratio, valley-to-total ratio, and daily average load as cluster analysis indicators for normalization: In the formula, x i is the initial data, x max is the maximum daily load, x' i is the normalized value; Perform smoothing to generate a smoothed load series: Step 2: Preprocess the data based on the time series feature stratification method; Step 3: Establish a classification model for the industry to which power users belong based on the fully automatic density method; Step 4: Industry mapping and matching correction; Step 5: Clustering performance analysis.

2. The method for classifying and analyzing power users by industry based on a fully automatic density clustering algorithm according to claim 1 is characterized in that: Step 2 includes: Step 2.1, segment the time series of power load: For a time series X of length I = {x1, x2, ... x i } to split the time series into m segments and find m-1 cut points Iteratively introduce time series data point by point into the dynamically growing window and calculate the squared error SSE within the window q , the maximum standard error is SEP max , when SSE q >SEP max When , the current point is marked as the split point t i , and then end the current segment, and each time the error is detected to exceed the threshold, the cut point t i For the record, the initial segment starts at t0=1. When t1=n is detected, it means that the first n data points belong to the first segment. The cut point set {t1, t2, …t m-1 }Define fragment range: fragment q: [t q-1 +1,t q ]; Step 2.2, segment mapping and grouping the segmented fragments: Each segment is mapped into an array including polynomial fitting parameters and statistical features. Each segment is projected into l-dimensional space, where l is the length of the mapped segment. The polynomial coefficients p are directly obtained from the piecewise least squares fitting process. q , and calculate the following statistical features: variance Skewness γ 1q : Autocorrelation coefficient AC q : By combining polynomial fit parameters and statistical features: where p q are the polynomial approximation parameters of the approximation segment q; Divide the fragments into two categories, with the number of groups k = 2; Step 2.3, perform time series mapping on the grouped segments: For each time series X i , extract the corresponding cluster centers in: i∈{1,…,T} represents the time series index, T is the number of time series; j∈{1,…,k} represents the index of the cluster in each time series, and k is the number of clusters; For each cluster, C ij Extract the following: Cluster centers The average of all fragments in this cluster; extreme fragments The fragment with the largest variance in this cluster; The length of the mapped cluster is w = l × 2, where l is the length of the mapped segment. For each cluster Its mapping is represented as: Select one time series as a reference, calculate the cluster centers of all other time series, match them with the cluster center of the reference time series, and select the cluster closest to the reference center to correspond; When matching is completed, each time series will be converted into a mapped time series X i ′, the time series includes the extracted cluster features, and the length of the mapped time series is: (w×k)+v, where w×k is the sum of the feature dimensions of all clusters, and v=2 is two additional information, including the error difference and the number of fragments.

3. The method for classifying and analyzing power users by industry based on a fully automatic density clustering algorithm according to claim 1 is characterized in that: Step 3 includes: Step 3.1, calculate the radius Eps of the circular area around the specified point: For each data point P(x,y), calculate the distance between the point and its nearest neighbor, given the coordinates (x,y) of the data point and the coordinates (x c ,y c ), the distance D between the nearest neighbors of two points min (P(x,y)) is calculated by the following formula: For the mapped time series X i Apply formula (3-1) to each point in ′ to calculate the shortest distance from each point to its nearest neighbor, and store all the shortest distances in a list List D In ascending order: List D =Sorted(List D )(3-2) A statistical histogram is used to display the data distribution of the extracted distance, and the normalized distance list is calculated by formula (3-3): Where R is the sub-range value where the distance becomes similar after normalization, calculated by the following formula: R = int(N×C)(3-4) Where N is the number of sampling points in the data set, and C is the division factor; Divide the number of non-zero distances in the normalized histogram by the distance range R Av , thus obtaining a frequency value F, which represents the degree of noise and long distances in the data set: Where Hstg A Indicates the number of non-zero frequency points within the histogram range, and R Avg is the average of the data points within each histogram range; By calculating the maximum value P among the data points P max and the minimum value P min The distance range between them is then adjusted according to the characteristics of noise and outliers. The following formula is used to calculate the final radius Eps value: Eps=P min +((P min -P max )×F)(3-7) Step 3.2, calculate the minimum number of points Min based on the final radius Eps value Pts : Scan all points in the dataset with radius Eps and count the number of points within radius Eps: Where D Eps (p(x,y)) represents the number of points within the radius Eps around a given point P(x,y). Repeat this process for each point in the dataset and store the results in a list: Calculating Lists Eps Graph of the data and determine the most frequent value; Step 3.3, assign cluster labels to sample points in the data set: For each point P(x,y) in the data set, construct a circular area with a radius of Eps and calculate the number of points in the area: If the number of points in the area is greater than or equal to Min Pnts , then the point is considered to be a cluster core, and all points in its neighborhood are marked as the same cluster; If there is no labeled data point in the neighborhood, the maximum label L used previously is assigned to the current point max The next value of: L[P]=L max +1(3-12) If at least one neighboring data point has already been assigned a label, the current point is assigned the label of the point with the smallest label in the neighborhood: L[P]=min(L)(3-13) Where L is the label; If the number of points N within the radius Eps is less than Min Pnts If the value of is too low, the density of the area is considered too low to form a cluster; repeat the cluster label assignment step for all points in the data set to identify the high-density area as a potential cluster core; Process all unclustered points P in the data set and check whether there are marked points within their neighborhood radius Eps. If so, assign the current point to the nearest marked point in the neighborhood. Apply this process iteratively until all points are marked as members of a cluster. Merge clusters: Use the Union-Find method to identify clusters that are connected to each other and merge them into one cluster. The merged clusters are reassigned labels, and the merged new cluster is given the smallest label value to obtain the final clustering result.

4. The method for classifying and analyzing power users by industry based on a fully automatic density clustering algorithm according to claim 3 is characterized in that: Step 4 includes: Step 4.1, industry mapping and matching correction: Based on the clustering results in step 3, assign a corresponding industry ID to each cluster label and compare it with the user's registered industry information to ensure the accuracy of the classification; According to the characteristics of the assigned cluster label L and the industry characteristics, the closest industry classification is selected and the cluster label industry identifier Z is given j : Where Z j is the cluster label L i Industry logo; Z k is the industry feature template; Distane is the distance function for feature matching; Traverse all users and extract the industry information Z originally registered by the user u , and with the assigned industry identifier Z j Make a comparison; If Z u =Z j , then the match is considered successful and the industry classification of the user is recorded; If Z u ≠Z j , it indicates that the user's electricity usage has changed, the user's changed industry classification is recorded, and the user is marked as "pending manual verification"; Step 4.2, classification result statistics and user information registration: Count the number of users in each industry category and their main load characteristics, and generate a classification statistics table, including industry category, number of users, and average load characteristics, and then use bar charts and heat maps to display the classification results; Record all Z u ≠Z j The user information, including user ID, original industry category, current load characteristics, and recommended industry category, is collected and a notification is sent to the management personnel to verify the industry classification of the power user.

5. The method for classifying and analyzing power users by industry based on a fully automatic density clustering algorithm according to claim 1 is characterized in that: Step 5 includes: Step 5.1, Rand Index: Where, a: the number of identical point samples found in the same group of A and the same group of B; b: The number of samples with the same points found in different groups of A and different groups of B; c: the number of identical point samples found in the same set of A and in different sets of B; d: the number of identical point samples found in different groups of A and the same group of B; In the formula, A is the true value of the label, and B is the aggregation result; Step 5.2, Homogeneity: In the formula, H(C|K) represents the distribution degree of points belonging to different categories in each cluster, and H(C) represents the entropy of category C; Where n c,k is the number of samples belonging to category c and cluster k, n k is the total number of samples in cluster k, and N is the total number of samples in the data set; Step 5.3, Completeness: Where H(K|C) represents the distribution of samples assigned to different clusters in each category, and H(K) represents the entropy of cluster K; Where n c,k is the number of samples belonging to category c and cluster k, n c is the total number of samples of category c, and N is the total number of samples in the dataset; Step 5.4, harmonic mean: The harmonic mean is a weighted average of homogeneity and completeness, calculated using normalized mutual information: In the formula, H represents homogeneity and C represents completeness; Step 5.5, based on the original data and features provided in Steps 1 and 2, calculate the Rand index, homogeneity, completeness, and harmonic mean, and summarize the calculation results of each indicator; verify the rationality of the clustering results obtained in Step 3 and the effect of the industry classification in Step 4 through performance evaluation.