High-price low-connection user identification method based on PSO optimization clustering and Hausdorff distance discriminant analysis

By using the PSO optimization clustering and Hausdorff distance discriminant analysis methods, the problems of missed detection and computational complexity in the identification of high-priced, low-connection users were solved, efficient and accurate user identification and real-time monitoring were achieved, and the automation level of the power system was improved.

CN120706804APending Publication Date: 2025-09-26SHENYANG INST OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510831057.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies have low sensitivity and a high missed detection rate when identifying illegal electricity consumption behaviors such as high prices and low connections. In addition, existing algorithms have high computational complexity and poor real-time performance, making it difficult to meet the needs of refined supervision of the power system.

Method used

A method based on PSO optimization clustering and Hausdorff distance discriminant analysis is adopted. The K-means clustering is optimized by particle swarm optimization algorithm, combined with Hausdorff distance discriminant analysis to identify high-price low-connection users, reduce manual intervention, and improve clustering accuracy and real-time performance.

Benefits of technology

It significantly improves the accuracy and efficiency of identifying high-priced low-connection users, reduces computational complexity, enhances automated detection capabilities, reduces economic losses, and ensures stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706804A_ABST
    Figure CN120706804A_ABST
Patent Text Reader

Abstract

The invention relates to a user identification algorithm, in particular to a high-price low-answer user identification method based on PSO (Particle Swarm Optimization) clustering and Hausdorff distance discriminant analysis. The method enables a power enterprise to recognize illegal users more efficiently, thereby effectively reducing the economic loss caused by default power consumption, and guaranteeing the safe and stable operation of a power system. Comprising the following steps: S1, inputting a power consumer power utilization data set, and carrying out data preprocessing; s2, optimizing a K-means clustering algorithm based on particle swarm optimization, and generating a typical power consumption track of the commercial user; s3, carrying out Hausdorff distance discriminant analysis on the clustered typical electricity utilization tracks of the commercial users and the electricity utilization tracks of the common residential users so as to preliminarily identify a series of users with abnormal electricity utilization behaviors; and S4, according to a preset threshold value, judging whether the series of common resident users with the abnormal electricity consumption behaviors preliminarily identified in the step S3 are high-price low-price users, and performing early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a user identification algorithm, in particular to a high-price low-access user identification method based on PSO optimized clustering and Hausdorff distance discriminant analysis. Background Art

[0002] With the rapid evolution of smart grids, power companies have gradually established a two-way interactive system between power and information systems, achieving digital coverage of the entire power supply chain (from production to management). While this innovation has improved the convenience of electricity use for residents, it has also created a new management challenge: the increasingly prominent problem of non-technical power losses (NTL). Consumer violations have become a major driver of this problem, particularly illegal high-price, low-connection practices (i.e., users illegally connecting high-priced equipment to low-priced lines), which have resulted in significant economic losses for power companies. Due to the difficulty of oversight, low-voltage distribution areas are prone to illegal connections and rate evasion. Therefore, accurately detecting, identifying, and curbing these violations has become a core issue in smart grid management.

[0003] Current technology for detecting power users’ default on electricity usage faces multiple bottlenecks:

[0004] 1. Mainstream detection methods are less sensitive to highly concealed violations such as "high prices and low acceptance", and the problem of missed detection is prominent, making it difficult to meet the needs of refined supervision.

[0005] Second, technical solutions that rely on complex algorithms such as deep learning and time-frequency analysis are unable to adapt to real-time monitoring scenarios of large-scale electricity consumption data due to their large computational workload and slow processing speed, resulting in delayed detection of violations.

[0006] 3. Existing technologies often require manual intervention in data labeling, feature extraction and other links, which reduces the system's automation level and makes it difficult to meet the dynamic analysis needs of massive data.

[0007] 4. Traditional clustering algorithms are prone to clustering bias when processing electricity consumption data with multiple dimensions such as voltage, current, and load patterns, which affects the accuracy of detection results. Summary of the Invention

[0008] The present invention addresses the shortcomings of the existing technology and provides a method for identifying high-priced, low-connection users based on PSO optimized clustering and Hausdorff distance discriminant analysis. This method can accurately identify high-priced, low-connection users and avoid the missed detection problem of traditional methods. By optimizing the clustering algorithm through PSO, the clustering efficiency and accuracy are improved, the computational complexity is reduced, and higher real-time performance is ensured when processing large-scale power data. At the same time, this method greatly reduces the need for manual intervention and enhances the automated detection capability, allowing power companies to more efficiently identify illegal users, thereby effectively reducing the economic losses caused by illegal electricity use and ensuring the safe and stable operation of the power system.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: a method for identifying high-priced, low-connection users based on PSO optimization clustering and Hausdorff distance discriminant analysis, comprising the following steps:

[0010] S1. Input the electricity consumption data set of power users and perform data preprocessing;

[0011] S2, using the particle swarm optimization (PSO) algorithm to optimize the K-means clustering algorithm and generate typical electricity consumption trajectories of commercial users;

[0012] S3. Perform Hausdorff distance discriminant analysis on the clustered typical electricity usage trajectories of commercial users and ordinary residential users, and calculate the similarity between the two to preliminarily identify a series of users with abnormal electricity usage behaviors.

[0013] S4: Based on a preset threshold, determine whether a series of ordinary residential users with abnormal electricity usage behavior initially identified in S3 are high-price, low-connection users, and issue an early warning. The early warning information includes the user ID, the matching commercial user category, and the degree of deviation.

[0014] Furthermore, S1 includes:

[0015] S1.1. Data cleaning, including outlier processing and missing value recovery;

[0016] S1.2. Perform data standardization and Gaussian smoothing, and perform Gaussian smoothing on the standardized load data; the standardization process uses the Z-Score method, which includes calculating the mean and standard deviation at each time point, and normalizing the original load value to obtain the standardized load sequence Z′.

[0017] Furthermore, S1.1 includes:

[0018] S1.1.1. Obtain the monthly electricity consumption trajectory set M = {m1, m2, ..., m n}, where mj The monthly electricity consumption trajectory of the jth electricity user constitutes a power load time series; and the user's power load data is collected every 15 minutes;

[0019] S1.1.2. Perform outlier processing on the collected power load time series and delete the time series; if the negative value ratio is less than 20%, the negative value will be regarded as missing value;

[0020] S1.1.3. Identify and correct the error values ​​in the power load time series. The correction formula is:

[0021]

[0022] Where x i ; represents the user's power load data value, σ(X i ) represents vector X i The standard deviation of x is NaN. i ;

[0023] S1.1.4. Recover the missing values ​​marked in S1.1.2 using the following formula:

[0024]

[0025] Where, mean(X i ) represents vector X i The average value of .

[0026] Furthermore, S2 specifically includes:

[0027] S2 includes using the particle swarm optimization algorithm (PSO) to optimize the K-means clustering algorithm to determine the optimal number of clusters K, the initial cluster center, the number of algorithm iterations, and the number of initialization runs, and to obtain the optimal clustering result with the maximization of the silhouette coefficient as the objective function; specifically, it is divided into:

[0028] S21, initialize the particle swarm, each particle represents a group of cluster centers;

[0029] S22, calculate the fitness value of each particle and optimize it with the silhouette coefficient as the objective function;

[0030] S23, by iteratively updating the position and velocity of the particles, dynamically adjusting the inertia weight to find the optimal cluster center, and stopping the iteration when the fitness curve of the objective function tends to be flat;

[0031] S24. When the iteration reaches a stable state, the optimal clustering result and the corresponding typical power consumption trajectory are output. The stable state means that the objective function value converges when the number of iterations reaches 100.

[0032] Furthermore, S3 specifically includes:

[0033] S31. The typical electricity consumption trajectory of commercial users is represented as a point set A = {a1, a2, ..., a q}, the electricity consumption trajectory of ordinary residential users is represented by point set B n ={b1,b2,...,b q}, where n∈N + Representing each ordinary residential user;

[0034] Among them, a i ,b j Respectively represent the power consumption of each power consumption trajectory at a certain moment;

[0035] S32, calculate the one-way similarity distance h(A,B n ):

[0036]

[0037] represents the maximum value of the minimum distance between each point of the commercial trajectory and the resident trajectory;

[0038] And calculate the one-way similarity distance h(B n ,A);

[0039]

[0040] represents the maximum value of the minimum distance between each point of the resident trajectory and the commercial trajectory;

[0041] Among them, d{A(a i ),B n (b j )}=||A(a i )-B n (b j )|| represents a point on the typical electricity consumption trajectory A of a commercial user to the electricity consumption trajectory B of an ordinary residential user n The Euclidean distance of

[0042] S33: Take the maximum value of the two-way distance as the comprehensive similarity measure;

[0043] H(A,B n )=max{h(A,B n ),h(B n ,A)}

[0044] Where H(A,B n ) value is smaller, indicating a higher similarity; output the minimum H(A,B n ) value to S4 for threshold determination.

[0045] Furthermore, in S4, determining the preset threshold includes:

[0046] Extract typical electricity consumption trajectories of each commercial user category from the training set samples;

[0047] Calculate the Hausdorff distance between the electricity consumption trajectory of each user marked as abnormal in the training set and the typical trajectory of the corresponding commercial user category;

[0048] The maximum value of the Hausdorff distance of abnormal users corresponding to each type of commercial users is taken as the preset threshold ε of this category.

[0049] Compared with the prior art, the present invention has beneficial effects.

[0050] This method effectively overcomes the shortcomings of traditional detection methods, particularly in identifying high-price, low-cost connections, avoiding missed detections. It significantly improves clustering accuracy and processing efficiency, reduces the computational complexity of traditional clustering algorithms when processing large-scale data, and ensures real-time responsiveness. Furthermore, by combining the Hausdorff distance with discriminant analysis of electricity usage patterns, the method can accurately distinguish different types of users when processing complex electricity usage data, improving detection accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The present invention is further described below with reference to the accompanying drawings and specific embodiments. The scope of protection of the present invention is not limited to the following description.

[0052] Figure 1 It is the daily electricity consumption trajectory curve of power users.

[0053] Figure 2 It is a schematic diagram of Gaussian smoothing of the curve.

[0054] Figure 3 This is the PSO optimization k-means clustering flow chart.

[0055] Figure 4.1 It is the flow chart of Hausdorff distance discriminant analysis.

[0056] Figure 4.2 It is the PSO-kmeans clustering optimization curve.

[0057] Figure 4.3 This is the electricity consumption trajectory of four typical types of users.

[0058] Figure 4.4 These are the electricity consumption screening results for four categories of commercial users.

[0059] Figure 4.5 Optimization of the three-category commercial judgment threshold. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and beneficial effects of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0061] Specific preferred implementation plan: Step 1.1, data cleaning.

[0062] Example 1 uses electricity consumption data of urban residential users in a complete low-voltage area from March 2021 to February 2022, extracted from a power supply bureau's metering center and marketing system (Note: the electricity consumption types in the data source include commercial electricity and ordinary residential electricity). Among them, the training set is the electricity consumption data of 80% of residential users in this time period; the experimental set is the data of 20% of residential users in the same period. As of February 2022, the system has marked 43 defaulting users.

[0063] Table 2.1 Basic information of data samples.

[0064]

[0065] Example 1: M={m1,m2,…,m n} is the set of monthly electricity consumption trajectories of power users in the substation, where m j is the trajectory curve of monthly electricity consumption of the jth electricity user, such as Figure 1 The following figure shows the daily power consumption trajectory curves of all electricity users. The power load data of each user is collected every 15 minutes, which can be used to identify and analyze the patterns and trends of load changes.

[0066] Dealing with outliers in the dataset:

[0067] During equipment maintenance or replacement, the meter may restart counting from zero, resulting in anomalies such as negative values ​​when calculating daily electricity consumption. Example 1 handles time series containing negative values ​​as follows: if more than 20% of the time series are negative, the time series is deleted and not included as a sample in subsequent experiments; if less than 20% of the time series are negative, it is considered a missing value and handled as the default value in the next step.

[0068] Handling of default values ​​in datasets:

[0069] During the electricity load data collection process, due to factors such as software and hardware failures or special events, erroneous values ​​or missing values ​​may appear, which will affect the continuity of electricity consumption records. Therefore, it is necessary to preprocess the original data set.

[0070] For the recovery of error values, Example 1 selects the 3σ law for correction, and its calculation formula is as follows:

[0071]

[0072] Where x i ; represents the user's power load data value, σ(X i ) represents vector X i The standard deviation of x is NaN. i ; is not a number sign.

[0073] For missing values ​​in the data set, the following calculation formula is used to restore them

[0074]

[0075] Where, mean(X i ) represents vector X i The average value of .

[0076] Step 1.2: Data standardization and curve trajectory smoothing;

[0077] Because there are many different types of users in the low-voltage area, the electricity consumption behaviors of different users vary greatly. Therefore, before performing cluster analysis and identifying ordinary residential users, the data needs to be standardized. Example 2 uses the Z-Score normalization method to normalize the load data to remove the influence of base load and highlight the trend characteristics of variable load, while avoiding the influence of large differences in the order of magnitude of electricity consumption of each user on subsequent model training and convergence. The formula is:

[0078]

[0079] Z'=[Z'1,Z'2,...,Z' D ]

[0080] Z′ t is the load data Z at time t t The normalized value, f mean (Z t ) is Z at time t t The standard deviation of f std (Z t ) is the standard deviation at time t; Z∈N user ×D is the data set of user load in the area, N user is the number of users. The normalized Z′ is used as the input for subsequent clustering and distance discrimination.

[0081] In addition, in order to further improve the continuity and stability of the load curve, Example 2 also performs Gaussian smoothing on the normalized load curve. Figure 2Figure 2 shows a comparison of power user load curves before and after Gaussian smoothing. After Gaussian smoothing, the load curves become more continuous and complete along the time axis, effectively removing noise and sudden changes, making load trends and cyclical characteristics more distinct. This helps clustering algorithms more accurately identify user groups with similar electricity usage patterns or behavioral characteristics.

[0082] Step 2: K-means clustering optimization based on PSO. Due to the significant complexity and diversity of residential electricity consumption behavior, traditional clustering methods such as K-means suffer from low clustering accuracy and slow convergence in practical applications. Therefore, Example 3 proposes a K-means clustering method based on a particle swarm optimization (PSO) algorithm to improve the accuracy and stability of clustering results.

[0083] In this embodiment, PSO is used to optimize the key parameters of K-means clustering, including: the number of clusters K with the maximum silhouette coefficient, the number of algorithm iterations, the number of times the algorithm is initialized and run with different cluster centers, and the initial cluster center.

[0084] like Figure 3 As shown in the figure, the PSO optimization k-means clustering process specifically includes: first, the parameters are digitized, then the PSO algorithm particles are generated, and then the clustering operation of the PSO optimization scheduling is performed, and the silhouette coefficient is calculated to determine whether it converges. If it converges, the corresponding K-means cluster center position is output. If it does not converge, it returns to the clustering operation step of the PSO optimization scheduling.

[0085] In step 3, in the Hausdorff distance, let the two sets of points be:

[0086] A={a1,a2,...,a q}

[0087] B={b1,b2,...,b q}

[0088] Among them, A and B represent the sampling point sets on the two electricity consumption curves. i ,b j Represent the parameters of curves A and B respectively.

[0089] The Hausdorff distance between curves A and B can be defined as one-way and two-way:

[0090] Among them, the one-way Hausdorff distance from curve A to curve B is defined as:

[0091]

[0092] Among them, d{A(a i ),B(b j )}=||A(a i )-B(b j )|| represents the Euclidean distance from a point on curve A to curve B.

[0093] The one-way Hausdorff distance from curve B to curve A is:

[0094]

[0095] The bidirectional Hausdorff distance between curve A and curve B is defined as:

[0096] H(A,B)=max{h(A,B),h(B,A)}

[0097] The Hausdorff distance reflects the maximum deviation between two point sets. In power load data analysis, by calculating the Hausdorff distance between the user's power consumption curve and the typical trajectory, users with abnormal power consumption behavior can be effectively identified.

[0098] like Figure 4.1 In order to accurately identify high-price low-connection users, it is necessary to set a reasonable Hausdorff distance threshold ε to determine whether a user's electricity consumption curve has a sufficiently high similarity with the typical trajectory. The two electricity consumption curves are: F1 and F2, where c1 and c2 are the characteristic segments of the two curves, c1 = {k1, k2, ..., k m},c2={K1,K2,…,K m}, then the directed Hausdorff distance between point sets c1 and c2 is:

[0099]

[0100] Normally, the values ​​of h(c1, c2) and h(c2, c1) are not equal. However, since the data structures of the c1 and c2 point sets are basically the same, in this embodiment, they are considered equal. Therefore, the Hausdorff distance is defined as:

[0101] H(c1,c2)=max{h(c1,c2),h(c2,c1)}

[0102] H(c1,c2)≤ε

[0103] In step 4, a threshold ε is set based on the distance value H to determine the similarity of the curves. A smaller ε indicates a higher similarity between the two curves. In this embodiment, the threshold ε is set by taking the maximum Hausdorff distance between the anomalous user and the typical trajectory of each class of labeled training set samples as the threshold for that class of anomalous users. Finally, the ε obtained in the training set is applied to the experimental set to further verify the recognition accuracy of the model.

[0104] Example 1: Analyze the daily electricity consumption data of 1,560 residential users in a low-voltage area in a certain region. Figure 4.2 The curve shows how the clustering quality changes with the increase in the number of iterations. Figure 4.2 It can be seen from the figure that: in the early stage of iteration, the fitness of the objective function will drop rapidly, which represents the process of the algorithm continuously optimizing the cluster center. As the number of iterations increases, the objective function value gradually converges to a stable level, indicating that the algorithm has found a set of better cluster centers. Finally, after 100 iterations, the curve will tend to be flat, indicating that the algorithm has converged and achieved a stable clustering result. The clustering results show that after cluster analysis, commercial users generate four typical categories of users, namely: Category 1 commercial users, Category 2 commercial users, Category 3 commercial users, and Category 4 commercial users. Their electricity consumption trajectories are as follows: Figure 4.3 As shown. Combined Figure 4.3 The user's electricity consumption trajectory is used to analyze the user's electricity consumption behavior and obtain the electricity consumption behavior characteristics of different categories of users, as shown in Table 4.1.

[0105] Table 4.1 Electricity consumption characteristics of four types of users

[0106]

[0107] As shown in 4.3, commercial users are clustered and analyzed. Based on the clustering results, we can conclude that: (1) the number of users in the four clusters is reasonably distributed and the clustering is effective; (2) the typical electricity consumption trajectories of the four types of commercial users are quite different, with obvious cluster center characteristics and strong representativeness.

[0108] This paper uses electricity load data from low-voltage power users in the low-voltage area from February 2021 to February 2022 to validate the high-price, low-connection user identification and detection method. As of the end of February 2022, the number of abnormal users in the experimental set marked by the marketing system using the power anomaly recognition model was 176.

[0109] The Hausdorff distance was used to discriminate the electricity consumption trajectories of the 176 abnormal residential electricity users and the typical electricity consumption trajectories of commercial users in the above experimental set, and then the identification and screening were carried out in the experimental set. The screening results are as follows: Figure 4.4 shown.

[0110] Depend on Figure 4.4 It can be seen that when performing discriminant analysis on the three types of commercial users, the high-price low-connection users cannot be identified due to the problem of threshold selection. The threshold selection principle of the present invention increases the threshold H3 from 1.337 to 1.470 under the premise that the audit workload does not exceed 5%. Figure 4.5 As shown in Figure 1, after the threshold is increased, the number of audited users increases from 8 to 13, which meets the threshold selection condition. Finally, the final abnormality judgment is made by combining on-site audit.

[0111] To evaluate the accuracy of the model, a confusion matrix is ​​introduced to evaluate the accuracy and recall of the Hausdorff distance discriminant model used in this paper for identifying high-price, low-acceptance users. The definition of the confusion matrix for high-price, low-acceptance user detection and recognition is shown in Table 4.2.

[0112] Tab.4.2 Confusion matrix of high-price low-access user detection and recognition

[0113]

[0114] In Table 4.2, the correct classification results and incorrect classification results of the recognition test are represented by the letters T and F respectively; the predicted abnormality and predicted normality are represented by the letters P and N respectively. The recall rate and precision rate can be calculated based on the confusion matrix:

[0115]

[0116] R stands for recall; ATP represents the number of users whose predicted results were abnormal but whose actual results were also abnormal; AFN represents the number of users whose predicted results were normal but whose actual results were abnormal. A higher recall (R) indicates better detection model performance. P stands for precision; AFP represents the number of users whose predicted results were abnormal but whose actual results were normal. A higher precision (P) indicates a lower false positive rate and better model performance.

[0117] Figure 4.4 The middle is a visualization of the abnormal detection effect of high-price low-acceptance default users in the corresponding category based on the experimental set samples and the threshold set above. Figure 4.4 It can be seen that: (1) All abnormal users were accurately identified in the abnormal identification of high-price low-connection default users in categories 1 and 4; (2) There was one normal user identified as an abnormal user (AFP) in category 2 and category 3, and one abnormal user identified as a normal user (ATP) in category 3 of business.

[0118] The on-site audit results ultimately determined that the number of users who violated the high-price, low-connection policy was 25. To verify the accuracy and effectiveness of this research method, we compared it with several other cluster audit models. The comparison details are shown in Table 4.3.

[0119] It can be seen from Table 4.3 that the high-price low-access user detection and identification method of the present invention is significantly better than several other clustering detection methods in all aspects. The accuracy of the audit method of the present invention is as high as 96%, and the accuracy of the density clustering audit model is 92.31%. This shows the superiority of the recognition detection model of the present invention.

[0120] Tab.4.3Is compared with the statistical results of the traditional audit model

[0121]

[0122] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "preferred embodiments," "specific implementations," or "preferred implementations" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0123] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. Therefore, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A method for identifying high-priced, low-receiving users based on PSO optimization clustering and Hausdorff distance discriminant analysis, characterized in that: The following steps are involved: S1. Input the electricity consumption data set of power users and perform data preprocessing; S2. Optimize the K-means clustering algorithm based on the particle swarm optimization algorithm to generate typical electricity consumption trajectories of commercial users; S3. Perform Hausdorff distance discriminant analysis on the typical electricity consumption trajectories of clustered commercial users and ordinary residential users to calculate the similarity between the two. To preliminarily identify a series of users with abnormal electricity usage behavior; S4. Based on a preset threshold, determine whether a series of ordinary residential users with abnormal electricity consumption behaviors initially identified in S3 are high-price low-connection users, and issue an early warning.

2. The method according to claim 1, characterized in that S1 includes: S1.

1. Data cleaning, including outlier processing and missing value recovery; S1.

2. Perform data standardization and Gaussian smoothing, and perform Gaussian smoothing on the standardized load data; the standardization process uses the Z-Score method, which includes calculating the mean and standard deviation at each time point, and normalizing the original load value to obtain the standardized load sequence Z′.

3. The method according to claim 2, characterized in that S1.1 includes: S1.1.

1. Obtain the monthly electricity consumption trajectory set M = {m1, m2, ..., m n }, where m j The monthly electricity consumption trajectory of the jth electricity user constitutes a power load time series; and the user's power load data is collected every 15 minutes; S1.1.

2. Perform outlier processing on the collected power load time series and delete the time series; if the negative value ratio is less than 20%, the negative value will be regarded as missing value; S1.1.

3. Identify and correct the error values ​​in the power load time series. The correction formula is: Where x i ; represents the user's power load data value, σ(X i ) represents vector X i The standard deviation of x is NaN. i ; S1.1.

4. Recover the missing values ​​marked in S1.1.2 using the following formula: Where, mean(X i ) represents vector X i The average value of .

4. The method according to claim 1, wherein S2 specifically includes: S2 includes using the particle swarm optimization algorithm to optimize the K-means clustering algorithm to determine the optimal number of clusters K, the initial cluster center, the number of algorithm iterations, and the number of initialization runs, and to obtain the optimal clustering result with the maximization of the silhouette coefficient as the objective function; specifically, it is divided into: S21, initialize the particle swarm, each particle represents a group of cluster centers; S22, calculate the fitness value of each particle and optimize it with the silhouette coefficient as the objective function; S23, by iteratively updating the position and velocity of the particles, dynamically adjusting the inertia weight to find the optimal cluster center, and stopping the iteration when the fitness curve of the objective function tends to be flat; S24. When the iteration reaches a stable state, the optimal clustering result and the corresponding typical power consumption trajectory are output. The stable state means that the objective function value converges when the number of iterations reaches 100.

5. The method according to claim 1, wherein S3 specifically includes: S31. The typical electricity consumption trajectory of commercial users is represented as a point set A = {a1, a2, ..., a q }, the electricity consumption trajectory of ordinary residential users is represented by point set B n ={b1,b2,...,b q }, where n∈N + Representing each ordinary residential user; Among them, a i ,b j Respectively represent the power consumption of each power consumption trajectory at a certain moment; S32, calculate the one-way similarity distance h(A,B n ): represents the maximum value of the minimum distance between each point of the commercial trajectory and the resident trajectory; And calculate the one-way similarity distance h(B n ,A); represents the maximum value of the minimum distance between each point of the resident trajectory and the commercial trajectory; Among them, d{A(a i ),B n (b j )}=||A(a i )-B n (b j )|| represents a point on the typical electricity consumption trajectory A of a commercial user to the electricity consumption trajectory B of an ordinary residential user n The Euclidean distance of S33: Take the maximum value of the two-way distance as the comprehensive similarity measure; H(A,B n )=max{h(A,B n ),h(B n ,A)} Where H(A,B n ) value is smaller, indicating a higher similarity; output the minimum H(A,B n ) value to S4 for threshold determination.

6. The method according to claim 1, wherein In S4, determining the preset threshold includes: Extract typical electricity consumption trajectories of each commercial user category from the training set samples; Calculate the Hausdorff distance between the electricity consumption trajectory of each user marked as abnormal in the training set and the typical trajectory of the corresponding commercial user category; The maximum value of the Hausdorff distance of abnormal users corresponding to each type of commercial users is taken as the preset threshold ε of this category.