An automated electricity theft detection method and an electricity metering box

By constructing the selection coefficient, the SMOTE algorithm is improved for balancing processing, the problem of unbalanced power stolen data is solved, the recognition accuracy of power stolen detection is improved, and the effective distinction between different power stolen methods is achieved.

CN119782755BActive Publication Date: 2025-07-18广东佰林电气设备厂有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510278947.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-18
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing power stolen detection methods based on machine learning and neural networks have low recognition accuracy due to the imbalance of power stolen data. Traditional SMOTE algorithms have failed to effectively distinguish power stolen data generated by different power stolen methods.

Method used

By calculating the overlapping area distance and the importance of the power stolen data, the selection coefficient is constructed, the SMOTE algorithm is improved for balanced processing, and the random forest classifier is trained to perform power stolen detection.

Benefits of technology

The identification accuracy of the power stolen detection method for different power stolen data generated by power stolen methods is improved, and the synthetic samples are avoided being distributed in boundaries or sparse areas, which enhances the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782755B_ABST
    Figure CN119782755B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power theft detection, and specifically relates to an automated power theft detection method and a power metering box. The method includes: obtaining all power consumption data of each user within a historical detection period to form a power consumption dataset; calculating all typical power theft detection feature indicators of each power consumption data; screening out power theft data from the power consumption dataset using power theft records; constructing a selection coefficient for the power theft data based on the power consumption dataset; improving the method for calculating sample distances in the SMOTE algorithm using the selection coefficient to perform a balancing process on the power consumption dataset; and training a random forest classifier based on the balanced user power consumption samples to detect power theft in the user power consumption data. The purpose of the present application is to improve the recognition accuracy of power theft data in the obtained power consumption data by the subsequent power theft detection method used.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of power theft detection, and specifically relates to an automated power theft detection method and a power metering box. Background Art

[0002] A power metering box is a power device used for measuring electrical energy. By combining an automated power theft detection method with a power metering box, the safety and reliability of the power system can be effectively improved, and the economic losses caused by power theft can be reduced.

[0003] Since the number of power theft users is relatively small compared to all users and the workload of searching for power theft users household by household is large, it is difficult to collect power theft data. This will cause the power consumption data set used in existing power theft detection methods based on machine learning and neural networks to have an imbalance in power theft data, thereby affecting the accuracy of these methods. To solve this imbalance in power theft data, oversampling techniques are usually used to balance the power consumption data set. For example, the power theft detection method, system and device based on oversampling and improved random forest disclosed in publication number CN114610706B use the SMOTE algorithm to oversample the power theft data in the data set used in the power theft detection method based on random forest.

[0004] However, the traditional SMOTE algorithm does not consider that the power theft data in the obtained power consumption data set is likely to have uneven sample distribution due to different power theft methods. This will cause the newly synthesized power theft data samples to be easily distributed at the boundary of positive and negative samples and there will be problems with the distribution of minority samples, making the balanced power consumption data set unable to effectively distinguish the power theft data generated by different power theft methods in the power consumption data set, thereby affecting the recognition accuracy of the power theft detection method based on random forest for the power theft data generated by each power theft method in the power consumption data. Summary of the Invention

[0005] To solve the above technical problems, the present application provides an automated power theft detection method and a power metering box, and the specific technical solutions adopted are as follows:

[0006] In the first aspect, an embodiment of the present application provides an automated power theft detection method, and the method includes the following steps:

[0007] Step 1: Obtain all the power consumption data of each user within a historical detection period and form a power consumption data set; calculate all the typical feature indicators for power theft detection of each power consumption data; and screen out the power theft data from the power consumption data set using power theft records.

[0008] Step 2: Construct a selection coefficient for power theft data based on the power consumption data set; specifically including:

[0009] S1. Utilize all typical feature indicators for electricity theft detection to map all electricity consumption data in the electricity consumption dataset into the data space. Use a support vector machine to perform binary classification on all electricity consumption data in the data space according to whether it is electricity theft data, and output the linear equation of the separation hyperplane.

[0010] S2. Count the number of positive and negative values of the results after substituting all electricity theft data into the linear equation to determine the overlapping set of electricity theft data. And judge whether each electricity theft data in the data space belongs to the overlapping set of electricity theft data, so as to determine the overlapping area distance of each electricity theft data.

[0011] S3. Cluster all electricity theft data in the data space. With any electricity theft data that is not the cluster center in any clustering cluster as the center, set a hypersphere with a preset radius. Record the number of all electricity theft data belonging to the clustering cluster where the electricity theft data is located within the hypersphere as the local electricity theft density of the electricity theft data.

[0012] S4. Utilize the distance between the electricity theft data and the cluster center of its clustering cluster and the local electricity theft density of the electricity theft data to determine the importance degree of the electricity theft characteristics of the electricity theft data.

[0013] S5. Utilize the overlapping area distance and the importance degree of the electricity theft characteristics to determine the selection coefficient of the electricity theft data.

[0014] Step 3: Utilize the selection coefficient to improve the method of calculating the sample distance in the SMOTE algorithm to balance the electricity consumption dataset. And train a random forest classifier based on the balanced user electricity consumption samples to detect electricity theft in user electricity consumption data.

[0015] Preferably, the electricity consumption data includes three-phase current and three-phase voltage; the typical feature indicators for electricity theft detection include voltage unbalance degree, rated voltage deviation degree, current unbalance degree and power factor.

[0016] Preferably, the method of mapping all electricity consumption data in the electricity consumption dataset into the data space by utilizing all typical feature indicators for electricity theft detection is as follows:

[0017] Construct the data space of the electricity consumption dataset with each typical feature indicator for electricity theft detection of the electricity consumption data in the electricity consumption dataset as a dimension, and map all electricity consumption data in the electricity consumption dataset into the data space.

[0018] Preferably, the method for determining the overlapping set of electricity theft data is as follows:

[0019] Substitute the coordinates of each electricity theft data in the data space into the linear equation of the separation hyperplane, and record the result of the equation after substitution as the symbolic distance corresponding to the electricity theft data; respectively obtain the electricity theft data sets composed of all electricity theft data with symbolic distances greater than 0 and less than 0 in the data space, and record the electricity theft data set with the smallest number of elements in the set as the electricity theft data overlapping set.

[0020] Preferably, the method for determining the overlapping region distance of each electricity theft data is as follows:

[0021] When the electricity theft data belongs to the electricity theft data overlapping set, assign the overlapping region distance of the electricity theft data to 0;

[0022] Otherwise, assign it to the absolute value of the symbolic distance corresponding to the electricity theft data.

[0023] Preferably, the method for determining the importance degree of the electricity theft characteristics of the electricity theft data is as follows:

[0024] Take the ratio result of the distance between the electricity theft data and the cluster center of its cluster and the local electricity theft density of the electricity theft data as the importance degree of the electricity theft characteristics of the electricity theft data.

[0025] Preferably, the selection coefficient of the electricity theft data is determined by the product result of the overlapping region distance and the importance degree of the electricity theft characteristics of the electricity theft data.

[0026] Preferably, the method of using the selection coefficient to improve the method of calculating the sample distance in the SMOTE algorithm to balance the electricity consumption data set includes:

[0027] Use the SMOTE algorithm to perform interpolation processing on each cluster in the data space respectively to generate the interpolation data set of each cluster;

[0028] Take the union of the interpolation data sets of all clusters and the electricity consumption data set as the balanced electricity consumption data set after balancing the electricity consumption data set;

[0029] Among them, when using the SMOTE algorithm to perform interpolation processing on each cluster in the data space respectively, take the cluster center of the cluster as the root sample in the SMOTE algorithm, and replace the Euclidean distance between the electricity theft data in the cluster and the root sample with the reciprocal of the selection coefficient of the electricity theft data.

[0030] Preferably, based on the balanced user electricity consumption samples, train a random forest classifier to detect electricity theft in user electricity consumption data, including:

[0031] Divide the balanced electricity consumption data set into a training set E1 and a test set E2;

[0032] Construct a random forest classifier, and use the training set E1 and the test set E2 as the training set and the test set of the random forest classifier respectively;

[0033] Substitute the user electricity consumption data collected in real time by the power metering box into the trained random forest classifier to output the electricity theft detection result of the user.

[0034] In a second aspect, another embodiment of the present application provides a power metering box, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned automatic electricity theft detection method are implemented.

[0035] The embodiments of the present application at least have the following beneficial effects:

[0036] The present application obtains the distance of the index overlapping area, which can accurately evaluate the distance between the overlapping areas between the electricity theft data in the data space of the obtained electricity consumption dataset and the electricity theft data and the normal electricity consumption data respectively; the present application obtains the importance of the electricity theft feature index, which can accurately evaluate whether the electricity theft data in the data space of the electricity consumption dataset contains special data distribution features that are difficult to be recognized by subsequent electricity theft detection methods in the electricity theft data distribution features corresponding to its corresponding electricity theft method; the present application constructs a selection coefficient by combining the overlapping area distance and the electricity theft feature importance to select the nearest neighbor samples of the root samples in the SMOTE algorithm, which can effectively avoid the situation that the newly synthesized electricity theft data samples are distributed at the positive and negative sample boundaries, and avoid the situation that the newly synthesized electricity theft data samples are difficult to be distributed in the sparse area containing its special data distribution features in the electricity theft data area where a certain electricity theft method in the obtained electricity consumption dataset is located, thereby improving the recognition accuracy of the subsequent electricity theft detection method for the electricity theft data in the obtained electricity consumption data. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a flowchart of the steps of an automatic electricity theft detection method provided by an embodiment of the present application;

[0039] Figure 2 It is a flowchart of the construction process of the selection coefficient of the electricity theft data provided by an embodiment of the present application. Detailed Embodiments

[0040] To further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following provides a detailed description of a method for automatic electricity theft detection and an electricity metering box proposed according to this application, including their specific implementation manners, structures, features, and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0042] The following specifically describes the specific solutions of a method for automatic electricity theft detection and an electricity metering box provided by this application in combination with the accompanying drawings.

[0043] Please refer to Figure 1 , which shows a step flow chart of a method for automatic electricity theft detection provided by an embodiment of this application. The method includes the following steps:

[0044] Step 1: Obtain all electricity consumption data of each user within a historical detection period and form an electricity consumption dataset; calculate all typical electricity theft detection characteristic indexes of each electricity consumption data; and screen out electricity theft data from the electricity consumption dataset using electricity theft records.

[0045] This application uses the power load management terminal in the electricity metering box to obtain the electricity consumption data of each user within a historical detection period. The electricity consumption data includes three-phase current and three-phase voltage. The detection period and sampling time interval of the electricity consumption data are set to 1 month and 15 minutes respectively, and can be set by the implementer himself.

[0046] Preprocess the obtained electricity consumption data. The preprocessing includes using the KNN interpolation method to fill in the missing values of the electricity consumption data, using the Grubbs criterion to process the outliers in the electricity consumption data, and using the Min-Max normalization method to normalize the electricity consumption data. The methods used in the preprocessing process are all well-known technologies, and the specific process will not be elaborated.

[0047] Denote the dataset constructed from all preprocessed electricity consumption data as electricity consumption dataset A, and mark the electricity consumption data in electricity consumption dataset A as electricity theft data and normal electricity consumption data according to whether there is an electricity theft record.

[0048] Furthermore, calculate all electricity theft detection characteristic indexes of each electricity consumption data in electricity consumption dataset A. The typical electricity theft detection characteristic indexes in this embodiment include voltage unbalance degree, rated voltage deviation degree, current unbalance degree, and power factor. The calculation of the typical electricity theft detection characteristic indexes is a well-known technology, and the specific process will not be elaborated.

[0049] Step 2: Construct the selection coefficient of power theft data based on the power consumption dataset.

[0050] In this application, the flow chart of the construction process of the selection coefficient of power theft data is as shown in the appendix Figure 2 and specifically includes:

[0051] S1. Utilize all typical feature indicators of power theft detection to map all power consumption data in the power consumption dataset into the data space; use the support vector machine to perform binary classification on all power consumption data in the data space according to whether it is power theft data, and output the linear equation of the separation hyperplane.

[0052] In the SMOTE algorithm, the synthesis of new samples depends on the selection of root samples and neighbor samples. The traditional SMOTE algorithm takes all minority samples as root samples and uses the Euclidean distance to take the K-nearest neighbor samples of the root samples as neighbor samples. Although the existing method uses the cluster center of the clustering cluster as the root sample to alleviate the situation where the newly synthesized power theft data samples are distributed at the boundary of positive and negative samples in the SMOTE algorithm, it does not consider that the selection of neighbor samples will also cause the newly synthesized power theft data samples to be distributed at the boundary of positive and negative samples.

[0053] Therefore, to avoid the situation where the newly synthesized power theft data samples are distributed at the boundary of positive and negative samples, the following processing is carried out.

[0054] S11. Specifically, construct the data space V of the power consumption dataset A with each power theft detection feature indicator of the power consumption data in the power consumption dataset A as a dimension, and map all power consumption data in the power consumption dataset A into the data space V.

[0055] S12. Further, use the support vector machine (SVM) algorithm to perform binary classification on all power consumption data in the data space V, where the labels of the power theft data and normal power consumption data in the power consumption data are set to 1 and 0 respectively, and output the linear equation P of the separation hyperplane. The SVM algorithm is a well-known technology, and the specific process will not be elaborated.

[0056] S2. Count the number of positive and negative values of the results after all power theft data are brought into the linear equation to determine the overlapping set of power theft data; and judge whether each power theft data in the data space belongs to the overlapping set of power theft data, so as to determine the overlapping region distance of each power theft data.

[0057] Since there are usually large differences in the data distribution characteristics between power theft data and normal power consumption data, only a small number of power theft data have relatively similar power consumption data distribution characteristics to normal power consumption data, so that most of the power theft data in the data space V are on one side of the separation hyperplane, and a small number of power theft data are on the other side.

[0058] Therefore, in order to accurately evaluate the distance between the electricity theft data in the data space V and the overlapping regions between the electricity theft data and the normal electricity consumption data, and to avoid the SMOTE algorithm synthesizing too many new samples of electricity theft data in the overlapping regions, the following processing is carried out.

[0059] S21. Specifically, substitute the coordinates of each electricity theft data in the data space V into the linear equation P of the separation hyperplane, and record the result of the substituted equation as the signed distance of each electricity theft data. The smaller the absolute value of the signed distance, the smaller the distance from the electricity theft data to the separation hyperplane. When the signed distance is greater than 0, the electricity theft data is on the positive normal vector direction side of the separation hyperplane; when the signed distance is less than 0, the electricity theft data is on the negative normal vector direction side of the separation hyperplane; when the signed distance is equal to 0, the electricity theft data is located on the separation hyperplane.

[0060] Respectively obtain the electricity theft data sets composed of all electricity theft data with signed distances greater than 0 and less than 0 in the data space V, and record the electricity theft data set with the smallest number of set elements as the electricity theft data overlapping set B, which is used to represent the set composed of all electricity theft data in the overlapping regions between the electricity theft data and the normal electricity consumption data in all electricity theft data in the data space V.

[0061] S22. Further, taking the electricity theft data a in the data space V as an example, obtain the overlapping region distance D(a) of the electricity theft data a, which is used to represent the distance between the electricity theft data a and the overlapping regions between the electricity theft data and the normal electricity consumption data in the data space V:

[0062] , where D1(a) represents the signed distance of the electricity theft data a; B represents the electricity theft data overlapping set.

[0063] It should be understood that the smaller the distance between the electricity theft data a and the overlapping regions between the electricity theft data and the normal electricity consumption data in the data space V, that is, the smaller the overlapping region distance D(a), the less the electricity theft data a can be used as the neighbor sample of the root sample in the SMOTE algorithm to avoid the SMOTE algorithm synthesizing too many new samples of electricity theft data in the overlapping regions.

[0064] S3. Cluster all the electricity theft data in the data space; taking any electricity theft data that is not the cluster center in any cluster as the center, set a hypersphere with a preset radius; record the number of all electricity theft data belonging to the cluster of this electricity theft data within the hypersphere as the local electricity theft density of this electricity theft data.

[0065] Secondly, since the collected electricity theft data will have different data distribution characteristics due to the different electricity theft methods used by the electricity theft users, the electricity theft data in the acquired electricity consumption data set will easily have uneven sample distribution due to the different electricity theft methods. This will cause the new electricity theft data samples synthesized by the SMOTE algorithm using nearest neighbor samples to be mainly distributed in the dense area where the cluster center of the cluster is located, but it is difficult to synthesize more new electricity theft data samples in the sparse area containing the special data distribution characteristics of a certain electricity theft method in the electricity consumption data set. As a result, it is difficult for the subsequent electricity theft detection method based on random forest to effectively extract the special data distribution characteristics of the electricity theft method in its sparse electricity theft data area, thereby affecting the recognition accuracy of the electricity theft detection method for the electricity theft data generated by the electricity theft method.

[0066] Therefore, in order to avoid the situation where the synthesized new electricity theft data samples are difficult to be distributed in the sparse area containing the special data distribution characteristics of a certain electricity theft method in the acquired electricity usage data, the following processing is performed.

[0067] S31. Specifically, in this embodiment, the neighbor propagation clustering algorithm (AP) is used to classify all electricity theft data in the data space V to obtain multiple clusters in the data space V and the cluster center of each cluster. Each cluster corresponds to an electricity theft data area in the data space V corresponding to an electricity theft method in the applied electricity data set A. The damping factor, preference parameter, number of clustering iterations, and maximum number of iterations in the AP clustering algorithm all take default values. The AP clustering algorithm is a well-known technology, and the specific process will not be repeated here.

[0068] S32. Further, taking the j-th electricity theft data C(i,j) in the i-th cluster C(i) in the data space V (excluding the electricity consumption data corresponding to the cluster center in the cluster C(i)) as an example, a hypersphere g(i,j) is set with the electricity theft data C(i,j) as the center, and the radius of the hypersphere is set to 3 unit lengths, which is a preset radius in this embodiment and can be set by the implementer.

[0069] The number of all electricity theft data belonging to the cluster C(i) within the hypersphere g(i,j) is recorded as the local electricity theft density of the electricity theft data C(i,j), which is used to characterize the number of electricity theft data in the cluster C(i,j) in the local space where the electricity theft data C(i,j) is located.

[0070] S4, determining the importance of the electricity theft feature of the electricity theft data by using the distance between the electricity theft data and the cluster center of the cluster to which it belongs and the local electricity theft density of the electricity theft data.

[0071] Further, the importance U(i,j) of the electricity theft feature of the electricity theft data C(i,j) is obtained by using the distance between the electricity theft data C(i,j) and the cluster center of its corresponding cluster C(i) and the local electricity theft density of the electricity theft data C(i). It is used to characterize the possibility that the electricity theft data distribution feature carried in the local area where the electricity theft data C(i,j) is located belongs to the special data distribution feature that is difficult to be recognized by subsequent electricity theft detection methods among the electricity theft data distribution features corresponding to the electricity theft method of the cluster C(i):

[0072] , where d(i,j) represents the Euclidean distance between the electricity theft data C(i,j) and the cluster center of its corresponding cluster C(i); ρ(i,j) represents the local electricity theft density of the electricity theft data C(i,j).

[0073] It should be understood that the farther the distance between the electricity theft data C(i,j) and the regional center of its corresponding cluster C(i), that is, the larger d(i,j), the greater the difference between the electricity theft data distribution feature carried in the local area where the electricity theft data C(i,j) is located and the main electricity theft data distribution feature corresponding to the electricity theft method of the cluster C(i), and the fewer the number of electricity theft data within its cluster contained in the local space where the electricity theft data C(i,j) is located, that is the smaller ρ(i,j), the more likely it is that the local area where the electricity theft data C(i,j) is located carries the special data distribution feature that is difficult to be recognized by subsequent electricity theft detection methods among the electricity theft data distribution features corresponding to the electricity theft method of the cluster C(i), that is, the greater the importance U(i,j) of the electricity theft feature. To reduce the situation that the synthesized new sample distribution is difficult to be distributed in the sparse area containing its special data distribution feature in the electricity theft data area where a certain electricity theft method is located in the obtained electricity consumption data due to the improper selection of the nearest neighbor samples in the SMOTE algorithm, the electricity theft data C(i,j) should be used as the nearest neighbor sample of the root sample in the SMOTE algorithm more.

[0074] S5. Use the overlapping region distance and the importance of the electricity theft feature to determine the selection coefficient of the electricity theft data.

[0075] Further, the selection coefficient W(i,j) of the electricity theft data C(i,j) is obtained by using the overlapping region distance D(i,j) and the importance U(i,j) of the electricity theft feature of the electricity theft data C(i,j). It is used to characterize the possibility that the electricity theft data C(i,j) is selected as the nearest neighbor sample when using the SMOTE algorithm to synthesize new electricity theft data samples in the cluster C(i): W(i,j) = D(i,j) * U(i,j), where D(i,j) and U(i,j) respectively represent the overlapping region distance and the importance of the electricity theft feature of the electricity theft data C(i,j).

[0076] Step 3: Improve the method for calculating sample distance in the SMOTE algorithm using the selection coefficient to balance the electricity consumption dataset; and train a random forest classifier based on the balanced user electricity consumption samples to detect electricity theft in user electricity data.

[0077] Further, taking the clustering cluster C(i) as an example, the SMOTE algorithm is used to perform interpolation processing on the clustering cluster C(i) to generate an interpolation dataset , where the cluster center of the clustering cluster C(i) is used as the root sample in the SMOTE algorithm, and the Euclidean distance between the electricity theft data and the root sample in the clustering cluster C(i) is replaced by the reciprocal of the selection coefficient of the electricity theft data. The SMOTE algorithm is a well-known technology, and the specific process will not be elaborated here.

[0078] Using the same method as the interpolation dataset Interpolated datasets for each clustering cluster in the data space V are generated respectively. The union of all the obtained interpolated datasets recombined with the electricity consumption dataset A is used as the balanced electricity consumption dataset E after balancing the electricity consumption dataset A, and the balanced electricity consumption dataset is divided into a training set E1 and a test set E2.

[0079] Further, a random forest classifier is constructed, where the training set E1 and the test set E2 are used as the training set and the test set of the random forest classifier respectively. The specific training process is a well-known technology and will not be elaborated here.

[0080] Substitute the electricity consumption data of a certain user collected in real-time by the power metering box into the trained random forest classifier, and output the electricity theft detection result of this user to complete the automated electricity theft detection.

[0081] An embodiment of the present application also proposes a power metering box, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned automated electricity theft detection methods are implemented. Since an automated electricity theft detection method is described in detail above, it will not be elaborated here.

[0082] The embodiments in the present application are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are to illustrate the differences from other embodiments.

[0083] It should be noted that, unless otherwise specified or limited, terms such as "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the article or device including the said element. In addition, the term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0084] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not invented by the present application.

[0085] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. An automated electricity theft detection method, characterized in that, The method includes the following steps: Step 1: Obtain all the power consumption data of each user within a historical detection period and form a power consumption data set; calculate all types of typical feature indexes for power theft detection of each power consumption data; and screen out the power theft data from the power consumption data set by using the power theft records. Step 2: Construct a selection coefficient for the power theft data based on the power consumption data set. Specifically, it includes: S1. Use all types of typical feature indexes for power theft detection to map all the power consumption data in the power consumption data set into the data space; use a support vector machine to perform binary classification on all the power consumption data in the data space according to whether it is power theft data, and output the linear equation of the separation hyperplane. S2. Count the number of positive and negative values of the results after all the power theft data are brought into the linear equation to determine the overlapping set of power theft data; and judge whether each power theft data in the data space belongs to the overlapping set of power theft data, so as to determine the overlapping region distance of each power theft data. S3. Cluster all the power theft data in the data space; take any power theft data that is not the cluster center in any cluster as the center, and set a hypersphere with a preset radius; record the number of all the power theft data belonging to the cluster where the power theft data is located within the hypersphere as the local power theft density of the power theft data. S4. Use the distance between the power theft data and the cluster center of the cluster where it is located and the local power theft density of the power theft data to determine the importance degree of the power theft characteristics of the power theft data. S5. Use the overlapping region distance and the importance degree of the power theft characteristics to determine the selection coefficient of the power theft data. Step 3: Use the selection coefficient to improve the method for calculating the sample distance in the SMOTE algorithm to balance the power consumption data set; and train a random forest classifier based on the balanced user power consumption samples to detect power theft in the user power consumption data. The method of using all types of typical feature indexes for power theft detection to map all the power consumption data in the power consumption data set into the data space is: Take each type of typical feature index for power theft detection of the power consumption data in the power consumption data set as a dimension respectively, construct the data space of the power consumption data set, and map all the power consumption data in the power consumption data set into the data space. The method of using the selection coefficient to improve the method for calculating the sample distance in the SMOTE algorithm to balance the power consumption data set includes: Use the SMOTE algorithm to perform interpolation processing on each cluster in the data space respectively to generate the interpolation data set of each cluster. Take the union of the interpolation data sets of all the clusters and the power consumption data set as the balanced power consumption data set after balancing the power consumption data set. Among them, when using the SMOTE algorithm to perform interpolation processing on each cluster in the data space respectively, take the cluster center of the cluster as the root sample in the SMOTE algorithm, and replace the Euclidean distance between the power theft data in the cluster and the root sample with the reciprocal of the selection coefficient of the power theft data.

2. The automated electricity theft detection method according to claim 1, characterized in that, The power consumption data includes three-phase current and three-phase voltage; the typical feature indexes for power theft detection include voltage unbalance degree, rated voltage deviation degree, current unbalance degree and power factor.

3. An automated electricity theft detection method according to claim 1, characterized in that The method for determining the overlapping set of power theft data is: Substitute the coordinates of each electricity theft data in the data space into the linear equation of the separating hyperplane, and record the result of the equation after substitution as the signed distance of the corresponding electricity theft data; respectively obtain the electricity theft data sets composed of all electricity theft data with signed distances greater than 0 and less than 0 in the data space, and record the electricity theft data set with the smallest number of elements in the set as the electricity theft data overlapping set.

4. The automated electricity theft detection method according to claim 3, wherein The method for determining the overlapping area distance of each electricity theft data is as follows: When the electricity theft data belongs to the electricity theft data overlapping set, assign the overlapping area distance of the electricity theft data to 0; Otherwise, assign it to the absolute value of the signed distance of the corresponding electricity theft data.

5. An automated electricity theft detection method according to claim 1, characterized in that The method for determining the importance degree of the electricity theft feature of the electricity theft data is as follows: Take the ratio of the distance between the electricity theft data and the cluster center of its cluster and the local electricity theft density of the electricity theft data as the importance degree of the electricity theft feature of the electricity theft data.

6. An automated electricity theft detection method according to claim 1, characterized in that, The selection coefficient of the electricity theft data is determined by the product of the overlapping area distance and the importance degree of the electricity theft feature of the electricity theft data.

7. An automated electricity theft detection method according to claim 1, characterized in that, And based on the balanced user electricity consumption samples, train a random forest classifier to detect electricity theft in user electricity consumption data, including: Divide the balanced electricity consumption data set into a training set E1 and a test set E2; Construct a random forest classifier, and use the training set E1 and the test set E2 as the training set and the test set of the random forest classifier respectively; Substitute the user electricity consumption data collected in real time by the power metering box into the trained random forest classifier, and output the electricity theft detection result of the user.

8. An electric energy metering box, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it realizes the steps of an automatic electricity theft detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Electricity theft detection method, system and device based on oversampling and improved random forest

    CN114610706B

  • Electricity stealing detection method, system and device based on oversampling and improved random forest

    CN114610706A