A detection method for identifying electricity theft based on cumulative probability distribution
Through the identification method based on cumulative probability distribution, combined with clustering and the longest common subsequence algorithm, the user's electricity consumption data is analyzed, and the problem of low efficiency in identifying and collecting abnormal electricity consumption data in the prior art is solved, and accurate identification and timely discovery of electricity theft behavior is achieved.
Patent Information
- Application Number
- CN202211697313.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-12-28
AI Technical Summary
The prior art is inefficient when identifying and collecting abnormal electricity use data and cannot guarantee timeliness and accuracy. Especially when the user information data is large, it is difficult to effectively identify power theft behavior.
The identification method based on the cumulative probability distribution is adopted to establish the cumulative probability distribution function through the user's daily power law, and combine the clustering algorithm and the longest common subsequence algorithm to analyze and judge the user's electricity consumption data to identify theft of electricity.
It improves the accuracy and efficiency of abnormal electricity use for special-transform users and low-voltage platform users, can detect power theft in a timely manner, and reduces manpower and material consumption.
Smart Images

Figure CN115859208B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of abnormal electricity consumption of special transformer users and low-voltage substation area users, and particularly to a detection method for identifying electricity theft based on cumulative probability distribution. Background Art
[0002] In addition to conventional data, the electricity consumption information acquisition system also contains abnormal electricity consumption data. When the special metering device discovers abnormal electricity consumption behavior, the characteristic parameters such as voltage, current, and power in the device will have a large gap from the normal data, which is abnormal electricity consumption data. At present, power companies mainly collect abnormal electricity consumption data through data mining monitoring and manual sampling methods. However, there are numerous power users, and using the above methods to collect abnormal electricity consumption data will consume a large amount of manpower and material resources. Moreover, due to the large amount of user information data, the screening efficiency of abnormal electricity consumption information is low, and it cannot ensure the timeliness and accuracy of the collection of abnormal electricity consumption data. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a detection method for identifying electricity theft based on cumulative probability distribution. Based on the daily power law of users, the daily power distribution result of users is given through the daily power determination algorithm of cumulative probability distribution, and then different power distribution results are obtained by clustering the determination results through a clustering algorithm. Further, it can be inferred whether there is an abnormality in the metering device or user electricity theft.
[0004] To achieve the above purpose, the present invention adopts the following technical scheme: A detection method for identifying electricity theft based on cumulative probability distribution, comprising the following steps:
[0005] Step 1: Obtain the electricity consumption data of users through the electricity consumption information acquisition system to obtain the power values of users;
[0006] Step 2: Establish the probability distribution range of the user's daily power scatter points;
[0007] Step 3: Add the probability values of the active power points P i corresponding to different time intervals in different power intervals to obtain the corresponding cumulative probability;
[0008] Step 4: Establish a corresponding fuzzy matrix according to the cumulative probability function value of the user's daily power distribution and the maximum and minimum index of the daily power;
[0009] Step 5: Calculate the density of all power data in the power dataset;
[0010] Step 6: Calculate the average distance density value of the electricity consumption power data;
[0011] Step 7: Classify different cluster distances C X If w XyFor the applied electric power data density ρ Xmax is the maximum value, it is defined as maxd xy ; if w Xy corresponding to the electric power data density ρ Xma x is not the maximum value, it is defined as mind xy , that is, C X ∈[mind xy , maxd xy ;
[0012] Step Eight: According to the power data set given in Step Four, repeat Steps Five to Six, and end the clustering when the clustering data set is an empty set;
[0013] In Step Nine: Update the data center;
[0014] Step Ten: According to the set minimum threshold φ xy , when the clustering center is within the threshold σ x operating range and changes little, that is, φ xy ≤σ x , this center is regarded as the clustering center, and the clustering center sequence C'={C 1 , C 2 ,..., C x} is obtained;
[0015] Step Eleven: Use the Longest Common Subsequence Algorithm LCSS to calculate the similarity of each clustering center;
[0016] Step Twelve: Measure the similarity of the curves by calculating the number of trajectory points between the matching points of two clustering centers;
[0017] Step Thirteen: Calculate the matching rate between each pair of the clustering center sequence C'={C 1 , C 2 ,..., C x} respectively, accumulate the matching degrees between every two sets C i and C j , and set a judgment threshold η through experience. When the accumulated result is lower than η, that is, ∑ρ(C i , C j )<η, it is considered that the user's electricity consumption is abnormal.
[0018] In a preferred embodiment, after the electricity consumption data in the first step is subjected to data cleaning and filling of missing values, the power values of 96 measurement points for each day of the user are obtained, that is:
[0019] P={P 1 , P 2 ,..., P 96} (1).
[0020] In a preferred embodiment, in step two, the daily active power P is equally divided into M intervals. For a certain small interval (P i , P i+1 ), assuming that the width of the active power in each interval is H, then P i+1 - P i = H, that is, H is:
[0021]
[0022] Since the width of the small interval of active power is much smaller than the daily power magnitude, each power within the interval can be regarded as the same point value P i .
[0023] In a preferred embodiment, in step three, the cumulative probability is used to obtain the cumulative distribution curve of each interval through the distribution cumulative probability function. The expression of the distribution cumulative probability function is as follows:
[0024]
[0025]
[0026] Where α is the scale parameter and β is the shape parameter.
[0027] In a preferred embodiment, in step four, according to the user's daily power distribution cumulative probability function value and the daily power maximum and minimum value indicators, assuming that the number of clusters to be divided is X, and the number of indicators considered for each cluster center is Y, the corresponding fuzzy matrix is established as:
[0028]
[0029] In the formula, a 1 , a 2 , …, a X are the number of clusters, a X is the number of indicators of the i-th cluster, w x1 , w x2 , …, w xY are the Y indicator coefficients corresponding to a i , and Y = M + 2.
[0030] In a preferred embodiment, in step five: Calculate the density ρ X (w Xy , D Xav ) of all power data in the power dataset. ρ X is defined as the sample area with w Xy as the center and D Xav as the radius. Using the maximum value of the power data density value ρ XmaxAs the first clustering center, the expression for the corresponding radius is:
[0031]
[0032] where x, z ∈ X.
[0033] In a preferred embodiment, in step six: calculate the average distance density value d of the power consumption data X , specifically:
[0034]
[0035] In a preferred embodiment,
[0036] In step nine: update the data center, with ∑mind' xy being the smallest as the new clustering center. According to the new clustering center, use K-medoids clustering to cluster the probability distribution curve. The specific steps are as follows:
[0037] Calculate the distance D(d xy ′ i ′ nx ) from the power consumption data to ∑mind'
[0038] D(d′ inx ) = angmin||d′ in xy - ∑mind′ xy || 2 (8)
[0039] where: d′ in xy is the data value in the x-th group of clusters; ∑mind′ xy is the distance of the clustering center;
[0040] Select the average value of d′ in xy as the new clustering center and repeat the above formula.
[0041] In a preferred embodiment, in step eleven, use the longest common subsequence algorithm LCSS to calculate the similarity of each clustering center. Now, taking two clustering center curves as an example, let the sequence point sets of the two clustering center curves C 1 and C 2 be C 1 = {a 1 , a 2 ,..., a Y} and C 2 = {b 1 , b 2 ,..., b Y}, then the length of the longest common subsequence is:
[0042]
[0043] where t = 1, 2, 3, …, Y; i = 1, 2, 3, …, Y; γ is the set distance threshold; LCSS(a t , b i ) is the length of the longest common subsequence of sequence C 1 before the trajectory point a t and sequence C 2 before the trajectory point b i ; dist(a t , b t ) is the distance between the a 1 -th point in the C t trajectory and the b 2 -th point in the C i sequence.
[0044] In a preferred embodiment, in step twelve, the similarity of the curves is measured by calculating the number of trajectory points between two clustering center coincidence points that meet the distance threshold; the calculation method of the matching rate ρ of two sets of sequence points C 1 , C 2 is as follows:
[0045]
[0046] where ρ(C 1 , C 2 ) ∈ [0, 1]; and the larger the value of ρ(C 1 , C 2 ), the more similar the C 1 sequence is to the C 2 sequence.
[0047] Compared with the prior art, the present invention has the following beneficial effects: Through the present invention, a detection method for identifying equal-ratio electricity theft based on cumulative probability distribution can be realized. By analyzing the historical power data of a user over a period of time, the probability distribution curve of the daily power is obtained through the cumulative probability function of the cumulative probability, and then a set of clustering center sequences is obtained through the clustering method, and the longest common subsequence algorithm is used to judge the abnormal electricity consumption of the user. Considering that the daily load switching of special transformer users is relatively regular, this method can greatly improve the accuracy of the probability of identifying abnormal conditions of special transformer users. Similarly, this method can also identify the equal-ratio electricity theft situation of low-voltage substation area users. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the implementation flowchart of the preferred embodiment of the present invention;
[0049] Figure 2User power scatter plot of the preferred embodiment of the present invention, where (a) is the normal user power scatter plot and (b) is the power theft user power scatter plot;
[0050] Figure 3 User cumulative distribution curve of the preferred embodiment of the present invention, where (a) is the normal user cumulative distribution curve and (b) is the power theft user cumulative distribution curve;
[0051] Figure 4 Schematic diagram of the clustering result of the user cumulative distribution curve of the preferred embodiment of the present invention, where (a) is the clustering result of the normal user cumulative distribution curve and (b) is the clustering result of the power theft user cumulative distribution curve. Detailed implementation manners
[0052] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0053] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0054] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0055] A detection method for power theft identification based on cumulative probability distribution, referring to Figure 1 , includes the following steps:
[0056] Step 1: Obtain the power consumption data of users through the power consumption information collection system. After data cleaning and filling the missing values of the data, the power values of 96 measurement points per day for each user are obtained, that is:
[0057] P = {P 1 , P 2 ,..., P 96} (1)
[0058] Step 2: Establish the daily power scatter probability distribution range P of users. The daily active power P is equally divided into M intervals. For a certain small interval (P i , P i+1 ), assuming the width of the active power of each interval is H, then P i+1 - P i = H, that is, H is:
[0059]
[0060] Since the width of the active power small interval is much smaller than the daily power magnitude, each power within the interval can be regarded as the same point value P i ;
[0061] Step 3: Add the probabilities of the active power point values P i corresponding to different time intervals in different power intervals to obtain the corresponding cumulative probability, and obtain the cumulative distribution curve of each interval through the cumulative distribution probability function. The expression of the cumulative distribution probability function is as follows:
[0062]
[0063]
[0064] where α is the scale parameter and β is the shape parameter.
[0065] Step 4: According to indicators such as the cumulative distribution probability function value of the user's daily power distribution, the maximum and minimum values of the daily power, etc., assuming that the number of clusters to be divided is X, and the number of indicators considered for each cluster center is Y, establish the corresponding fuzzy matrix as:
[0066]
[0067] In the formula, a 1 , a 2 , …, a X are the number of clusters, a X is the number of indicators of the i-th cluster, w x1 , w x2 , …, w xY are the Y indicator coefficients corresponding to a i , and Y = M + 2.
[0068] Step 5: Calculate the density ρ X (w Xy , D Xav ) of all power data in the power dataset. ρ X is defined as the sample region centered on w Xy with D Xav as the radius. Take the maximum value ρ Xmax of the power data density as the first cluster center, and the expression of the corresponding radius is:
[0069]
[0070] where x, z ∈ X.
[0071] Step 6: Calculate the average distance density value d of the power consumption data X , specifically:
[0072]
[0073] Step 7: Classify different cluster distances C X . If w Xy corresponds to the power consumption data density ρ Xmax is the maximum value, it is defined as maxd xy ; if w Xy corresponds to the power consumption data density ρ Xmax is not the maximum value, it is defined as mind xy , that is, C X ∈[mind xy , maxd xy .
[0074] Step 8: According to the power data set given in Step 4, repeat Steps 5 - 6, and end the clustering when the clustering data set is an empty set.
[0075] Step 9: Update the data center. Take ∑mind' xy as the smallest as the new clustering center. According to the new clustering center, use K-medoids clustering to cluster the probability distribution curve. The specific steps are as follows:
[0076] Calculate the distance D(d′ xy ) from the power consumption data to ∑mind' inx . The specific expression is as follows:
[0077] D(d′ inx )=angmin||d′ in xy - ∑mind′ xy || 2 (8)
[0078] In the formula: d′ in xy is the data value in the x-th group of clusters; ∑mind′ xy is the clustering center distance.
[0079] Select the average value in xy of d′ as the new clustering center and repeat the above formula.
[0080] Step 10: According to the set minimum threshold φ xy , when the clustering center is within the threshold σ x operating range and changes little, that is, φ xy ≤σ x , regard this center as the clustering center to obtain the clustering center sequence C'={C1 ,C 2 ,...,C x}.
[0081] Step 11: Use the longest common subsequence algorithm (LCSS) to calculate the similarity of each cluster center. Now take two cluster center curves as an example. Let the two cluster center curves C 1 and C 2 The sequence point set is C 1 ={a 1 ,a 2 ,...,a Y} and C 2 = {b 1 ,b 2 ,...,b Y}, then the length of the longest common subsequence is:
[0082]
[0083] Where t = 1, 2, 3, ..., Y; i = 1, 2, 3, ..., Y; γ is the set distance threshold; LCSS (a t , b i ) is the sequence C 1 At trajectory point a t and sequence C 2 At trajectory point b i The maximum common subsequence length before; dist(a t , b t ) is C 1 The first t Points and C 2 The bth i The distance between points.
[0084] Step 12: Measure the similarity of the curves by calculating the number of trajectory points whose two cluster centers meet the distance threshold between points. 1 , C 2 The calculation method of the matching rate ρ is:
[0085]
[0086] Where ρ(C 1 , C 2 )∈[0,1];and ρ(C 1 , C 2 ) value is larger, indicating that C 1 Sequence and C 2 The more similar the sequences are.
[0087] Step 13: Calculate the cluster center sequence C' = {C 1 ,C2 ,..., C x The matching rate between each pair, accumulate the matching degrees between every two sets C i and C j . Set a judgment threshold η through experience. When the accumulated result is lower than η, that is, when ∑ρ(C i , C j ) < η, it is considered that the user's power consumption is abnormal.
[0088] Specifically,
[0089] (1) Obtain the user's power consumption data through the power consumption information collection system, and get the power values of 96 measurement points for each day of the user. As Figure 2 shown, in the figure, (a) is a normal user and (b) is a power theft user. The abscissa is the collection days of 50 days, and the ordinate is the power value corresponding to 96 measurement points per day.
[0090] (2) Establish the probability distribution range P of the user's daily power scatter points, add up the probability values of the active power points P i corresponding to different time intervals in different power intervals to obtain the corresponding cumulative probability, and obtain the cumulative distribution curve of each interval through the distribution cumulative probability function. As Figure 3 shown.
[0091] (3) According to indicators such as the cumulative probability function value of the user's daily power distribution, the maximum and minimum values of the daily power, assuming that the number of clusters to be divided is 2, use K-medoids clustering to cluster the probability distribution curve. The clustering result is as Figure 4 shown.
[0092] Use the longest common subsequence algorithm to calculate the similarity of each cluster center, and measure the similarity of the curve by calculating the number of track points of the distance threshold between two cluster center coincidence points. Set a judgment threshold through experience to further judge whether the user's power consumption data is abnormal.
Claims
1. A detection method for identifying electricity theft based on cumulative probability distribution, characterized in that it includes the following steps: Step 1: Obtain the electricity consumption data of users through the electricity information collection system to obtain the power values of users; Step 2: Establish the probability distribution range of user daily power scatter points; Step 3: Add the probability values of the active power points P within different time intervals corresponding to different power intervals to obtain the corresponding cumulative probability; i Step 4: Establish a corresponding fuzzy matrix according to the cumulative probability function value of the user daily power distribution and the maximum and minimum index of daily power; Step 5: Calculate the density of all power data in the power dataset; Step 6: Calculate the average distance density value of the electricity consumption power data; Step Seven: Classify different cluster distances C X If w Xy for the applied electric power data density ρ Xmax is the maximum value, it is defined as maxd xy ; if w Xy for the corresponding electric power data density ρ Xmax is not the maximum value, it is defined as mind xy , that is, C X ∈[mind xy , maxd xy ; Step 8: According to the power dataset given in Step 4, repeat Steps 5 - 6, and end the clustering when the clustering dataset is an empty set; Step 9: Update the data center; Step Ten: According to the set minimum threshold φ xy , when the cluster center is within the running range of the threshold σ x and changes little, that is, when φ xy ≤σ x , this center is regarded as the cluster center, and the cluster center sequence C' = {C 1 , C 2 ,..., C x} Step 11: Use the longest common subsequence algorithm LCSS to calculate the similarity of each clustering center; Step 12: Measure the similarity of the curves by calculating the number of trajectory points of the distance threshold between the coincidence points of the two clustering centers; Step 13: Calculate the matching rates between each pair of the clustering center sequences C' = {C 1 , C 2 ,..., C x}, and accumulate the matching degrees between every two sets C i and C j . Set a judgment threshold η through experience. When the accumulated result is lower than η, that is, when ∑ρ(C i , C j ) < η, it is considered that the user's electricity consumption is abnormal.
2. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 1, characterized in that, after the electricity consumption data in Step 1 is subjected to data cleaning and filling of missing values, the power values of 96 measurement points per day of users are obtained, that is: P = {P 1 , P 2 ,..., P 96} (1).
3. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 2, characterized in that, In step 2, the daily active power P is equally divided into M intervals. For a certain small interval (P i , P i+1 ), assuming the width of the active power of each interval is H, then P i+1 -P i =H, that is, H is: Since the width of the small interval of active power is much smaller than the daily power magnitude, each power value within the interval can be regarded as the same point value P i .
4. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 3, characterized in that, In Step 3, the cumulative probability is used to obtain the cumulative distribution curve of each interval through the distribution cumulative probability function, and the expression of the distribution cumulative probability function is as follows: where α is the scale parameter, β is the shape parameter, and Y represents the number of indicators considered for each clustering center.
5. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 4, characterized in that, In Step 4, according to the cumulative probability function value of the user daily power distribution and the maximum and minimum index of daily power, assuming that the number of clusters to be divided is X, and the number of indicators considered for each clustering center is Y, the corresponding fuzzy matrix is established as: where a 1 , a 2 , …, a X is the number of clusters, a X is the number of the i-th clustering index, w x1 , w x2 , …, w xy are the coefficients of Y indicators corresponding to ai, where Y = M + 2.
6. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 5, characterized in that, In step five: calculate the density ρ of all power data in the power dataset X (w Xy ,D Xav ), ρ X is defined as a sample region centered on w xy with a radius of D Xav . Using the maximum value of the power consumption data density ρ Xmax as the first clustering center, the expression for the corresponding radius is: where x, z ∈ X.
7. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 6, characterized in that, In Step 6: Calculate the average distance density value d of the electricity power data X , specifically as follows:
8. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 7, characterized in that, In Step Nine: Update the data center with ∑mind' xy Take the minimum as the new cluster center. According to the new cluster center, perform clustering on the probability distribution curve using K-medoids clustering. The specific steps are as follows: Calculate the distance D(d′ xy ) from the electricity consumption data to ∑mind' inx , and the specific expression is as follows: D(d′ inx ) = angmin||d′ inxy - ∑mind′ xy || 2 (8) where: d' inxy is the data value in the x-th group of clusters; ∑mind' xy is the cluster center distance; Select d' inxy The average value of As the new clustering center, repeat the above formula.
9. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 8, characterized in that, In Step Eleven, the Longest Common Subsequence Algorithm (LCSS) is used to calculate the similarity of each cluster center. Taking two cluster center curves as an example, let the two cluster center curves C 1 and C 2 sequence point sets be C 1 ={a 1 , a 2 ,..., a Y} and C 2 ={b 1 , b 2 ,..., b Y}, then the length of the longest common subsequence is: where t = 1, 2, 3, …, Y; i = 1, 2, 3, …, Y; γ is the set distance threshold; LCSS(a t , b i ) is the length of the longest common subsequence of sequence C 1 before the trajectory point a t and sequence C 2 before the trajectory point b i ; dist(a t , b t ) is the distance between the a 1 -th point in the C t trajectory and the b 2 -th point in the C i sequence.
10. A detection method for identifying electricity theft based on cumulative probability distribution according to claim 9, characterized in that, In Step Twelve, the similarity of the curves is measured by calculating the number of trajectory points between two clustering center matching point distance thresholds; the matching rate ρ of two sequence point sets C 1 and C 2 is calculated as follows: where ρ(C 1 , C 2 ) ∈ [0, 1]; and the larger the value of ρ(C 1 , C 2 ), the more similar the C 1 sequence is to the C 2 sequence.
Citation Information
Patent Citations
Clustering and density estimation fused electricity stealing detection method and system
CN111539840A
Method and system for enhancing photovoltaic electricity larceny data based on NICE model
CN113919408A