Electricity stealing user detection method based on combined model
Through the combined model-based detection method of power stolen users, the problem of difficult to detect power stolen users under different power stolen methods in the prior art is solved, and efficient detection without pre-training is achieved, which is suitable for a variety of power stolen scenarios.
Patent Information
- Application Number
- CN202510257294.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively detect and identify power stolen users under different power stolen methods, and the existing methods require pre-training with the help of training sets, and the detection accuracy of different power stolen methods is poor.
The power stolen user detection method is adopted based on the combined model. By dividing the power stolen means into linear and nonlinear classes, setting up a multi-view identification model to form a combined detection framework, using smart power meter and electricity consumption information to collect system data, without pre-training, and all power stolen means have good detection accuracy.
It realizes efficient detection of different power theft methods, avoids interference from the imbalance of power theft samples, has strong adaptability and high detection efficiency, and provides reference for the research on anti-power stolen technology of power enterprises.
Smart Images

Figure CN120180329A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system big data analysis and anti-electricity theft, and specifically relates to a method for detecting electricity theft users based on a combined model. Background Art
[0002] With the continuous expansion of the user scale and the increasing diversity of load types, some users have committed electricity theft by tampering with meter readings, privately connecting lines, and signal interference in order to pay less or no electricity bills. Such phenomena are emerging in an endless stream, seriously damaging the interests of power companies, undermining the fairness of power transactions, and may cause potential safety hazards in electricity use, endangering people's lives and property. Due to the rapid development of science and technology, the means of electricity theft are becoming more and more diverse, including undervoltage method, undercurrent method, phase shift method, and bypass method. Different methods have certain differences in the mechanism of electricity theft, mathematical models, etc., and it is urgent to carry out research on electricity theft detection methods suitable for multiple scenarios based on the characteristics of each method.
[0003] With the popularization of smart electricity meters and the commissioning of platforms such as electricity consumption information collection systems and integrated line loss systems, power companies can obtain massive multi-dimensional spatiotemporal data such as users' time-of-day electricity consumption and area line losses, laying a solid foundation for the research of electricity theft detection methods based on intelligent algorithms. This technical route effectively improves the accuracy of traditional manual inspection methods and reduces related workload, and has become an important trend in the development of anti-electricity theft technology. Summary of the invention
[0004] In view of the problem that there are various existing means of electricity theft, and different means have certain differences in electricity theft mechanism, mathematical model, etc., the present invention provides a method for detecting electricity theft users based on a combined model. According to the different means of electricity theft mechanism and mathematical model, they are divided into linear and nonlinear categories, and multi-view recognition models are set for each type to form a combined detection framework. The proposed method makes full use of the data that can be collected by platforms such as smart electricity meters and electricity consumption information collection systems, does not require pre-training with the help of training sets, and has good detection accuracy for each electricity theft method, providing a reference for the research and application of anti-electricity theft technology for power companies.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for detecting electricity theft users based on a combined model, the method comprising the following steps:
[0007] Step 1: Obtain the data sequence of the power consumption of users under the jurisdiction of the substation and the line loss data of the substation to which they belong during the same period, segment and preprocess the power consumption sequence of the users to be inspected and the line loss sequence of the substation to which they belong by using a sliding window, and analyze the local characteristics of each curve;
[0008] Furthermore, the specific operations of step 1 are:
[0009] Step 1.1: For the missing part of the sequence, Lagrange interpolation is used to supplement the data, and the relevant formula is as follows:
[0010]
[0011] In the formula: n is the number of participating points, t is the serial number of the missing point, t i 、t j 、q are the serial numbers and electric quantities of two different known points respectively; q(t) represents the electric quantity after completion;
[0012] Step 1.2: Use several sliding windows of equal length to segment the user's electricity consumption curve and line loss curve. Among them, the window width is w metering points, and the sliding distance each time is s points. The schematic diagram is as Figure 1 shown.
[0013] Step 2: For the user's electricity consumption sequence after segmentation, extract the sample entropy features to extract the user's short-term electricity consumption pattern, and realize the extraction of non-linear electricity theft features through horizontal and vertical comparison of users;
[0014] Furthermore, the specific operation of Step 2 is as follows:
[0015] For a certain user to be inspected, calculate the sample entropy of the electricity consumption sequence in each window respectively to describe its short-term electricity consumption pattern. Then, conduct horizontal and vertical comparison on the statistical results of the sample entropy, that is, comparison of the same user in different periods and comparison of users in the same substation area in the same period, to analyze whether there are users with abnormal electricity consumption behaviors; when there are abnormal situations, mark the corresponding users and the time periods corresponding to their abnormal sample entropy values;
[0016] Step 2.1: Calculate the sample entropy of the electricity consumption sequence in each sliding window. Sample Entropy (SampEn) is a statistical method for users to measure the unpredictability and complexity of time series, which can effectively quantify the generation rate of new patterns in the electricity consumption sequence and describe its dynamic characteristics. According to the ratio of the electricity theft amount of the user to the electricity consumption in the same period, the electricity theft methods corresponding to the time-varying ratio and the constant ratio are classified as non-linear electricity theft and linear electricity theft respectively. The sample entropy statistics are especially suitable for non-linear electricity theft detection. The sample entropy calculation process is as follows: First, determine the embedding dimension m and the tolerance p, and convert the sequence to be inspected into an m-dimensional embedding vector and obtain the matching probability B m (p):
[0017]
[0018] In the formula, The subscript i of corresponds to the element serial number in the vector ;
[0019] Increase the embedding dimension to m + 1 to continue calculating the matching probability B m+1 (p), and obtain the sample entropy s(m, p):
[0020]
[0021] Step 2.2: For the area to be inspected, conduct horizontal and vertical comparisons on the sample entropy calculation results of the power consumption sequences of its subordinate users. Assume the number of users in the area is K, and the number of windows required for segmenting the power consumption sequences of each user is V, and obtain the horizontal comparison dataset S of the sample entropy of each user in the area at a certain moment h and the longitudinal verification dataset S of a certain user z :
[0022]
[0023] In the formula: the upper and lower subscripts of s are the user number and the window number respectively; represents the sample entropy of the power consumption sequence of user K within window v; represents the sample entropy of the power consumption sequence of user k within window V;
[0024] Step 2.3: Based on the Isolation Forest algorithm, conduct outlier detection on the above datasets respectively. The core idea of this algorithm is to construct multiple random binary trees by randomly selecting features and partitioning the data space, and then observe the depth of each sample in the tree to identify outliers. It is considered that due to the alienation of outliers from most samples, they are usually assigned to leaf nodes earlier and are closer to the root node. For the horizontal comparison dataset S h , the calculation formula for the outlier score h(s, K) of each user is as follows:
[0025]
[0026] v(K) = 2[ln(K - 1) + γ] - 2(K - 1) / K
[0027] In the formula: E(l(s)) is the average path length of the sample in each isolation tree, γ is the Euler constant; v(K) represents the expected path length; when the score is greater than the threshold δ1, it is considered that the corresponding user has a suspicion of non - linear power theft;
[0028] For the longitudinal verification dataset S z , the calculation method for the outlier score h(s, V) of each user is the same as that of the horizontal comparison dataset S h .
[0029] Step 3: Regarding the order of magnitude difference between the user power consumption sequence and the substation area line loss sequence, perform a linear transformation on each user power consumption sequence based on the least squares estimation, so that the cumulative distance between the transformed sequence and the line loss sequence in the same period is minimized, for subsequent analysis of the similarity degree between sequences and abnormal identification steps;
[0030] The least squares estimation target is to minimize the cumulative distance between each user power consumption sequence and the line loss sequence of its affiliated substation area. After this processing step, the comparability between the transformed power consumption sequence and the line loss sequence is significantly enhanced, which helps to more accurately evaluate the specific impact of different user power consumption patterns on the substation area line loss, and further deeply analyze the similarity degree and potential causal relationship between power consumption and line loss.
[0031] Further, the specific operation of Step 3 is as follows:
[0032] Since the subsequent analysis step involves calculating the morphological similarity between the user power consumption sequence and the substation area line loss sequence, and considering that the similarity calculation result is non-normalized and there is a significant difference in the order of magnitude between the user power consumption and the substation area line loss value, a linear transformation is performed on the power consumption sequence based on the least squares estimation to minimize the sum of the point-to-point distances between it and the line loss sequence; let the power consumption sequence and the substation area line loss sequence in the same period be X = [x t , t = 1, 2,..., T] and Y = [y t , t = 1, 2,..., T], respectively, then the transformed sequence and the objective function G are as follows:
[0033]
[0034] In the formula: is the transformed power consumption sequence; the least squares estimation target is to minimize the objective function G; a represents the sequence transformation magnification; b represents the sequence translation distance; T represents the time series length.
[0035] Step 4: Based on the transformed power consumption sequence and the line loss sequence, calculate the similarity degree between the two sequences by means of dynamic time warping, and extract potential linear power theft users through the isolation forest algorithm.
[0036] Further, the specific operation of Step 4 is as follows:
[0037] Use dynamic time warping (DTW) to mine the optimal matching relationship between the transformed power consumption sequence and the line loss sequence to alleviate the negative impact of data acquisition and transmission delay, and then quantitatively calculate the similarity degree between the sequences based on the cumulative distance between the matching data pairs. Mark the users with abnormal similarity degree as suspected linear power theft users, and obtain the comprehensive detection conclusion by integrating the identification results of multiple models.
[0038] The similarity calculation is implemented based on the dynamic time warping algorithm, which can find the most similar matching path between sequences in the presence of time warping or desynchronization between sequences, achieving the best alignment of the two sequences in time, thus greatly alleviating the negative impacts brought by data acquisition and transmission delays. The abnormal similarity identification is implemented based on the isolation forest algorithm. Users with electricity consumption sequences that have a high correlation with the line loss sequence or exhibit abnormal fluctuations are regarded as users with abnormal line loss contributions, that is, users suspected of linear electricity theft.
[0039] Step 4.1: For the substations with electricity theft users, the increased part of the line loss is mainly contributed by the electricity theft amount of the electricity theft users. Therefore, compared with most normal users, the electricity consumption sequences of electricity theft users should have a higher similarity with the substation line loss sequences, and this feature is particularly evident in linear electricity theft users. To sum up, based on dynamic time warping, the similarity analysis is carried out on the reconstructed electricity consumption sequences and line loss sequences of each user. The minimum distance between the two sequences is calculated by finding the optimal alignment path between them, and this is used as the basis for similarity evaluation. The specific steps are as follows:
[0040] Construct a distance matrix D, where the element D[i][j]:
[0041]
[0042] In the formula, represents the value of the i-th point in the linearly transformed electricity consumption sequence; y j represents the value of the j-th point in the line loss sequence;
[0043] Find the optimal path through dynamic programming to minimize the cumulative distance and initialize the boundary conditions:
[0044] k[i][j] = D[i][j] + min(k[i - 1][j], k[i][j - 1], k[i - 1][j - 1])
[0045] k[1][1] = D[1][1]
[0046] In the formula, k[i][j] represents the filled value in the i-th row and j-th column of the distance matrix; k[i - 1][j] represents the filled value in the (i - 1)-th row and j-th column of the distance matrix; k[i][j - 1] represents the filled value in the i-th row and (j - 1)-th column of the distance matrix;; k[i - 1][j - 1] represents the filled value in the (i - 1)-th row and (j - 1)-th column of the distance matrix; Read the data at k[T][T] as the similarity calculation result, and obtain the optimal matching path through backtracking; Compared with normal users, the similarity between the reconstructed electricity consumption sequence and the line loss sequence of electricity theft users is higher, and the cumulative distance is smaller;
[0047] Step 4.2: Based on the Isolation Forest algorithm, anomaly identification is carried out on the similarity calculation results. This analysis step is carried out from two aspects: the similarity amplitude and the fluctuation situation. At time t, for a certain power distribution area, except for nonlinear power stealing users, the similarity calculation results of the remaining users are where Z is the number of users to be measured. Anomaly detection of the above users by the Isolation Forest algorithm obtains the amplitude detection result and combines the threshold δ2 to extract abnormal users;
[0048] Step 4.3: Compare the change of the anomaly scores of each user at two consecutive measurement times, and combine the sigmoid function to achieve local feature amplification and obtain the sequence similarity fluctuation detection result The calculation formula is as follows:
[0049]
[0050] In the formula, represents the amplified result of the anomaly score of user z at measurement time t.
[0051] When the fluctuation amplitude exceeds the threshold δ3, it is considered that there is a power stealing critical point, that is, the user has changed from power stealing to non-power stealing, or from non-power stealing to power stealing during this period; furthermore, the relevant users are marked as linear power stealing suspected users, and finally the analysis results of each sub-model are summarized to generate the final detection conclusion. To sum up, the detection process of the combined model is as Figure 2 shown.
[0052] Furthermore, the thresholds δ1, δ2, and δ3 are set to 0.75, 0.75, and 0.3 respectively.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] 1) The models involved in the method all adopt unsupervised algorithms and do not need to be pre-trained with the help of a large amount of data, including power consumption feature extraction based on fuzzy entropy and power line loss analysis based on DTW. This effectively avoids the interference of the power stealing sample imbalance problem on detection, and has strong adaptability and high detection efficiency.
[0055] 2) The method sets different detection technical routes according to the differences in power stealing methods. The overall detection framework is multi-level and multi-perspective, and has good detection effects in different power stealing scenarios.
[0056] 3) The detection sub-model adopts the DTW algorithm for linear power stealing detection. This algorithm can effectively obtain the optimal matching relationship between power line loss sequences and alleviate the negative impacts caused by data collection and transmission delays. Description of the Drawings
[0057] Figure 1 It is a schematic diagram of sliding window segmentation;
[0058] Figure 2 Detection flow chart for the combined model;
[0059] Figure 3 A vertical comparison result chart for a certain user;
[0060] Figure 4 It is a horizontal comparison result diagram between users in a certain area at a certain time;
[0061] Figure 5 This is the result diagram of the similarity analysis of the power line loss sequence;
[0062] Figure 6 This is the result diagram of score amplitude anomaly detection;
[0063] Figure 7 This is the result diagram of abnormal score fluctuation detection. DETAILED DESCRIPTION
[0064] In order to gain a deeper understanding of the present invention, we will provide a comprehensive and detailed description of the present invention. However, the present invention has multiple implementations and is not limited to the specific examples listed herein. The presentation of these examples is intended to deepen the comprehensive understanding of the disclosure of the present invention.
[0065] A method for detecting electricity theft users based on a combined model, the method comprising the following steps:
[0066] Step 1: Obtain the data sequence of the power consumption of users under the jurisdiction of the substation and the line loss data of the substation to which they belong during the same period, segment and preprocess the power sequence of the users to be inspected and the line loss sequence of the substation to which they belong using a sliding window, and analyze the local characteristics of each curve;
[0067] Furthermore, the specific operations of step 1 are:
[0068] Step 1.1: For the missing part of the sequence, Lagrange interpolation is used to supplement the data. The relevant formula is as follows:
[0069]
[0070] Where: n is the number of participating points, t is the number of missing points, t i ,t j , q are the serial numbers and electric quantities of two different known points respectively; q(t) represents the electric quantity after completion;
[0071] Step 1.2: Use several sliding windows of equal length to segment the user power curve and line loss curve, where the window width is w metering points and the sliding distance each time is s points. The schematic diagram is as follows: Figure 1 shown.
[0072] Step 2: For the segmented user power consumption sequence, extract its sample entropy features to extract the short-term power consumption patterns of users, and realize the extraction of non-linear power theft features through horizontal and vertical comparisons of users;
[0073] Further, the specific operation of Step 2 is as follows:
[0074] For a user to be inspected, calculate the sample entropy of the power consumption sequence in each window respectively to describe its short-term power consumption pattern. Then, make horizontal and vertical comparisons of the sample entropy statistical results, that is, compare different periods of a single user and compare the same period among users in the same substation area to analyze whether there are users with abnormal power consumption behaviors; when there are abnormal situations, mark the corresponding users and the time periods corresponding to their abnormal sample entropy values;
[0075] Step 2.1: Calculate the sample entropy of the power consumption sequence in each sliding window. Sample Entropy (SampEn) is a statistical method for measuring the unpredictability and complexity of time series by users, which can effectively quantify the generation rate of new patterns in the power consumption sequence and describe its dynamic characteristics. According to the ratio of the power theft amount of the user to the power consumption during the same period, the power theft methods corresponding to the time-varying ratio and the constant ratio are classified as non-linear power theft and linear power theft respectively. Sample entropy statistics are especially suitable for non-linear power theft detection. The sample entropy calculation process is as follows: First, determine the embedding dimension m and the tolerance p, and convert the sequence to be inspected into an m-dimensional embedding vector and obtain the matching probability B m (p):
[0076]
[0077] In the formula, the subscript i of corresponds to the element serial number in the vector ;
[0078] Increase the embedding dimension to m + 1 to continue calculating the matching probability B m+1 (p), and obtain the sample entropy s(m, p):
[0079]
[0080] Step 2.2: For the substation area to be inspected, make horizontal and vertical comparisons of the sample entropy calculation results of the power consumption sequences of the users under its jurisdiction. Assume that the number of users in the substation area is K, and the number of windows required for segmenting the power consumption sequences of each user is V, and obtain the horizontal comparison dataset S of the sample entropy of each user in the substation area at a certain moment h and the longitudinal verification dataset S of a certain user z :
[0081]
[0082] In the formula: the superscript and subscript of s are the user number and window number respectively; represents the sample entropy of the user K’s electricity sequence in window v; represents the sample entropy of the user k’s electricity sequence in window V;
[0083] Step 2.3: Perform outlier detection on each of the above datasets based on the isolation forest algorithm. The core idea of the algorithm is to construct multiple random binary trees by randomly selecting features and dividing the data space, and then observe the depth of each sample in the tree to identify outliers. It is believed that due to the alienation of outliers from most samples, they are usually classified as leaf nodes earlier and closer to the root node. For the horizontal comparison dataset S h , the calculation formula for each user's abnormal score h(s,K) is as follows:
[0084]
[0085] v(K)=2[ln(K-1)+γ]-2(K-1) / K
[0086] Where: E(l(s)) is the average path length of the sample in each isolated tree, γ is the Euler constant; v(K) represents the expected path length; when the score is greater than the threshold δ1, it is considered that the corresponding user is suspected of nonlinear electricity theft;
[0087] For the longitudinal validation dataset S z The calculation method of each user's abnormal score h(s,V) is the same as the horizontal comparison dataset S h .
[0088] Step 3: Based on the order of magnitude difference between the user power sequence and the station area line loss sequence, each user power sequence is linearly transformed based on the least squares estimation, so that the cumulative distance between the transformed sequence and the line loss sequence of the same period is minimized, so as to facilitate the subsequent similarity analysis and anomaly identification steps between sequences;
[0089] The least squares estimation goal is to minimize the cumulative distance between each user's power sequence and the line loss sequence of the substation to which it belongs. After this processing step, the comparability between the transformed power sequence and the line loss sequence is significantly enhanced, which helps to more accurately evaluate the specific impact of different user power consumption patterns on the substation line loss, and then deeply analyze the similarity and potential causal relationship between power and line loss.
[0090] Furthermore, the specific operations of step 3 are:
[0091] Since the subsequent analysis steps involve calculating the morphological similarity between the user power consumption sequence and the substation area line loss sequence, and considering that the similarity calculation result is non-normalized and there are significant differences in the order of magnitude between the user power consumption and the substation area line loss value, a linear transformation is performed on the power consumption sequence based on the least squares estimation to minimize the sum of the point-to-point distances between it and the line loss sequence; let the power consumption sequence and the substation area line loss sequence in the same period be X = [x t , t = 1, 2,..., T] and Y = [y t , t = 1, 2,..., T], respectively, then the transformed sequence and the objective function G are as follows:
[0092]
[0093] In the formula: is the transformed power consumption sequence; the least squares estimation target is to minimize the objective function G; a represents the sequence transformation magnification; b represents the sequence translation distance; T represents the time series length.
[0094] Step 4: Based on the transformed power consumption sequence and the line loss sequence, calculate the similarity degree between the two sequences by means of dynamic time warping, and extract potential linear power theft users through the isolation forest algorithm.
[0095] Furthermore, the specific operation of Step 4 is as follows:
[0096] Use dynamic time warping (DTW) to mine the optimal matching relationship between the transformed power consumption sequence and the line loss sequence to alleviate the negative impact of data collection and transmission delay, and then quantitatively calculate the similarity degree of the sequences based on the cumulative distance between the matching data pairs. Mark the users with abnormal similarity degree as linear power theft suspect users, and obtain the comprehensive detection conclusion by integrating the identification results of multiple models.
[0097] The similarity calculation is implemented based on the dynamic time warping algorithm. This algorithm can find the most similar matching path between sequences in the case of time warping or out-of-sync between sequences, so that the two sequences are optimally aligned in time, thus greatly alleviating the negative impact brought by data collection and transmission delay. Abnormal similarity identification is implemented based on the isolation forest algorithm. The users corresponding to the power consumption sequences with high correlation or abnormal fluctuations with the line loss sequence are regarded as users with abnormal line loss contribution, that is, linear power theft suspect users.
[0098] Step 4.1: For the substations with electricity theft users, the increased part of the line loss is mainly contributed by the electricity theft amount of the electricity theft users. Therefore, compared with most normal users, there should be a higher similarity between the electricity consumption sequence of the electricity theft users and the line loss sequence of the substation, and the above characteristics are particularly significant for linear electricity theft users. To sum up, based on dynamic time warping, the similarity degree analysis is carried out on the reconstructed electricity consumption sequence and line loss sequence of each user. By finding the optimal alignment path between the two sequences, the minimum distance between them is calculated, and this is used as the basis for similarity evaluation. The specific steps are as follows:
[0099] Construct a distance matrix D, where the element D[i][j]:
[0100]
[0101] In the formula, represents the value of the i-th point in the electricity consumption sequence after linear transformation; y j represents the value of the j-th point in the line loss sequence;
[0102] Find the optimal path through dynamic programming to minimize the cumulative distance and initialize the boundary conditions:
[0103] k[i][j] = D[i][j] + min(k[i - 1][j], k[i][j - 1], k[i - 1][j - 1])
[0104] k[1][1] = D[1][1]
[0105] In the formula, k[i][j] represents the filled value in the i-th row and j-th column of the distance matrix; k[i - 1][j] represents the filled value in the (i - 1)-th row and j-th column of the distance matrix; k[i][j - 1] represents the filled value in the i-th row and (j - 1)-th column of the distance matrix; k[i - 1][j - 1] represents the filled value in the (i - 1)-th row and (j - 1)-th column of the distance matrix; Read the data at k[T][T] as the similarity calculation result, and obtain the optimal matching path through backtracking; Compared with normal users, the similarity between the reconstructed electricity consumption sequence and the line loss sequence of electricity theft users is higher, and the cumulative distance is smaller;
[0106] Step 4.2: Based on the isolation forest algorithm, anomaly identification is carried out on the similarity calculation results. This analysis step is carried out from two aspects: the amplitude and fluctuation of the similarity. Let the similarity calculation results of the users other than the non-linear electricity theft users at a certain substation at time t be where Z is the number of users to be measured, and the amplitude detection results are obtained by performing isolation forest anomaly detection on the above users and the abnormal users are extracted in combination with the threshold δ2;
[0107] Step 4.3: Compare the change in the anomaly scores of each user at two consecutive metering moments, and combine the sigmoid function to achieve local feature amplification and obtain the detection result of the volatility of sequence similarity The calculation formula is as follows:
[0108]
[0109] In the formula, represents the amplified result of the anomaly score of user z at the t metering moment.
[0110] When the fluctuation amplitude exceeds the threshold δ3, it is determined that there is a critical point of electricity theft, that is, the user has changed from electricity theft to non-electricity theft, or from non-electricity theft to electricity theft during this period; furthermore, mark the relevant users as suspected users of linear electricity theft, and finally summarize the analysis results of each sub-model to generate the final detection conclusion. To sum up, the detection process of the combined model is as Figure 2 shown.
[0111] Furthermore, the thresholds δ1, δ2, and δ3 are set to 0.75, 0.75, and 0.3 respectively.
[0112] Case study:
[0113] Construct a simulation substation area and simulate the electricity consumption behavior of users to verify the effectiveness of the proposed combined detection framework. The time span of the user electricity quantity and line loss sequences is 30 days, and the data acquisition frequency is once every 15 minutes. The width w of the sliding window is set to 3 days, that is, 288 acquisition points, and the sliding distance s each time is 96 points; the thresholds δ1, δ2, and δ3 are set to 0.75, 0.75, and 0.3 respectively.
[0114] Input the user electricity quantity data into the non-linear electricity theft detection model, and some results are as Figure 3 and Figure 4 shown. Furthermore, combine the user electricity quantity and the line loss of the substation area to conduct linear electricity theft detection, and obtain the detection results as Figure 5 、 Figure 6 and Figure 7 shown. The results show that the proposed combined model can accurately identify the preset electricity theft users under the substation area to be measured from multiple perspectives, that is, users User14 and 19. Although there are normal users User6 and 8 that are misdetected, the overall accuracy of the model still reaches 93%, showing good detection effect and applicability.
[0115] To sum up, the proposed method is for electricity users in low-voltage distribution substation areas, and can extract electricity consumption behavior characteristics through multi-perspective comprehensive analysis, accurately identify potential electricity theft situations, provide technical support for substation area line loss management work and the improvement of the level of compliance with electricity use, and contribute to the development of anti-electricity theft technology and the construction of a new power system.
[0116] Matters not covered by this invention are well-known technologies.
[0117] The content not described in detail in the specification of this invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of this technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
Claims
1. A method for detecting electricity theft users based on a combined model, characterized in that: The method comprises the following steps: Step 1: Obtain the data sequence of the power consumption of users under the jurisdiction of the substation and the line loss data of the substation to which they belong during the same period, and use the sliding window to segment and preprocess the data of the power consumption sequence of the users to be inspected and the line loss sequence of the substation to which they belong; Step 2: For the user power sequence after segmentation, sample entropy feature extraction is performed to extract the user's short-term power consumption pattern, and nonlinear power theft feature extraction is achieved through horizontal and vertical comparison of users; Step 3: Based on the order of magnitude difference between the user power sequence and the station area line loss sequence, each user power sequence is linearly transformed based on the least squares estimation, so that the cumulative distance between the transformed sequence and the line loss sequence of the same period is minimized; Step 4: Based on the transformed power sequence and line loss sequence, the similarity between the two sequences is calculated with the help of dynamic time warping, and potential linear electricity theft users are extracted through the isolation forest algorithm.
2. The method for detecting electricity theft users based on a combined model according to claim 1, characterized in that: The specific operations of step 1 are: Step 1.1: For the missing part of the sequence, Lagrange interpolation is used to supplement the data. The relevant formula is as follows: Where: n is the number of participating points, t is the number of missing points, t i ,t j , q are the serial numbers and electric quantities of two different known points respectively; q(t) represents the electric quantity after completion; Step 1.2: Use several sliding windows of equal length to segment the user power curve and the line loss curve, where the window width is w metering points and the sliding distance each time is s points.
3. The method for detecting electricity theft users based on a combined model according to claim 2, characterized in that: The specific operations of step 2 are: For a user to be inspected, the sample entropy of the electricity sequence in each window is calculated to describe its short-term electricity consumption pattern, and then the sample entropy statistical results are compared horizontally and vertically, that is, the single user is compared in different periods and the users in the same area are compared in the same period to analyze whether there are users with abnormal electricity consumption behavior; when there is an abnormal situation, the corresponding user and the time period corresponding to the abnormal sample entropy value are marked; Step 2.1: Calculate the sample entropy of the electricity sequence in each sliding window. According to the ratio of the user's stolen electricity and the electricity consumption in the same period, classify the electricity theft corresponding to the time-varying ratio and the constant ratio as nonlinear electricity theft and linear electricity theft respectively. The sample entropy calculation process is as follows: first determine the embedding dimension m and tolerance p, and convert the sequence to be tested into an m-dimensional embedding vector And get the template matching probability B m (p): In the formula, The subscript i corresponds to the vector The sequence numbers of the elements in Increase the embedding dimension to m+1 to continue calculating the matching probability B m+1 (p), and obtain the sample entropy s(m,p): Step 2.2: For the area to be inspected, the sample entropy calculation results of the user power sequence under its jurisdiction are compared horizontally and vertically. Assume that the number of users in the area is K, and the number of windows required for segmenting the power sequence of each user is V. The horizontal comparison data set S of the sample entropy of each user in the area at a certain time is obtained. h And a user longitudinal validation dataset S z : In the formula: the superscript and subscript of s are the user number and window number respectively; represents the sample entropy of the user K’s electricity sequence in window v; represents the sample entropy of the user k’s electricity sequence in window V; Step 2.3: Perform outlier detection on each of the above datasets based on the isolation forest algorithm. h , the calculation formula of each user's abnormal score h(s,K) is as follows: v(K)=2[ln(K-1)+γ]-2(K-1) / K Where: E(l(s)) is the average path length of the sample in each isolated tree, γ is the Euler constant; v(K) represents the expected path length; when the score is greater than the threshold δ1, it is considered that the corresponding user is suspected of nonlinear electricity theft; For the longitudinal validation dataset S z The calculation method of each user's abnormal score h(s,V) is the same as the horizontal comparison dataset S h .
4. The method for detecting electricity theft users based on a combined model according to claim 3 is characterized in that: The specific operations of step 3 are: Based on the least squares estimation, the power sequence is linearly transformed to minimize the sum of the point-to-point distances between it and the line loss sequence. Suppose the reconstructed power sequence and the area line loss sequence in the same period are X = [x t ,t=1,2,...,T] and Y=[y t ,t=1,2,…,T], then the transformed sequence and objective function G are: Where: is the transformed electricity series; the least squares estimation objective is to minimize the objective function G; a represents the sequence transformation ratio; b represents the sequence translation distance; T represents the length of the time series.
5. The method for detecting electricity theft users based on a combined model according to claim 4, characterized in that: The specific operations of step 4 are: Step 4.1: Based on dynamic time warping, the similarity analysis is performed on the reconstructed power sequence and line loss sequence of each user. The minimum distance between the two sequences is calculated by finding the optimal alignment path between the two sequences, and this is used as the basis for similarity evaluation. The specific steps are as follows: Construct a distance matrix D, where the element D[i][j] is: In the formula, represents the value of the i-th point in the electric quantity sequence after linear transformation; y j Represents the value of the jth point in the line loss sequence; Use dynamic programming to find the optimal path to minimize the cumulative distance and initialize the boundary conditions: k[i][j]=D[i][j]+min(k[i-1][j],k[i][j-1],k[i-1][j-1]) k[1][1]=D[1][1] In the formula, k[i][j] represents the value filled in the i-th row and j-th column of the distance matrix; k[i-1][j] represents the value filled in the i-1th row and j-th column of the distance matrix; k[i][j-1] represents the value filled in the i-th row and j-1th column of the distance matrix; k[i-1][j-1] represents the value filled in the i-1th row and j-1th column of the distance matrix; read the data at k[T][T] as the similarity calculation result, and obtain the optimal matching path by backtracking; Step 4.2: Based on the isolation forest algorithm, anomaly identification is performed on the similarity calculation results. Assume that at time t, except for linear electricity theft users in a certain area, the similarity calculation results of other users are Where Z is the number of users to be tested, and the above users are subjected to isolated forest anomaly detection to obtain the amplitude detection results. And combined with the threshold δ2 to extract abnormal users; Step 4.3: Compare the changes in the anomaly scores of each user at two consecutive measurement moments, combine the sigmoid function to achieve local feature amplification and obtain the sequence similarity volatility detection results The calculation formula is as follows: In the formula, represents the amplified result of the abnormal score of user z at measurement time t; When the fluctuation amplitude exceeds the threshold δ3, it is determined that there is a critical point of electricity theft, that is, the user has a transition from electricity theft to non-electricity theft, or non-electricity theft to electricity theft during this period; then the relevant users are marked as linear electricity theft suspects, and finally the analysis results of each sub-model are summarized to generate the final detection conclusion.
6. The method for detecting electricity theft users based on a combined model according to claim 5, characterized in that: The thresholds δ1, δ2 and δ3 are set to 0.75, 0.75 and 0.3 respectively.