A power stealing analysis method and system based on rough set and counter-inductive learning
By combining rough set theory and inverse learning, the electricity theft analysis process is optimized, solving the problem that existing technologies fail to effectively utilize expert experience and achieving more efficient and accurate electricity theft identification.
Patent Information
- Application Number
- CN202211629561.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-12-19
Smart Images

Figure CN116304791B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power systems, and particularly relates to a power stealing analysis method and system based on rough set and counter-inductive learning. BACKGROUND
[0002] Power stealing identification and processing is a long-standing topic in the power field, and has always been the focus of analysis by power companies. In some areas, power stealing is the main reason for high line loss in addition to technical line loss. How to identify power stealing more quickly and accurately is of great significance. The development of the current smart grid is very rapid, and a large number of smart Internet of Things devices are connected to the smart grid, especially smart meters with remote communication. More and more related data are collected, making it possible to identify power stealing. The current mainstream power stealing analysis methods are as follows:
[0003] 1. Artificial discrimination method: mainly through artificial discrimination by relevant staff. This method has high accuracy, but low efficiency and high cost.
[0004] 2. Power consumption and line loss analysis method: the characteristic data are mainly power consumption and line loss rate. This method has high accuracy for devices with obvious power consumption change characteristics and power consumption and line loss correlation characteristics, but the utilization rate of other characteristics is not high, the discrimination method is single, and is only suitable for some scenes.
[0005] 3. Machine learning method: machine learning uses powerful computer computing power to identify power stealing by identifying data anomalies. However, in the process of machine learning, only the characteristics of the data itself are used, without considering the use of expert experience related to power stealing analysis background.
[0006] 4. Unsupervised classification method: most of the data in the power stealing analysis process are unlabeled unsupervised data, which need to be classified and solved by unsupervised methods. In general, this method can identify some abnormal data, but the accuracy is low. SUMMARY
[0007] To solve the above problems, the present application provides a power stealing analysis method and system based on rough set and counter-inductive learning, which can solve the problem of difficult discrimination in the prior art, only considering the characteristics of the data itself, without considering the power stealing analysis related background. In addition, the formation of new rules and the deletion of old rules are realized by rough set, and the update of the background knowledge base can be realized.
[0008] The technical scheme adopted by the present application is as follows:
[0009] A power stealing analysis method based on rough set and counter-inductive learning, comprising the following steps:
[0010] S1. Classification analysis: according to whether the label data in the electricity stealing analysis background is sufficient, it is divided into supervised and unsupervised two cases, if it is supervised, then jump to step S4; otherwise, execute step S2;
[0011] S2. Sample augmentation: learning the electricity data set to obtain a new set of features, and augmenting the electricity data training samples based on the features; the electricity data set includes electricity stealing related background knowledge and electricity data training samples;
[0012] S3. Data processing: classify the original data, and compare the results with the results obtained by the domain knowledge, if both the users identified by the two are electricity stealing, then keep the electricity stealing label; otherwise, correct the label to non-electricity stealing, thereby forming a new corrected supervised sample;
[0013] S4. Machine learning: obtain the electricity stealing model hypothesis space through abductive learning, and obtain the hypothesis model for judging electricity stealing in the hypothesis space, test the hypothesis model based on the electricity data training samples, and judge whether to accept the hypothesis model according to the test result, if not, repeat this step;
[0014] S5. Data update: use the domain knowledge to generate a symbolic hypothesis of the logical fact of the electricity data training sample, use the hypothesis model to sample and observe the electricity data training sample, test the symbolic hypothesis based on the abductive logic reasoning and observation results, if consistent, output the symbolic hypothesis result; otherwise, modify the errors in the symbolic hypothesis, and continue to test according to the hypothesis model until the result is consistent with the domain knowledge; perform attribute reduction process of rough set, and reduce, modify and expand the domain knowledge to form a new domain knowledge base;
[0015] S6. Electricity stealing analysis: input the electricity data to be analyzed into the hypothesis model, and combine the domain knowledge base to perform electricity stealing analysis, and finally output the electricity stealing analysis result.
[0016] Further, the process of abductive logic reasoning includes: inputting a training data set D, when processing a training data sample e i =(x i ,y i ), identifying basic concept information p(x i ), y i and domain knowledge for reasoning and refining knowledge model Δ c .
[0017] Further, in the process of abductive logic reasoning, if the basic concept information p(x i ) is a pseudo label, the pseudo label is corrected by abductive logic reasoning, thereby obtaining new basic concept information p(x i), and apply it to the process after the inductive learning.
[0018] Further, in the process of inductive learning, the supervision information needs to be optimized simultaneously with the set P(X) of pseudo-labels generated in the learning process and the final knowledge model Δ i = (x i ,y i ) generated in the learning process. c
[0019] Further, the optimization goal is to maximize the number of consistent samples between the hypothesis model and the training data set D.
[0020] Further, the attribute reduction of data features using the neighborhood rough set can form four rules, two certainty rules and two possibility rules, respectively. Based on the four rules formed in the attribute reduction, the reduction, modification and expansion of the domain knowledge are further realized.
[0021] Further, in the process of attribute reduction, the attribute reduction is first performed by using the available condition attributes and the core condition attributes. After several cycles, the reduction that meets the reduction condition is obtained, and then the reducible or deletable condition attributes are substituted, and the reduction is continued until the reduction condition is met and the cycle is ended.
[0022] Further, the available condition attributes are the features in the original data, including the event code and the power factor; and the core condition attributes include the domain knowledge.
[0023] Further, after the cycle is ended, the pseudo-labels and the domain knowledge base are modified based on the knowledge of the local rough set theory and the pseudo-label neighborhood decision rough set theory.
[0024] A power stealing analysis system based on rough set and inductive learning, comprising:
[0025] A classification analysis module is used to divide the power stealing analysis background into two cases of supervised and unsupervised according to whether the label data is sufficient, and if it is supervised, a machine learning module is executed; otherwise, a sample augmentation module is executed.
[0026] A sample augmentation module is used to learn the power consumption data set to obtain a new set of features, and to augment the power consumption data training samples based on the features; the power consumption data set includes the background knowledge related to power stealing and the power consumption data training samples.
[0027] A data processing module is used to classify and process the original data, and compare the results with the results obtained from the domain knowledge. If both the identified users are power stealing, the power stealing label is retained; otherwise, the label is modified to be non-power stealing, thereby forming a new modified supervised sample.
[0028] The machine learning module is used to obtain a power stealing model hypothesis space through reverse deduction learning, to obtain a hypothesis model for judging power stealing through machine learning in the hypothesis space, to test the hypothesis model based on the power consumption data training sample, and to judge whether to accept the hypothesis model according to the test result, and if not, to repeat the step;
[0029] The data updating module is used to generate a symbolic hypothesis of logical facts for the power consumption data training sample by using domain knowledge, to sample and observe the power consumption data training sample by using the hypothesis model, to test the symbolic hypothesis based on reverse deduction logic reasoning and the observation result, to output a symbolic hypothesis result if consistent, otherwise, to modify errors in the symbolic hypothesis and continue to test according to the hypothesis model until the result is consistent with the domain knowledge, to perform an attribute reduction process of a rough set, and to delete, modify and expand the domain knowledge to form a new domain knowledge base;
[0030] The power stealing analysis module is used to input the power consumption data to be analyzed into the hypothesis model, to combine the domain knowledge base to perform power stealing analysis, and finally to output a power stealing analysis result.
[0031] The present application has the following advantages:
[0032] 1. The problem of combining expert knowledge and machine learning in the power stealing analysis problem is solved optimally, so that the operation is more efficient and more accurate.
[0033] 2. The problem of updating the domain knowledge base in the reverse deduction learning process is solved optimally, so that useless knowledge can be deleted and useful knowledge can be added.
[0034] 3. The work mode of manually judging one household at a time in the power stealing analysis process is solved optimally, which greatly improves the work efficiency and provides a basis for judging power stealing.
[0035] 4. The problem of only considering the data itself in the analysis process, which is separated from the power stealing analysis background, is solved optimally, and the problem nature is fully utilized. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a power stealing analysis method flow chart based on rough set and reverse deduction learning of embodiment 1 of the present application.
[0037] Figure 2 is a basic framework diagram of reverse deduction learning in the field of power stealing analysis of embodiment 1 of the present application. DETAILED DESCRIPTION
[0038] In order to make the technical features, objectives and effects of the present application clearer, the specific embodiments of the present application are described. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, that is, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0039] Embodiment 1
[0040] As Figure 1 shown, the present embodiment provides a power stealing analysis method based on rough set and abductive learning, comprising the following steps:
[0041] S1. Classification analysis: according to whether the label data in the power stealing analysis background is sufficient, it is divided into supervised and unsupervised two cases, if it is supervised, it is jumped to step S4; otherwise, step S2 is executed;
[0042] S2. Sample augmentation: learning the power consumption data set to obtain a new set of features, and augmenting the power consumption data training samples based on the features; the power consumption data set includes background knowledge related to power stealing and power consumption data training samples;
[0043] S3. Data processing: classifying the original data, and comparing the results with the results obtained from the domain knowledge, if both the users identified by the two are power stealing, the power stealing label is retained; otherwise, the label is corrected to non-power stealing, thereby forming new corrected supervised samples;
[0044] S4. Machine learning: obtaining a power stealing model hypothesis space through abductive learning, performing machine learning in the hypothesis space to obtain a hypothesis model for judging power stealing, testing the hypothesis model based on the power consumption data training samples, and judging whether to accept the hypothesis model according to the test result, if not, repeating the step;
[0045] S5. Data update: using domain knowledge to generate symbolic hypotheses of logical facts from power consumption data training samples, using the hypothesis model to sample and observe the power consumption data training samples, testing the symbolic hypotheses based on abductive logic reasoning and observation results, if consistent, output the symbolic hypothesis result; otherwise, modify the errors in the symbolic hypotheses, and continue to test according to the hypothesis model until the result is consistent with the domain knowledge; perform attribute reduction process of rough set, and reduce, modify and expand the domain knowledge to form a new domain knowledge base;
[0046] S6. Power stealing analysis: inputting the power consumption data to be analyzed into the hypothesis model, and combining the domain knowledge base to perform power stealing analysis, and finally outputting the power stealing analysis result.
[0047] Preferably, step S2 specifically involves: augmenting the samples using domain knowledge; based on an existing expert system, learning from the electricity dataset B∪E (where B represents background knowledge related to electricity theft and E represents training samples of electricity data) to obtain a new set of features u={u1,…,u i},in:
[0048]
[0049] Finally, the electricity consumption data will be used to train the samples {<x1,y1> … <x i ,y i Augmenting {x1, u1, y1>…<(x1, u1), ... i ,u i ),y i >}.
[0050] Preferably, after obtaining the new corrected supervised samples in step S3, they are then fed into a supervised version of the rough set-based inverse learning framework to obtain a classification model. This process of sample augmentation and inverse learning is then repeated until a model result satisfying the termination condition is obtained.
[0051] Preferably, for supervised learning, the learning process is constrained by domain knowledge. This involves using reverse logical reasoning to constrain the hypothesis space of the machine learning task based on the domain knowledge expressed in first-order logic language. Then, learning is performed in the constrained hypothesis space through statistical induction. The learning objective function is:
[0052]
[0053] s.tCon(KB∪{h})
[0054] Where m is the number of samples, L is the loss function, and h is the classification model. Under the assumption that the domain knowledge KB must be consistent with h, there will be no contradiction.
[0055] Preferably, such as Figure 2 As shown, the inverse learning framework adopted in this implementation mainly consists of three parts: machine learning, inverse reasoning (logical reasoning), and consistency optimization.
[0056] The machine learning part is mainly used to learn the recognition model p of the basic concept P. Obviously, since the input part required by the subsequent part, especially the abductive reasoning part, must be in the form of a logical language, the output part of the machine learning part must also be in the form of a logical language, and the learning of the logical language is generally supervised learning, but since there is no supervision information about the recognition model p in the input data set D, obtaining the supervision information about the recognition model p also needs to be achieved through abductive reasoning.
[0057] The abductive reasoning part is used to optimize the domain knowledge and obtain the basic concepts that may appear in the training samples through abductive reasoning, that is, to obtain the supervision information required by the recognition model p in machine learning. When machine learning is just started, since the supervision information is not perfect, the basic concept information p(x i ) learned from it is not perfect and may contain some errors, and such learned imperfect basic concept information p(x i ) is pseudo-labeled. The specific process of the abductive reasoning part is as follows: first, input the training data set D, and when processing the training data sample e i =(x i ,y i ), use the recognition model p obtained by the machine learning part to recognize the basic concept information p(x i ), y i and the domain knowledge KB for reasoning and refining the knowledge model Δ c ; in the process, if the basic concept information p(x i ) is a pseudo-labeled case, the abductive reasoning part can correct the pseudo-labeled through abductive logical reasoning, thereby obtaining new basic concept information p(x i ) and applying it to the subsequent process of abductive learning. Since the obtained basic concept information p(x i ) is uncertain, and the result of abductive logical reasoning is not necessarily unique, the learning process is relatively complex.
[0058] The consistent optimization part is used to solve the uncertainty of the pseudo-labeled and the non-uniqueness of the abductive logical reasoning result. As can be seen from the brief description of the abductive logical reasoning in the previous abductive logical reasoning program, if the abductive logical reasoning wants to get a unique result, it needs a constraint condition, so there will be a problem of non-unique abductive reasoning result. The only supervision information that can be used in the whole abductive learning process is the sample label Y={y1…y i} in the input data set D and the domain knowledge KB, and in the learning process, the supervision information needs to be used as the basis for processing the sample e i =(x i ,y iA set of pseudo-labels P(X) generated in the learning process and the final knowledge model Δ to be obtained c Both are optimized simultaneously, and the optimization target is to meet the optimization target, and finally actually realize the optimization of the hypothesis model H = p U Δc. The optimization target selected is to maximize the number of samples consistent with the hypothesis model H and the training set D, that is:
[0059]
[0060] Finally, the required model is obtained.
[0061] Then, the rough set is substituted. Considering that the neighborhood rough set can be used for attribute reduction of data features, and four rules can be formed, two certainty rules and two possibility rules. Based on the four rules formed in the attribute reduction, the expansion, deletion and modification of expert experience are the key breakthroughs of the present application. The specific process is that the rough set is substituted into the consistent optimization process, and the feature knowledge is divided into conditional attributes and decision attributes. The decision attribute is the pseudo-label obtained by unsupervised learning and the pseudo-label to be modified by domain knowledge or the label item in supervised learning. The conditional attribute is divided into two parts, one part is the feature in the original data such as event code, power factor, etc. This part is called available conditional attribute, and the other part is composed of domain knowledge and called core conditional attribute. The other part is called reducible or deletable conditional attribute.
[0062] In the attribute reduction process, the attribute reduction is first performed by the available conditional attribute and the core conditional attribute. After several cycles, the reduction satisfying the reduction condition is obtained, and then the reducible or deletable conditional attribute is substituted, and the reduction is continued until the reduction condition is satisfied, and the cycle is ended. Based on the knowledge of the local rough set theory and the pseudo-label neighborhood decision rough set theory, the pseudo-label and the knowledge base are modified.
[0063] It should be noted that:
[0064] 1. The three parts of conditional attributes can only modify adjacent parts in one cycle.
[0065] 2. The three parts of conditional attributes can only modify one part in one cycle, and the modification is in the order of available conditional attribute first, core conditional attribute second, and reducible or deletable conditional attribute third.
[0066] 3. The core knowledge attribute obtained each time is used as domain knowledge to enter the logical reasoning part to identify the pseudo-labels to be modified.
[0067] Preferably, the present embodiment provides an anti-inductive learning algorithm based on rough set optimization, as follows.
[0068]
[0069]
[0070] Preferably, the embodiment also provides a pseudo-label neighborhood optimization method algorithm, specifically as follows.
[0071]
[0072] Specifically, the electricity stealing field knowledge used in the embodiment is as follows:
[0073] ① Disassembling the face cover to steal electricity
[0074] Electricity stealing method description: the user opens the tail cover of the electric meter, then opens the face cover of the electric meter, and changes the internal hardware connection of the electric meter to achieve the purpose of stealing electricity.
[0075] Event condition: tail cover open, face cover open.
[0076] Time condition: the tail cover opening time is t1, and the face cover opening time is t2
[0077] t1 must be earlier than t2
[0078] 36h≥t2-t1≥2min
[0079] Electricity condition: after the event occurs, the next day's daily frozen increment is 0 (the sampling circuit is damaged), or the daily frozen collection fails (the sampling chip is damaged), or the next day's daily frozen increment is less than 1kwh (the sampling circuit is short-circuited, the electric meter can be measured, but the measured value is less than the true value).
[0080] ② Short-circuit electricity stealing
[0081] Electricity stealing method description: the user opens the tail cover of the electric meter, short-circuits the live line in and out of the electric meter, and then connects the house line from the zero live line. At this time, the electric meter still has electricity.
[0082] Event condition: tail cover open, current imbalance.
[0083] Time condition: the tail cover opening time is t1, and the current imbalance occurrence time is t2
[0084] t1 must be earlier than t2
[0085] 36h≥t2-t1≥2min
[0086] Electricity condition: after the event occurs, the next day's daily frozen increment is not 0.
[0087] ③ User's electricity stealing due to arrears
[0088] Electricity stealing method description: the user pulls the line from somewhere to supply power to his home after the electric meter is pulled off.
[0089] Event condition: relay pull the grid after
[0090] Current imbalance event occurs
[0091] Daily settlement increment ≥ 0.5kwh
[0092] Time condition: none.
[0093] (4) Fire zero reverse connection electricity stealing
[0094] Electricity stealing method description: the electric meter fire zero reverse connection, when the relay is pulled, the user's fire line still has electricity, and the user connects the zero line from elsewhere, which can realize electricity stealing.
[0095] Event condition: fire zero reverse connection event occurs
[0096] And relay pull the grid state
[0097] Current imbalance event occurs
[0098] Time condition: none.
[0099] (5) CT meter electricity stealing
[0100] Electricity stealing method description: the voltage line and the current line of three-phase are staggered. (Similar to Ua*Ib, Uc*Ia)
[0101] Event condition: the electric meter reports the reverse phase sequence event phase sequence reversal
[0102] Or the electric meter power factor is negative
[0103] Or the power factor is greater than 1
[0104] Time condition: none.
[0105] Example 2
[0106] This embodiment is based on example 1:
[0107] This embodiment provides an electricity stealing analysis system based on rough set and counter-inductive learning, which includes a classification analysis module, a sample augmentation module, a data processing module, a machine learning module, a data updating module and an electricity stealing analysis module, wherein:
[0108] The classification analysis module is used to divide into two cases of supervised and unsupervised according to whether the label data in the electricity stealing analysis background is sufficient, if it is supervised, the machine learning module is executed; otherwise, the sample augmentation module is executed.
[0109] The sample augmentation module is used to learn the electricity data set to obtain a new set of features, and augment the electricity data training sample based on the features; the electricity data set includes background knowledge related to electricity stealing and electricity data training sample.
[0110] The data processing module is configured to classify the original data, compare the result with a result obtained by the domain knowledge, and retain the electricity stealing label if both the results identify the user as stealing electricity; otherwise, correct the label to non-electricity stealing, thereby forming a new corrected supervised sample.
[0111] The machine learning module is configured to obtain a hypothesis space of the electricity stealing model by inductive learning, perform machine learning in the hypothesis space to obtain a hypothesis model for judging electricity stealing, test the hypothesis model based on the electricity consumption data training sample, and determine whether to accept the hypothesis model according to a test result; if not, repeat the step.
[0112] The data updating module is configured to generate a symbolic hypothesis of a logical fact of the electricity consumption data training sample by using the domain knowledge, perform sampling observation on the electricity consumption data training sample by using the hypothesis model, test the symbolic hypothesis based on the inductive logic reasoning and the observation result, output a result of the symbolic hypothesis if the result is consistent; otherwise, modify an error in the symbolic hypothesis, and continue to test the symbolic hypothesis according to the hypothesis model until the result is consistent with the domain knowledge; perform an attribute reduction process of a rough set, and reduce, modify and expand the domain knowledge, thereby forming a new domain knowledge base.
[0113] The electricity stealing analysis module is configured to input the electricity consumption data to be analyzed into the hypothesis model, perform electricity stealing analysis in combination with the domain knowledge base, and finally output an electricity stealing analysis result.
[0114] It should be noted that, for the method embodiments described above, in order to simplify the description, the method embodiments are described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the order of actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
Claims
1. A method for analyzing electricity theft based on rough set theory and inverse learning, characterized in that, Includes the following steps: S1. Classification Analysis: Based on whether the tag data in the background of electricity theft analysis is sufficient, it is divided into two cases: supervised and unsupervised. If it is a supervised case, then proceed to step S4. Otherwise, proceed to step S2; S2. Sample Augmentation: The electricity consumption dataset is learned to obtain a new set of features, and the electricity consumption data training samples are augmented based on the features; the electricity consumption dataset includes background knowledge related to electricity theft and electricity consumption data training samples; S3. Data Processing: Classify the raw data and compare the results with those obtained from domain knowledge. If both identify users as electricity thieves, retain the electricity thief label; otherwise, correct the label to non-electricity thief, thus forming a new corrected supervised sample. S4. Machine Learning: Obtain the hypothesis space of the electricity theft model through inverse learning, perform machine learning in the hypothesis space to obtain the hypothesis model used to judge electricity theft, test the hypothesis model based on the training samples of electricity consumption data, and determine whether to accept the hypothesis model based on the test results. If not, repeat this step. S5. Data Update: Using domain knowledge, generate symbolic hypotheses of logical facts from the electricity consumption data training samples. Use the hypothetical model to sample and observe the electricity consumption data training samples. Test the symbolic hypotheses based on reverse logical reasoning and observation results. If they are consistent, output the symbolic hypothesis results; otherwise, correct the errors in the symbolic hypotheses and continue to test them according to the hypothetical model until the results are consistent with the domain knowledge. Perform rough set attribute reduction and reduce, modify and expand the domain knowledge to form a new domain knowledge base. S6. Electricity Theft Analysis: Input the electricity consumption data to be analyzed into the hypothetical model, combine it with the domain knowledge base to perform electricity theft analysis, and finally output the electricity theft analysis results.
2. The electricity theft analysis method based on rough set and inverse learning according to claim 1, characterized in that, The reverse logical reasoning process includes: inputting a training dataset. When processing training data samples At that time, among them For feature labeling, Label the samples to identify basic conceptual information. Sample labeling Domain knowledge is used for reasoning and refining knowledge models. .
3. The electricity theft analysis method based on rough set and inverse learning according to claim 2, characterized in that, In the process of reverse logical reasoning, if basic conceptual information is encountered... In the case of pseudo-labels, reverse logical reasoning is used to correct the pseudo-labels, thereby obtaining new basic conceptual information. And apply it to the process after reverse learning.
4. The electricity theft analysis method based on rough set and inverse learning according to claim 3, characterized in that, In the inverse learning process, supervised information needs to be used as a basis for training data samples. The collection of pseudo-labels generated during the learning process. and the knowledge model that should ultimately be obtained Both are optimized simultaneously.
5. The electricity theft analysis method based on rough set and inverse learning according to claim 4, characterized in that, The goal of optimization is to achieve the hypothesis model and training dataset. Maximize the number of consistent samples.
6. The electricity theft analysis method based on rough set and inverse learning according to any one of claims 1-5, characterized in that, Using neighborhood rough sets for attribute reduction of data features can form four rules: two deterministic rules and two probabilistic rules. Based on these four rules formed in attribute reduction, the domain knowledge can be reduced, modified, and expanded.
7. The electricity theft analysis method based on rough set and inverse learning according to claim 6, characterized in that, In the process of attribute reduction, attribute reduction is first performed using available conditional attributes and core conditional attributes. After several iterations, when a reduction condition that meets the reduction conditions is obtained, the reduceable or deleteable conditional attributes are substituted in to continue the reduction process until the reduction conditions are met, and the loop ends.
8. The electricity theft analysis method based on rough set and inverse learning according to claim 7, characterized in that, The available conditional attributes are features in the original data, including event codes and power factors; the core conditional attributes include domain knowledge.
9. The electricity theft analysis method based on rough set and inverse learning according to claim 7, characterized in that, After the loop ends, the pseudo-labels and domain knowledge base are corrected based on the knowledge of local rough set theory and pseudo-label neighborhood decision rough set theory.
10. A power theft analysis system based on rough set theory and inverse learning, characterized in that, include: The classification analysis module is used to classify the electricity theft analysis background into two cases: supervised and unsupervised, depending on whether the tag data is sufficient. If it is a supervised case, the machine learning module is executed. Otherwise, execute the sample augmentation module; The sample augmentation module is used to learn from the electricity consumption dataset, obtain a new set of features, and augment the electricity consumption data training samples based on the features; the electricity dataset includes background knowledge related to electricity theft and electricity consumption data training samples; The data processing module is used to classify the raw data and compare the results with the results obtained from domain knowledge. If both identify users as electricity thieves, the electricity thief label is retained; otherwise, the label is corrected to non-electricity thief, thus forming a new corrected supervised sample. The machine learning module is used to obtain the hypothesis space of the electricity theft model through inverse learning, perform machine learning in the hypothesis space to obtain the hypothesis model used to judge electricity theft, test the hypothesis model based on the training samples of electricity consumption data, and determine whether to accept the hypothesis model based on the test results. If not, repeat this step. The data update module is used to generate symbolic hypotheses of logical facts from electricity consumption data training samples using domain knowledge. It then uses the hypothesis model to sample and observe the electricity consumption data training samples, and tests the symbolic hypotheses based on reverse logical reasoning and observation results. If they match, the symbolic hypothesis result is output; otherwise, errors in the symbolic hypothesis are corrected, and the testing continues based on the hypothesis model until the result matches the domain knowledge. Finally, it performs rough set attribute reduction and deletes, modifies, and expands the domain knowledge to form a new domain knowledge base. The electricity theft analysis module is used to input the electricity consumption data to be analyzed into the hypothetical model, combine it with the domain knowledge base to perform electricity theft analysis, and finally output the electricity theft analysis results.
Citation Information
Patent Citations
Electricity stealing user discrimination method based on label augmentation
CN113379322A
Automatic electricity stealing identification method based on data mining technology
CN113408658A