Intelligent Electric Meter Operating State Evaluation Method Based on Entropy Weight Method and Random Forest Model

Through the smart meter evaluation method combined with the entropy weight method and the random forest model, the problems of low efficiency and poor adaptability of the traditional evaluation method are solved, and the efficient and accurate state evaluation of the smart meter is achieved, ensuring the safe operation of the power grid.

CN114757282BActive Publication Date: 2025-07-25GUANGXI POWER GRID CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210401585.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-07-25
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The prior art is difficult to monitor and accurately evaluate the operating status of smart meters in real time. The traditional periodic inspection is inefficient and cannot adapt to the randomness of the power metering devices of different manufacturers, resulting in challenges in the safe operation of the power grid.

Method used

The smart meter operating state evaluation method based on entropy weight method and random forest model is adopted, and the weight is optimized through data preprocessing, RandomForest model construction and entropy weight method to generate the final evaluation results of the smart meter.

Benefits of technology

It realizes timely status evaluation of smart meters, improves the accuracy and pertinence of evaluation, and ensures the stable operation of the smart grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114757282B_ABST
    Figure CN114757282B_ABST
Patent Text Reader

Abstract

Intelligent electricity meter operation status evaluation method based on entropy weight method and random forest model, including preprocessing the business system data of intelligent electricity meters; establishing a RandomForest model to generate a preliminary evaluation result of the intelligent electricity meter; obtaining a final evaluation result through weight calculation and processing, characterized in that: the evaluation method adds a step of further optimizing the weight of the preliminary calculation of the RandomForest model through the entropy weight method to obtain a final evaluation result, and the specific steps of further optimizing the weight are as follows: S31. Perform data standardization processing on multiple indicators screened by feature engineering; S32. Calculate the information entropy of the standardized data by means of the entropy weight method; S33. Fuse the weight initially calculated by the RandomForest model with the information entropy obtained in step 2 to obtain an optimized weight, and then obtain a training model, and evaluate the operation status of the intelligent electricity meter through the training model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent electricity meter operation status evaluation, and particularly relates to an intelligent electricity meter operation status evaluation method using an integrated learning algorithm, especially an intelligent electricity meter operation status evaluation method based on the entropy weight method and the random forest model. Background Art

[0002] The number of intelligent electricity meters and intelligent terminals installed within the power grid company is increasing. It is very difficult to achieve full-coverage periodic on-site inspections under the existing human conditions. On the other hand, the traditional periodic on-site inspection has low work efficiency, weak detection rigor and pertinence, cannot perform real-time monitoring of intelligent electricity meters, and it is difficult to timely discover operation failures and causes that occur between two on-site inspections. In addition, this fixed-cycle inspection mode does not consider the randomness of the operation status of power metering devices of different batches and different manufacturers, and it is difficult to meet the requirements of power metering technology and the refined management of power companies. Therefore, how to accurately and reasonably evaluate the status of intelligent electricity meters is an important issue for the safe operation of the power grid.

[0003] The traditional evaluation index system artificially sets indexes according to factors such as reliability maintenance, safety domain, family defects, and error characteristics. Its index system is simply set up and is not very adaptable to the status evaluation of tens of thousands of electricity meters of power grid companies. In terms of the dynamic evaluation index system, the decision tree method is used to adjust the electricity meter status evaluation index system, which can effectively solve the problem of effective production of the evaluation index set. However, since the data collected by the dynamic evaluation index system for electricity meters is from manual input and the weights of the evaluation index set cannot be adjusted, the quality of the intelligent electricity meter status evaluation is not high. Summary of the Invention

[0004] The present invention proposes an intelligent electricity meter operation status evaluation method based on the EntropyWeight-RandomForest algorithm, which uses a mature integrated learning algorithm to construct an intelligent electricity meter operation status evaluation system. First, operation data such as voltage, current, and power are associated to generate a feature sample data set. Secondly, a basic weak classifier is constructed using the random forest algorithm, and the weight parameters of the base classifier are optimized by the entropy weight method. Finally, the final classifier is obtained through weighted summation for the evaluation of the intelligent electricity meter operation status.

[0005] The technical solution of the present invention is an intelligent electricity meter operation status evaluation method based on the entropy weight method and the random forest model, including preprocessing the business system data of the intelligent electricity meter; establishing a RandomForest model to generate a preliminary evaluation result of the intelligent electricity meter; obtaining the final evaluation result through weight calculation and processing. The key is that the evaluation method adds a step of further optimizing the weight of the preliminary calculation of the RandomForest model through the entropy weight method to obtain the final evaluation result. The specific steps of further optimizing the weight are as follows:

[0006] S31. Perform data standardization processing on multiple indicators screened by feature engineering;

[0007] S32. Calculate the information entropy of the standardized data by means of the entropy weight method;

[0008] S33. Fuse the weight initially calculated by the RandomForest model with the information entropy obtained in step 2 to obtain the optimized weight, and then obtain the training model, and evaluate the operation status of the intelligent electricity meter through the training model.

[0009] Furthermore, the specific steps of preprocessing the business system data of the intelligent electricity meter are as follows:

[0010] S11. For missing values, if the missing ratio is greater than or equal to 40% and less than or equal to 70%, convert the field to an index field, with missing being 1 and non-missing being 0. If the missing ratio is greater than 70%, this field is not included in the model calculation; if the missing ratio is less than 40%, fill the discrete data with the mode and fill the continuous numerical values with the mean;

[0011] S12. For abnormal data, statistically monitor the numerical attributes, calculate the mean and standard deviation of the field values, and identify abnormal fields and abnormal records according to the data interval of each field;

[0012] S13. For redundant data, eliminate the approximate duplicate record problem in the data set according to business logic and technical means, and delete or generate derivative fields for the redundant data.

[0013] Furthermore, the specific steps of establishing a RandomForest model and initially calculating the weight are as follows:

[0014] S21. Perform data fusion on the preprocessed data, including operation data, evaluation data, power outage data, fault data, etc., to form sample data;

[0015] S22. Screen out the feature factors with a correlation greater than 0.3, and perform index screening in combination with specific business guidelines;

[0016] S23. For the sample data after data preprocessing, construct the training set and test set required for model training through the following method: Use the relevant features affecting the intelligent electricity meter evaluation system, such as equipment ledger data, operation data, fault data, power outage data, etc., as the model input X, and the intelligent electricity meter evaluation result as the model output Y to construct the dataset (X, Y) required for the model. For the constructed dataset (X, Y), randomly extract data at a ratio of training set: test set of 7:3, and finally construct the training set and test set required for the model;

[0017] S24. Initialize the model parameters, set the number of weak classifiers to 32; set the standard for feature segmentation to "squared_error"; set the depth of the weak classifier to 64; set the minimum number of samples for splitting to 2; set the maximum number of features for sample splitting to "auto"; set the out-of-bag score to "True"; set the number of training parallelisms to -1;

[0018] S25. Random sampling. The total number of training samples is N, where N is a positive integer greater than 1. Each single decision tree randomly extracts n samples from the N training sets with replacement, where n is a positive integer less than N. The probability of each sample being sampled is 1 / N, so the probability of not being sampled is 1-(1 / N). If it is not sampled after N samplings, then 35-40% of the data in the training set is not sampled and can be used as an evaluation dataset to evaluate the out-of-bag score of the model;

[0019] S26. Feature attribute division. The number of input features of the training example is N, and n is much smaller than N. Then, when splitting at each node of each decision tree, randomly select m input features from M input features, and then select the best one from these m input features for splitting. M is a positive integer greater than 1, and m is a positive integer less than M. For the regression scenario, the attribute split is based on the sample variance. For the classification scenario, the Gini index is used as the basis for attribute division;

[0020] S27. Construction of n weak classifiers. Combine the n sample datasets obtained by sampling with replacement, and randomly select a certain number of optimal features from all the features to be selected in each data subset as the input features of the decision tree. The dimension of the input features is [n*m*k], where n is the n data subsets, m is the number of samples, and k is the number of features. Input the feature subset into the weak classifier to construct n basic decision tree models;

[0021] S28. Model combination. Combine the n weak classifiers obtained in step S27 and perform summation and averaging;

[0022] S29. Preliminary calculation of weights. Calculate the change amount of the Gini index obtained through step S26;

[0023] S210. Calculate the scoring index by normalizing the importance scores to obtain the final scoring index.

[0024] The beneficial effects of the present invention are as follows:

[0025] 1. By using machine learning methods, analyze the operation data, power outage data, fault data, evaluation data, etc. of smart meters. At the same time, use the Pearson correlation coefficient for feature extraction, and combine the entropy weight method and the RF model to obtain an evaluation system, reasonably judge the operation status of smart meters, timely handle abnormal electric energy meters, and ensure the stable operation of the smart grid.

[0026] 2. Combine big data analysis with machine learning to deeply mine the factor characteristics affecting the operation status of smart meters, analyze the weight sizes among various influencing factors, and carry out relevant work such as operation and maintenance in a targeted manner, providing an auxiliary decision-making basis for work such as the condition-based maintenance of electric energy meters and spare parts. Brief Description of the Drawings

[0027] Figure 1 It is a schematic flow diagram of the evaluation method of the present invention.

[0028] Figure 2 It is a schematic diagram of the construction of the RandomForest model in the present invention. Detailed Embodiments

[0029] Related English Explanations:

[0030] EntropyWeight: Entropy weight method; RandomForest model: Random forest model; Pearson correlation coefficient: It is used to measure whether two data sets are on the same line, and it is used to measure the linear relationship between interval variables.

[0031] The present invention provides an evaluation method for the operation status of smart meters based on the entropy weight method and the random forest model. Refer to Figure 1 and Figure 2 . The core technology of the present invention is to add a step of further optimizing the weight of the preliminary calculation of the RandomForest model through the entropy weight method, and combine Figure 1 and Figure 2 to specifically illustrate the overall evaluation method of the present invention.

[0032] The technical solution of the present invention includes three main steps: data preprocessing, construction of the RandomForest model, and entropy weight optimization of weights, which are specifically as follows:

[0033] 1. Data preprocessing

[0034] Data preprocessing mainly involves associating data from various business systems to form a modeling wide table. Data cleaning and transformation need to be carried out from the following aspects:

[0035] S11. Missing data processing: For missing values, if the missing proportion is around 70%, the field is converted into an indicator field, where 1 represents missing and 0 represents non - missing. If the missing proportion is greater than 90%, this field is not included in the model calculation. If the missing proportion is less than 40%, the mode can be used to fill discrete - type data, and the mean can be used to fill continuous numerical data.

[0036] S12. Abnormal data processing: Statistically monitor numerical attributes, calculate the mean and standard deviation of field values, and identify abnormal fields and records by considering the data range of each field segment.

[0037] S13. Redundant data cleaning: Based on business logic and technical means, eliminate approximate duplicate record problems in the data set, and delete redundant data or generate derivative fields, etc.

[0038] 2. RandomForest model construction

[0039] The data obtained after data preprocessing has the problem of class - data imbalance. Therefore, sampling of imbalanced data is required before modeling. After completing data preprocessing, a sample data modeling wide table is obtained. First, analyze the correlation between sample features and target variables using the Pearson correlation coefficient, select feature fields that meet the modeling requirements, and initially construct using the RandomForest algorithm.

[0040] S21: Use methods such as Python and SQL for data fusion, including operation data, evaluation data, power outage data, fault data, etc., to form sample data.

[0041] S22: Use the Pearson correlation coefficient to analyze the correlation size of sample features. Screen out feature factors with a correlation size greater than 0.3, and combine with specific business guidelines for index screening.

[0042] S33: Construct a training set and a test set. For the sample data after data preprocessing, construct the training set and test set required for model training through the following method: Use relevant features such as equipment ledger data, operation data, fault data, and power outage data that affect the intelligent electricity meter evaluation system as the model input X, and the intelligent electricity meter evaluation result as the model output Y to construct the required data set (X, Y) for the model. For the constructed data set (X, Y), randomly extract data in a ratio of training set:test set = 7:3, and finally construct the training set and test set required for the model.

[0043] S24: Initialize model parameters. The number of weak classifiers (n_estimators) is set to 32; the criterion for feature splitting is set to "squared_error"; the maximum depth of the weak classifier (max_depth) is set to 64; the minimum number of samples for splitting (min_samples_split) is set to 2; the maximum number of features during sample splitting (max_features) is set to "auto"; the out-of-bag score (oob_score) is set to "True"; the number of parallel training processes (n_jobs) is set to -1.

[0044] S25: Random sampling. If the total number of training samples is N, each single decision tree randomly draws n samples from the N training sets with replacement as the training samples for this single tree. The probability of each sample being sampled is 1 / N, so the probability of not being sampled is 1 - (1 / N). If it is not sampled after N samplings, the probability is:

[0045]

[0046] When N → ∞,

[0047] That is, approximately 36.8% of the data in the training set is not sampled and can be used as an evaluation dataset to evaluate the out-of-bag score (oob_score) of the model.

[0048] S26: Feature attribute splitting. If the number of input features of the training examples is N and n is much smaller than N, then when splitting at each node of each decision tree, randomly select m input features from the M input features, and then select the best one from these m input features for splitting.

[0049] For the regression scenario, attribute splitting is based on the sample variance, and its calculation formula is:

[0050]

[0051] where x i is the value of the i-th feature and μ is the sample mean.

[0052] For the classification scenario, the Gini index is used as the basis for attribute splitting, and its calculation formula is as follows:

[0053]

[0054] Δi = i(N L ) - i(N R ) Equation 4

[0055] where N represents the unsplit node, N L and NR represent the separated left and right nodes respectively, W i is the weight of class c samples, n i represents the number of samples of each class within the node, and Δi represents the reduction in feature impurity.

[0056] S27: Construction of n weak classifiers. Combining the n sample datasets obtained by sampling with replacement, a certain number of optimal features are randomly selected from all the candidate features of each data subset as the input features of the decision tree. The dimension of the input features is [n * m * k], where n is the n data subsets, m is the number of samples, and k is the number of features. The feature subset is input into the weak classifier to construct n basic decision tree models.

[0057] S28: Model combination. Combining the n weak classifiers obtained in S27, the sum is taken and the average is obtained to get the final result. The calculation rule is as follows:

[0058]

[0059] where f i (x) is the calculation result of the i-th weak classifier model.

[0060] S29: Weight calculation. Through the Gini index calculated in S26, calculate its change amount. The calculation formula is as follows:

[0061]

[0062] where k represents that the data has k classes, p mk represents the proportion of class k in node m.

[0063] Feature x j The importance in node m, that is, the change amount of the Gini index before and after branching of node m is:

[0064]

[0065] where GI l and GI r represent the Gini indices of the two new nodes after branching respectively.

[0066] S210: Scoring index calculation. If the node where feature x j appears in the decision tree i is in the set M, then the importance of x j in the i-th tree is:

[0067]

[0068] If there are n trees in total, then:

[0069]

[0070] Finally, the importance score is normalized to obtain the final score index, and its calculation formula is as follows:

[0071]

[0072] 3. Entropy weight optimization of weights

[0073] The entropy weight method is used to optimize the weight parameters of the model, adjust the RandomForest model, and optimize the model weights. Finally, the EntropyWeight-RandomForest intelligent electricity meter operation status evaluation model is obtained. The entropy weight method is a method for adjusting the weight of a certain value in the index system. By calculating the entropy value, the dispersion degree of a certain index can be judged. The smaller the entropy value, the greater the influence of the index on the index system.

[0074] S31: Data standardization. Standardize the k indicators screened by feature engineering, and its calculation formula is as follows:

[0075]

[0076] Where X i is the value of the i-th sample, max(X i ) is the maximum value of the X feature, and min(X i ) is the minimum value of the X feature.

[0077] S32: Information entropy calculation. Calculate the information entropy of the standardized data, and its calculation formula is as follows:

[0078]

[0079] Among them,

[0080] S33: Weight optimization. Integrate the weights of the RandomForest model and the weights calculated by the entropy weight method to obtain the optimized weights, and its calculation formula is as follows:

[0081]

[0082] Where E i is the information entropy of the i-th feature, and f(x i ) is the weight of the i-th feature calculated by the RF model.

[0083] Finally, use the trained model to evaluate the operation status of the intelligent electricity meter.

Claims

1. An intelligent electricity meter operation status evaluation method based on the entropy weight method and the random forest model, including preprocessing the business system data of the intelligent electricity meter; Build a RandomForest model to generate the preliminary evaluation results of smart meters; obtain the final evaluation results through weight calculation and processing, characterized in that: The evaluation method adds the step of further optimizing the weight of the preliminary calculation of the RandomForest model through the entropy weight method to obtain the final evaluation result. The specific steps of further optimizing the weight are as follows: S31. Perform data standardization processing on multiple indicators screened by feature engineering. Specifically, perform data standardization processing on k indicators screened by feature engineering. The calculation formula is as follows: wherein is the i-th sample value, is the maximum value of the X feature, is the minimum value of the X feature; S32. Calculate the information entropy for the standardized data by means of the entropy weight method. Specifically, calculate the information entropy for the standardized data. The calculation formula is as follows: Among them, ; S33. Fuse the weight initially calculated by the RandomForest model with the information entropy obtained in step 32 to obtain the optimized weight, and then obtain the training model. Evaluate the operating status of the smart meter through the training model; Specifically, fuse the weight of the RandomForest model with the weight calculated by the entropy weight method to obtain the optimized weight. The calculation formula is as follows: where is the information entropy of the i-th feature, is the weight of the i-th feature calculated by the RF model.

2. The intelligent electric meter operation state evaluation method based on the entropy weight method and the random forest model according to claim 1, characterized in that: The specific steps for preprocessing the business system data of the smart meter are: S11. For missing values, if the missing ratio is greater than or equal to 40% and less than or equal to 70%, convert this field into an indicator field, where missing is 1 and non-missing is 0. If the missing ratio is greater than 70%, this field is not included in the model calculation; if the missing ratio is less than 40%, fill the discrete data with the mode and fill the continuous numerical values with the mean. S12. For abnormal data, statistically monitor numerical attributes, calculate the mean and standard deviation of the field values, and identify abnormal fields and abnormal records according to the data interval of each field. S13. For redundant data, based on business logic and technical means, eliminate the problem of approximately duplicate records in the data set, and delete or generate derivative fields for the redundant data.

3. The intelligent electric meter operation state evaluation method based on the entropy weight method and the random forest model according to claim 1, characterized in that: The specific steps for building a RandomForest model and initially calculating the weight are: S21. Perform data fusion on the preprocessed data, including operation data, evaluation data, power outage data, and fault data, to form sample data; S22. Screen out the feature factors with a correlation greater than 0.3, and perform indicator screening in combination with specific business guidelines; S23. For the sample data after data preprocessing, construct the training set and test set required for model training through the following method: Use the equipment ledger data, operation data, fault data, power outage data, and related features affecting the smart meter evaluation system as the model input X, and the smart meter evaluation result as the model output Y to construct the data set (X, Y) required for the model. For the constructed data set (X, Y), randomly extract data at a ratio of training set:test set of 7:3, and finally construct the training set and test set required for the model. S24. Initialize the model parameters. Set the number of weak classifiers to 32; set the criterion for feature splitting to "squared_error"; set the depth of the weak classifier to 64; set the minimum number of samples for splitting to 2; set the maximum number of features for splitting to "auto"; set the out-of-bag score to "True"; set the number of training parallelisms to -1; S25. Random sampling. The total number of training samples is N, where N is a positive integer greater than 1. Each individual decision tree randomly draws n samples with replacement from the N training sets, where n is a positive integer less than N. The probability of each sample being sampled is 1 / N, so the probability of not being sampled is 1 - (1 / N). If a sample is not sampled after N samplings, then 35 - 40% of the data in the training set has not been sampled and can be used as an evaluation data set to evaluate the out-of-bag score of the model; S26. Feature attribute splitting. The number of input features of the training examples is N, and n is much less than N. When splitting at each node of each decision tree, randomly select m input features from M input features, and then select the best one from these m input features for splitting. M is a positive integer greater than 1, and m is a positive integer less than M. For the regression scenario, attribute splitting is based on the sample variance. For the classification scenario, the Gini index is used as the basis for attribute splitting; S27. Construction of n weak classifiers. Combine the n sample data sets obtained by sampling with replacement. Randomly select a certain number of optimal features from all the candidate features of each data subset as the input features of the decision tree. The dimension of the input features is [n * m * k], where n is the n data subsets, m is the number of samples, and k is the number of features. Input the feature subset into the weak classifier to construct n basic decision tree models; S28. Model combination. Combine the n weak classifiers obtained in step S27 and perform summation and averaging; S29. Preliminary calculation of weights. Calculate the change in the Gini index obtained through step S26; S210. Calculation of the scoring index. Normalize the importance score to obtain the final scoring index.

Citation Information

Patent Citations

  • Power grid real-time operation risk assessment method and system based on random forest

    CN111292020A