District line loss acquisition calculation factor abnormity diagnosis method and device based on machine learning

By extracting and processing communication data in low-voltage table areas, using random forest algorithms and linear interpolation methods, the problem of accurate analysis of line loss anomalies in low-voltage table areas is solved, and fast and intelligent fault identification and processing is achieved, improving line loss repair efficiency.

CN120337017APending Publication Date: 2025-07-18JURONG CITY POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510416443.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to accurately analyze the causes of line loss abnormalities in low-voltage table areas, especially abnormalities caused by acquisition and calculation factors, which lead to high line loss and long repair cycles, making it difficult to achieve intelligent and accurate evaluation and processing.

Method used

By extracting the communication radar diagram and collecting the data of the monitoring page, a communication data set for the failed power meter is formed, a random forest algorithm is used for model training, and a linear interpolation method is used to complete the missing data, calculate the correlation coefficient between the user type and the station name, and a Gini coefficient evaluation model is used to realize batch marking and automated processing of the cause of the failed power meter failure.

Benefits of technology

It significantly improves the response speed and accuracy of the model, and can quickly and accurately identify and handle line loss abnormalities caused by acquisition and calculation factors, reducing line high loss and repair cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The invention discloses a machine learning-based transformer area line loss acquisition calculation factor anomaly diagnosis method, which comprises the following steps of: extracting various information of a communication radar map, and forming an acquisition failure ammeter communication data set by various data; based on the collection failure user set, forming a collection failure ammeter electricity utilization information data set on the collection monitoring page; extracting a characteristic value of the failed user data as a data characteristic of machine learning; the data set is divided, an artificial intelligence model is trained, and a model training effect and model tuning are tested; carrying out model training by adopting a machine algorithm random forest, and carrying out model evaluation on a model which is better trained; and the trained model is used to carry out batch marking processing on the actual acquisition failure ammeter fault reasons. Compared with the prior art, the failure electricity meter communication data is summarized and then serves as a training material to train the model, and the response speed of the trained model in practical application is remarkably improved compared with that in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of analysis of abnormal reasons for low-voltage substation area line losses, and specifically relates to the construction of a model for abnormal reasons for low-voltage substation area line losses caused by acquisition and calculation factors based on machine learning. Background Art

[0002] The main reasons for high losses in low-voltage distribution networks are file relationship factors, acquisition and calculation factors, distributed photovoltaic access factors, and technical factors. Among them, the acquisition and calculation factors mainly consist of common factors such as failure of user meter acquisition, abnormal user metering wiring circuits, user transformer failures, and no metering for substation electricity consumption, which lead to high losses in the substation area.

[0003] Regarding the high losses in the substation area caused by acquisition and calculation factors, it mainly relies on on-site inspections by grass-roots employees. Due to factors such as a large number of users, complex equipment, long time consumption, diverse reasons for acquisition failures, and great difficulty in inspections, it is difficult to comprehensively, specifically, and intelligently evaluate the electricity consumption status of each user, thus leading to pain points in business operations such as high line losses, long repair cycles for substation area line loss rates, and low accuracy in reliability analysis, seriously restricting the construction process of a first-class distribution network.

[0004] The invention patent document with publication number CN202110963489 discloses a method for diagnosing abnormal electricity consumption of medium-voltage distribution network users based on machine learning. It is characterized in that the integrated algorithm random forest in the machine learning is used for model training, and the Gini coefficient is used as the division evaluation criterion for the sub-CART trees in the random forest for the trained model. The model evaluation indexes are accuracy, precision, recall, F1 score, and ROC value. However, the eigenvalue adopted by this model is mainly the correlation coefficient of the average value of forward active power and voltage value, etc., which is slightly insufficient for communication data, user types, file classification, etc. that affect line losses, and it is impossible to accurately analyze various types and large quantities of abnormal reasons for line losses in low-voltage substation areas. Summary of the Invention

[0005] The purpose of the present invention is to accurately judge the anomalies caused by acquisition and calculation factors in the abnormal line losses of low-voltage substation areas. By collecting datasets such as communication data and user types of relevant meters that affect the acquisition quality in line loss abnormal substation areas as eigenvalues, using the random forest algorithm, classifying and analyzing the abnormal types of acquisition and calculation factor failures, and predicting and classifying them, so as to achieve early discovery, early solution, and active processing.

[0006] The present invention is implemented by the following technical means: A method for diagnosing abnormal acquisition and calculation factors of substation area line losses based on machine learning, the method comprising:

[0007] Extract various information from the communication radar chart, where the information includes various data such as the terminal signal field strength, the collector parallel carrier table, the communication timeout situation, the collector meter hanging situation, etc., to form a set of communication data of the failed meter for collection;

[0008] Based on the set of users with collection failures, extract datasets such as the power bureau number, terminal asset number, user type, and the district manager to which they belong from the collection monitoring page to form a dataset of electricity consumption information of the failed meters for collection, and summarize and preprocess each piece of data;

[0009] Extract the characteristic values of the data of the users with collection failures as the data characteristics for machine learning, and add the three types of labels for the faulty meters, transmission lines, and master station terminals to the historical data;

[0010] Divide the dataset, train the artificial intelligence model, and test the training effect of the model and optimize the model;

[0011] Use the ensemble algorithm random forest in machine learning to train the model, and evaluate the model with a better training effect;

[0012] Use the trained model to perform batch marking on the reasons for the failures of the actual meters for collection;

[0013] The preprocessing of the communication data of the failed meters for collection includes: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data. The specific steps are as follows:

[0014] Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar chart data. If there are, go to Step 2; if not, go to the next step;

[0015] Step 2: If there are missing data, use the linear interpolation method to complete the communication radar chart data collected originally, and obtain the communication and user data after data completion;

[0016] The linear interpolation method is to complete the originally collected communication data so that the interpolation function can approximately replace the original function. The interpolation function is a first-degree polynomial class, and it is required that the interpolation error at each interpolation node is 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the originally sampled data. Now, a function is constructed by linear interpolation such that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is:

[0017]

[0018] Now, based on the constructed interpolation function If there is a data missing situation in the original data at i = m, that is, f(m) is a null value, then Complete the data missing situation of the original sampling.

[0019] The calculation method of the eigenvalue of the correlation coefficient between the user type and the substation area name is as follows:

[0020] Find the frequent item sets by searching layer by layer, and calculate the support (the probability of co-occurrence) and confidence (the probability of B occurring when A occurs):

[0021] Support: Support(A→B) = P(A∩B);

[0022] Confidence: Confidence(A→B) = P(B|A);

[0023] Lift: The value > 1 indicates a positive correlation;

[0024] The steps of partitioning the data set, training the artificial intelligence model, and testing the model training effect and model tuning are as follows:

[0025] Partition the data set according to 7:3. 70% is used as the training set to train the artificial intelligence model, and 30% is used as the validation set to test the model training effect and model tuning.

[0026] The method of using the ensemble algorithm random forest in machine learning for model training and model evaluation of the trained model is as follows:

[0027] Use the Gini coefficient as the partitioning evaluation criterion for the sub-CART trees in the random forest. The model evaluation indicators are accuracy, precision, recall, F1 score, and ROC value.

[0028] The steps of using the trained model to perform batch marking processing on the electricity users in the actual distribution network are as follows:

[0029] After each index meets the requirements, use the trained model to perform automated batch marking processing on the users who actually fail to collect data daily, that is, perform batch processing on the electricity meters that fail to collect data daily for each power supply station and calculate the required characteristic values, send them into the model for calculation, and finally output the fault reasons and recommended repair methods, mainly including three categories of meter problems, transmission lines, and master station terminals, as well as special categories such as regional power outages.

[0030] A device for abnormal diagnosis of factors in the collection and calculation of substation area line losses based on machine learning, including:

[0031] Extraction module: used to extract radar chart information and relevant data on the collection monitoring page;

[0032] Summary module: used to summarize the radar chart information and the monitoring page data for subsequent use;

[0033] Preprocessing module: used to preprocess the data of the summary module;

[0034] Training module: used to train the data and give relevant result parameters;

[0035] Labeling module: used to label the data obtained from training.

[0036] The communication radar chart information extracted by the extraction module includes various data such as terminal signal field strength, collector parallel carrier table, communication timeout situation, and collector meter hanging situation, forming a set of communication data of failed electricity meters. The data extracted from the acquisition monitoring page includes substation number, terminal asset number, user type, and name of the district manager in charge of the substation area.

[0037] The preprocessing of the communication data of failed electricity meters by the preprocessing module includes: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data. The specific steps are as follows:

[0038] Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar chart data. If there are, go to Step 2; if not, go to the next step, that is, extract the feature values and add three types of labels: faulty meter, transmission line, and master station terminal;

[0039] Step 2: If there are missing data, use the linear interpolation method to complete the original collected communication radar chart data to obtain the communication and user data after data completion;

[0040] The linear interpolation method is to complete the original collected communication data so that the interpolation function can approximately replace the original function. The interpolation function is a first-degree polynomial class, and it is required that the interpolation error at each interpolation node is 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the original sampled data. Now, a function is constructed by linear interpolation so that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is:

[0041]

[0042] Now, based on the constructed interpolation function If there is a data missing situation at i = m in the original data, that is, f(m) is a null value, then complete the data missing situation of the original sample.

[0043] The calculation method of the correlation coefficient eigenvalue between the user type and the substation area name by the preprocessing module is as follows:

[0044] Find the frequent item sets by searching layer by layer, and calculate the support, that is, the probability of co-occurrence, and the confidence, that is, the probability of the substation area name appearing when the user type appears:

[0045] Support: Support(A→B) = P(A∩B);

[0046] Confidence: Confidence(A→B) = P(B|A);

[0047] Lift: The value > 1 indicates a positive correlation, where A is the substation area name and B is the user type.

[0048] The training module divides the data set in a ratio of 7:3. 70% is used as the training set to train the artificial intelligence model, and 30% is used as the validation set to test the training effect of the model and optimize the model.

[0049] The training module uses the Gini coefficient as the division evaluation criterion for the sub-CART trees in the random forest to train the random forest model, and evaluates the trained model. The evaluation indicators are accuracy, precision, recall, F1 score, and ROC value.

[0050] Among them, the Gini coefficient formula is as follows:

[0051]

[0052] P k is the proportion of category k in the data set D. D is the set of communication data of failed meters, and category k is 3 category labels and special category labels.

[0053] The labeling module uses the trained model to perform batch labeling processing on the electricity users in the actual distribution network, including the following steps:

[0054] Batch process the daily failed meters collected by each power supply station and calculate the required eigenvalues, send them into the model for calculation, and finally output the fault cause and recommended repair method. The fault causes include meter problems, transmission line problems, master station terminal problems, and special problems such as regional power outages.

[0055] Compared with the prior art, the present invention aggregates the communication data of failed meters and then uses it as training material to train the model. The trained model has a significant improvement in response speed in actual applications compared with before. Brief Description of the Drawings

[0056] Figure 1It is a flowchart of a method for diagnosing anomalies in factors for collecting and calculating line losses in a distribution transformer area based on machine learning according to the present invention.

[0057] Figure 2 It is a communication radar data diagram of a method for diagnosing anomalies in factors for collecting and calculating line losses in a distribution transformer area based on machine learning according to the present invention.

[0058] Figure 3 It is an intuitive data set formed by the results of a training model of a method for diagnosing anomalies in factors for collecting and calculating line losses in a distribution transformer area based on machine learning according to the present invention. Detailed implementation manners

[0059] The present invention will be further described in detail below with reference to the accompanying drawings of the specification:

[0060] The present invention relates to a method for diagnosing anomalies in factors for collecting and calculating line losses in a distribution transformer area based on machine learning, and the method includes:

[0061] Extracting various information from the communication radar diagram, where the information includes various data such as terminal signal field strength, collector-connected carrier meters, communication timeout situations, collector-meter hanging situations, etc. to form a data set;

[0062] Based on the set of users with collection failures, extracting data sets such as power bureau numbers, terminal asset numbers, user types, and affiliated distribution transformer area managers on the collection monitoring page, and summarizing and preprocessing each data item;

[0063] Extracting the characteristic values of the data of users with collection failures as data characteristics for machine learning, and adding three types of labels for faulty meters, transmission lines, and master station terminals to the historical data;

[0064] Dividing the data set, training an artificial intelligence model, and testing the training effect of the model and optimizing the model;

[0065] Using the ensemble algorithm random forest in machine learning to train the model, and evaluating the model with a better-trained model;

[0066] Using the trained model to perform batch labeling processing on the reasons for failures in actual collected meters.

[0067] The preprocessing of the communication data of the meters with collection failures includes: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data, and the specific steps are as follows:

[0068] Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar diagram data. If there are, go to Step 2; if not, go to the next step;

[0069] Step 2: If there are missing data, the linear interpolation method is used to complete the original collected communication radar map data to obtain the communication and user data after data completion;

[0070] The linear interpolation method is to complete the original collected communication data so that the interpolation function can approximately replace the original function. The interpolation function is a first-degree polynomial class, and it is required that the interpolation error at each interpolation node is 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the original sampled data. Now, a function is constructed by linear interpolation such that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is:

[0071]

[0072] Now, based on the constructed interpolation function If there is a data missing situation at i = m in the original data, that is, f(m) is a null value, then complete the data missing situation of the original sampling.

[0073] The calculation method of the correlation coefficient eigenvalue between the user type and the substation area name is as follows:

[0074] Find the frequent item sets by layer-by-layer search, and calculate the support (the probability of co-occurrence) and confidence (the probability that B appears when A appears):

[0075] Support: Support(A→B) = P(A∩B);

[0076] Confidence: Confidence(A→B) = P(B|A);

[0077] Lift: A value > 1 indicates positive correlation;

[0078] The steps of partitioning the data set, training the artificial intelligence model, and testing the model training effect and model tuning are as follows:

[0079] Partition the data set according to 7:3. 70% is used as the training set to train the artificial intelligence model, and 30% is used as the validation set to test the model training effect and model tuning.

[0080] The method of using the ensemble algorithm random forest in machine learning for model training and evaluating the trained model is as follows:

[0081] Use the Gini coefficient as the partitioning evaluation criterion for the sub-CART trees in the random forest. The model evaluation indicators are accuracy, precision, recall, F1 score, and ROC value.

[0082] Using the trained model to perform batch marking on electricity users in the actual distribution network, including the following steps:

[0083] After all indicators meet the requirements, use the trained model to perform automated batch marking on the users with failed daily collection in the actual situation, that is, batch process the meters with failed daily collection in each power supply station, calculate the required characteristic values, send them into the model for calculation, and finally output the cause of the fault and the recommended repair method, mainly including three categories: meter problems, transmission lines, and master station terminals, as well as special categories such as regional power outages.

[0084] The present invention also relates to a device for diagnosing anomalies in factors for collecting and calculating line losses in a transformer substation area based on machine learning, which includes:

[0085] Extraction module: used to extract radar chart information and relevant data on the collection monitoring page;

[0086] Summarization module: used to summarize the radar chart information and monitoring page data for subsequent use;

[0087] Preprocessing module: used to preprocess the data of the summarization module;

[0088] Training module: used to train the data and give relevant result parameters;

[0089] Marking module: used to mark the data obtained from training.

[0090] The communication radar chart information extracted by the extraction module includes various data such as terminal signal field strength, collector-connected carrier meters, communication timeout situations, and collector-meter hanging situations to form a communication data set of meters with failed collection. The data extracted from the collection monitoring page includes power bureau number, terminal asset number, user type, and name of the district manager to which it belongs.

[0091] The preprocessing module preprocesses the communication data of failed meters, including: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data. The specific steps are as follows:

[0092] Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar chart data. If there are, go to Step 2; if not, go to the next step, that is, extract the characteristic values and add three types of labels: fault meters, transmission lines, and master station terminals;

[0093] Step 2: If there are missing data, use the linear interpolation method to complete the communication radar chart data collected originally to obtain the communication and user data after data completion;

[0094] The linear interpolation method is used to complement the originally collected communication data, so that the interpolation function can approximately replace the original function. The interpolation function is of the first polynomial type, and the interpolation error at each interpolation node is required to be 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the originally sampled data. Now, a function is constructed by linear interpolation such that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is:

[0095]

[0096] Now, based on the constructed interpolation function If there is a data missing situation in the original data at i = m, that is, f(m) is a null value, then Complement the data missing situation in the originally sampled data.

[0097] The calculation method of the correlation coefficient eigenvalue between the user type and the substation area name by the preprocessing module is as follows:

[0098] Find the frequent item sets by layer-by-layer search, and calculate the support degree, that is, the probability of co-occurrence, and the confidence degree, that is, the probability of the substation area name appearing when the user type appears:

[0099] Support degree: Support(A→B) = P(A∩B);

[0100] Confidence degree: Confidence(A→B) = P(B|A);

[0101] Lift degree: The value > 1 indicates positive correlation, where A is the substation area name and B is the user type.

[0102] The training module divides the data set into 7:3. 70% is used as the training set to train the artificial intelligence model, and 30% is used as the validation set to test the training effect of the model and optimize the model.

[0103] The training module uses the Gini coefficient as the division evaluation criterion for the sub-CART trees in the random forest to train the random forest model, and evaluates the trained model. The evaluation indicators are accuracy, precision, recall, F1 score, and ROC value.

[0104] Among them, the Gini coefficient formula is as follows:

[0105]

[0106] P k is the proportion of category k in the data set D. D is the set of communication data of the electricity meters with collection failures, and category k is 3 types of labels and special labels.

[0107] The marking module uses the trained model to perform batch marking processing on electricity users in the actual distribution network, including the following steps:

[0108] Batch process the daily meter reading failure meters of each power supply station, calculate the required characteristic values, send them into the model for calculation, and finally output the cause of the failure and the recommended repair method. The causes of the failure include special problems such as meter problems, transmission line problems, master station terminal problems, and regional power outages.

Claims

1. A method for diagnosing anomalies in factors for collecting and calculating line losses in a transformer substation area based on machine learning, characterized in that, The method includes: Extracting various information from the communication radar chart, where the information includes data such as the terminal signal field strength, the collector parallel-connected carrier table, the communication timeout situation, and the collector meter hanging situation to form a set of communication data of meters with collection failures, and preprocessing the communication data of meters with collection failures; Based on the set of users with collection failures, extracting data sets such as the power bureau number, terminal asset number, user type, and name of the district manager to which they belong from the collection monitoring page and forming a set of electricity consumption information data of meters with collection failures, and summarizing and preprocessing each data item; Extracting the eigenvalue of the data of users with collection failures as the data feature of machine learning, and adding three types of labels, namely meter failure, transmission line failure, and master station terminal failure, and special class labels for the abnormal reasons of collection calculation factors to the historical data; Associating the user type with the name of the district in the data of users with collection failures to construct a correlation coefficient for special types of collection failures such as regional power outages. Dividing the data set, training an artificial intelligence model, and testing the training effect and model tuning of the model; Using the ensemble algorithm random forest in machine learning to train the model and evaluating the trained better model.

2. The abnormal diagnosis method for collection and calculation factors of substation line loss based on machine learning according to claim 1, characterized in that The preprocessing of the communication data of meters with collection failures includes: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data. The specific steps are as follows: Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar chart data. If there are, go to Step 2; if not, go to the next step, that is, extract the eigenvalue and add three types of labels: faulty meter, transmission line, and master station terminal; Step 2: If there are missing data, use the linear interpolation method to complete the original collected communication radar chart data to obtain the communication and user data after data completion; The linear interpolation method is used to complete the original collected communication data, so that the interpolation function can approximately replace the original function. The interpolation function is a first-degree polynomial class, and the interpolation error at each interpolation node is required to be 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the original sampled data. Now, a function is constructed by linear interpolation such that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is: Based on the constructed interpolation function If there is a data missing situation in the original data at i = m, that is, f(m) is a null value, then Complete the data missing situation of the original sampling.

3. The abnormal diagnosis method for factors of substation line loss acquisition and calculation based on machine learning according to claim 1, characterized in that The calculation method of the correlation coefficient eigenvalue between the user type and the name of the district is: Finding the frequent item sets by layer-by-layer search, and calculating the support degree, that is, the probability of simultaneous occurrence, and the confidence degree, that is, the probability of the name of the district appearing when the user type appears: Support degree: Support(A→B) = P(A∩B); Confidence degree: Confidence(A→B) = P(B|A); Lift: A value > 1 indicates a positive correlation, where A is the name of the substation area and B is the user type.

4. The abnormal diagnosis method for factors in the collection and calculation of substation area line loss based on machine learning according to claim 1, characterized in that, Dividing the data set according to 7:3, using 70% as the training set to train the artificial intelligence model, and 30% as the validation set to test the training effect and model tuning of the model.

5. A method for diagnosing anomalies in factors for collecting and calculating substation line losses based on machine learning according to claim 1, characterized in that, Using the Gini coefficient as the division evaluation criterion for the sub-CART trees in the random forest to train the random forest, and evaluating the trained model. The evaluation indicators are accuracy, precision, recall, F1 score, and ROC value. The specific steps are as follows: Step 1: Construct a single CART tree and use the Gini coefficient for division Among them, the Gini coefficient formula is as follows: P k is the proportion of class k in the dataset D, where D is the set of communication data of electricity meters with acquisition failures, and class k is three types of labels and special class labels. Node splitting rule:

1. For each possible splitting point of each feature, calculate the Gini coefficient after splitting; 2. Select the feature and splitting point with the smallest Gini coefficient (that is, the largest increase in purity); 3. Recursively generate the left and right subtrees until the stop condition is met, that is, the maximum depth is 10; Step 2: Random forest integration, training multiple trees (300 trees), and obtaining the final result through majority voting during prediction. Among them, each tree performs bootstrap sampling with replacement from the training set, and only considers the features of a random subset sqrt when each tree is split.

6. The abnormal diagnosis method for factors in the collection and calculation of substation area line loss based on machine learning according to claim 1, characterized in that, The batch marking process of electricity users in the actual distribution network using the trained model includes the following steps: Batch process the daily failed meters collected by each power supply station, calculate the required feature values, send them into the model for calculation, and finally output the cause of the failure and recommended repair methods. The causes of the failure include meter problems, transmission line problems, master station terminal problems, and special problems such as regional power outages.

7. A device for diagnosing anomalies in factors for collecting and calculating line losses in a transformer substation area based on machine learning as described in claim 1, characterized in that, It includes: Extraction module: used to extract radar chart information and relevant data on the acquisition monitoring page; Summarization module: used to summarize the radar chart information and monitoring page data for subsequent use; Preprocessing module: used to preprocess the data of the summarization module; Training module: used to train the data and give relevant result parameters; Marking module: used to mark the data obtained from training.

8. The abnormal diagnosis device for factors in collection and calculation of substation line loss based on machine learning according to claim 7, characterized in that The communication radar chart information extracted by the extraction module includes various data such as terminal signal field strength, collector parallel-connected carrier meters, communication timeout situations, and collector meter hanging situations to form a communication data set of failed meters. The data extracted from the acquisition monitoring page includes power bureau number, terminal asset number, user type, and name of the district manager to which it belongs.

9. The abnormal diagnosis device for factors of substation area line loss acquisition and calculation based on machine learning according to claim 7, characterized in that, The preprocessing module's preprocessing of the communication data of failed meters includes: using the linear interpolation method to interpolate and complete the missing values in the obtained feeder and user data. The specific steps are as follows: Step 1: After extracting the data, check whether there are missing values in the exported communication situation radar chart data. If there are, go to Step 2; if not, go to the next step, that is, extract the feature values and add three types of labels: faulty meters, transmission lines, and master station terminals. Step 2: If there are missing data, use the linear interpolation method to complete the communication radar chart data collected originally to obtain the communication and user data after data completion. The linear interpolation method is used to complete the original collected communication data, so that the interpolation function can approximately replace the original function. The interpolation function is a first-degree polynomial class, and the interpolation error at each interpolation node is required to be 0. Given the original function f(x i ), where xi (i = 0, 1, 2, 3,..., n), and n is the length of the original sampled data. Now, a function is constructed by linear interpolation such that the absolute value of the error |R(x)| is relatively small over the entire original data interval, that is: Based on the constructed interpolation function If there is a data missing situation in the original data at i = m, that is, f(m) is a null value, then Complete the data missing situation of the original sampling.

10. The abnormal diagnosis device for factors in the collection and calculation of substation line loss based on machine learning according to claim 7, characterized in that The calculation method of the correlation coefficient feature value between the user type and the district name by the preprocessing module is: Find the frequent item sets by layer-by-layer search, and calculate the support degree, that is, the probability of co-occurrence, and the confidence degree, that is, the probability of the district name appearing when the user type appears: Support degree: Support(A→B) = P(A∩B); Confidence degree: Confidence(A→B) = P(B|A); Lift: A value > 1 indicates a positive correlation, where A is the name of the substation area and B is the user type.

11. The abnormal diagnosis device for factors in the collection and calculation of substation line losses based on machine learning according to claim 7, characterized in that, The training module divides the data set into 7:

3. 70% is used as the training set to train the artificial intelligence model, and 30% is used as the validation set to test the training effect of the model and optimize the model.

12. The abnormal diagnosis device for factors in the collection and calculation of substation line loss based on machine learning according to claim 7, characterized in that, The training module uses the Gini coefficient as the division evaluation criterion for the sub-CART trees in the random forest to train the random forest model, and evaluates the trained model. The evaluation indicators are accuracy, precision, recall, F1 score, and ROC value. Among them, the Gini coefficient formula is as follows: P k is the proportion of class k in the dataset D, where D is the set of communication data of electricity meters with acquisition failures, and class k is three types of labels and special class labels.

13. The abnormal diagnosis device for factors of substation area line loss acquisition and calculation based on machine learning according to claim 7, characterized in that, The marking module uses the trained model to perform batch marking processing on electricity users in the actual distribution network, including the following steps: Batch process the daily meter reading failure meters of each power supply station, calculate the required characteristic values, send them into the model for calculation, and finally output the cause of the failure and the recommended repair method. The causes of the failure include meter problems, transmission line problems, master station terminal problems, and special problems such as regional power outages.

Citation Information

Patent Citations

  • A method for diagnosing abnormal power consumption of medium voltage distribution network users based on machine learning

    CN113866552B