Intelligent Line Loss Diagnosis Method and Device Based on Comprehensive Analysis of Machine Learning Methods
Through a comprehensive analysis method based on machine learning, a variety of line loss analysis models are constructed, combined with SHAP algorithm and Bayesian network model, the existing line loss intelligent diagnosis methods are solved in terms of accuracy, flexibility and adaptability, and more efficient line loss diagnosis and grid operation efficiency improvement are achieved.
Patent Information
- Application Number
- CN202411759352.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing intelligent linear loss diagnosis methods have shortcomings in accuracy, flexibility and adaptability, making it difficult to effectively diagnose complex linear loss problems, especially in complex and changeable situations in power grid operations.
A comprehensive analysis method based on machine learning methods is adopted, including building a classification model for unqualified characteristics in the table area, a main cause analysis model for unqualified line loss, a main cause analysis model for abnormal correlation, and an abnormal quantization impact analysis model. Multi-dimensional and multi-level line loss analysis is carried out through technical means such as SHAP algorithm and Bayesian network model.
It improves the accuracy, flexibility and adaptability of line loss diagnosis, and can efficiently focus on solving key abnormal problems when resources are limited, reduce line loss rate and improve grid operation efficiency.
Smart Images

Figure CN119226962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid monitoring, and in particular to a line loss intelligent diagnosis method and device based on comprehensive analysis using a machine learning method. Background Art
[0002] The line loss problem has always been a focus of attention in the power industry. With the continuous expansion of the scale of the power grid and the continuous growth of electricity demand, the line loss problem has become increasingly prominent, directly affecting the economic benefits and operating efficiency of the power grid. Accurate diagnosis and analysis of the causes of line loss are of great significance for reducing the line loss rate and improving the operating efficiency of the power grid. Therefore, the development of advanced line loss intelligent diagnosis algorithms has become a hot area in current power system research.
[0003] Traditional line loss analysis methods mainly rely on manual experience and simple statistical analysis. These methods have problems such as low analysis efficiency, lack of accuracy and objectivity, and difficulty in discovering complex nonlinear relationships. With the advent of the big data era, traditional methods can no longer meet the needs of real-time dynamic analysis of massive data, and new technical means are urgently needed to improve the efficiency and accuracy of line loss diagnosis.
[0004] In recent years, machine learning methods have been widely used in power system analysis. Decision trees and clustering algorithms have been applied to line loss prediction, pattern recognition, influencing factor analysis, and feature extraction. These methods have improved the efficiency and accuracy of line loss analysis to a certain extent, but a single algorithm is often difficult to fully and accurately diagnose complex line loss problems. In view of the complexity of line loss problems, it has become an inevitable trend to organically combine multiple machine learning methods.
[0005] The comprehensive application of multiple algorithms can give full play to the advantages of various algorithms, analyze line loss characteristics from different angles, improve the reliability of results through cross-validation, and realize multi-dimensional and multi-level analysis of line loss problems. This comprehensive research and judgment method is expected to break through the limitations of a single algorithm and provide a more comprehensive and accurate solution for intelligent diagnosis of line loss. However, in practical applications, intelligent diagnosis algorithms for line loss based on multiple machine learning methods still face many challenges.
[0006] The diagnostic results of the existing abnormal analysis and diagnosis models are relatively simple and cannot fully reflect the complex and changeable situations in power grid operations. They perform poorly when dealing with complex and changeable on-site conditions, and the diagnostic accuracy and reliability in complex scenarios are insufficient. The existing models are obviously insufficient in identifying nonlinear relationships in data. There are complex nonlinear relationships between various factors in the power grid system, and traditional models are mostly based on linear assumptions. It is difficult to accurately capture these nonlinear characteristics, thus affecting the accuracy of diagnosis. The existing diagnostic strategies lack sufficient intelligence and adaptability, and it is difficult to meet the growing demand for refined management. Summary of the invention
[0007] The purpose of the present invention is to overcome the defects of the prior art and propose an intelligent line loss diagnosis method based on comprehensive analysis of machine learning methods, which can improve the accuracy, flexibility and adaptability of line loss diagnosis.
[0008] To achieve the above object, the present invention adopts the following specific technical solutions:
[0009] The intelligent line loss diagnosis method based on comprehensive analysis of machine learning methods provided by the present invention includes the following steps:
[0010] S1. Construct a classification model for unqualified characteristics of the substation area based on line loss volatility and daily line loss rate, including dividing the daily line loss rate types according to the daily line loss rate results of the substation area and the theoretical value of one index per substation area, and calculating the volatility of the substation area according to the daily line loss results of the substation area, for classifying abnormal substation areas;
[0011] S2. Construct a main cause analysis model for unqualified line loss, including model training, abnormal sorting, and summary of sorting results, for obtaining the abnormal causes of unqualified line loss in the substation area;
[0012] S3. Construct an abnormal correlation main cause analysis model, and analyze the relationship between abnormal causes by constructing a Bayesian network model, for locating the main abnormal causes;
[0013] S4. Construct an abnormal quantitative impact analysis model, including an abnormal user impact quantification model and an abnormal substation area impact quantification model, for locating the main abnormal users and main abnormal substation areas to be governed.
[0014] Further, in step S1, dividing the daily line loss rate types according to the daily line loss rate results of the substation area and the theoretical value of one index per substation area includes ultra-high loss, high loss, small negative loss, negative loss, and non-calculable;
[0015] Ultra-high loss: daily line loss rate ≥ threshold upper limit + 20%, high loss: threshold upper limit ≤ daily line loss rate < threshold upper limit + 20%, small negative loss: -1% ≤ daily line loss rate < 0%, negative loss: daily line loss rate < -1%;
[0016] Calculating the volatility of the substation area according to the daily line loss results of the substation area includes stable, fluctuating, and sudden;
[0017] Stable: number of abnormal line loss days in a week ≥ 3 and abnormal discrete coefficient in the recent 60 days < 3, fluctuating: 1 ≤ number of abnormal line loss days in a week < 3 and abnormal discrete coefficient in the recent 60 days ≥ 3, sudden: abnormal line loss on the current day and not a stable or fluctuating substation area;
[0018] The calculation formula for the abnormal discrete coefficient is:
[0019] ,
[0020] where It represents the line loss rate of the i-th day in the substation area. N represents the data period required for analyzing the volatility of the substation area, with a default value of 60 days. The abnormal dispersion coefficient The higher it is, the greater the abnormal fluctuation of the daily line loss in the substation area.
[0021] Furthermore, in step S2, the model training is based on the unqualified characteristic classification model and the abnormal factor library of the substation area. For each type of substation area, the Pearson correlation coefficient and decision tree regression are combined to screen abnormal factors, improving the efficiency and stability of model training. The integrated learning XGBoost algorithm is used to train the model to achieve abnormal substation area line loss prediction, specifically as follows:
[0022] Define the training set as a data set containing m features n, , and the predicted output of the model trained by the integrated learning XGBoost algorithm is expressed as ; where, represents the overall prediction function of the model, which is composed of the sum of the predicted outputs of all K trees, represents the regression tree, K is the number of regression trees, represents the feature vector of an input sample, and the predicted value is the sum of the predicted values of K regression trees;
[0023] The objective function of XGBoost adds a regularization term to the GBDT loss function , that is, the sum of the regularization values of all regression trees. The objective function of the th tree is:
[0024] ;
[0025] where, is the L1 regularization coefficient of the number of leaf nodes, is the number of leaf nodes of the tree, is the L2 regularization coefficient of the leaf node weights, is the weight of the j-th leaf node of the t-th tree; is the actual line loss rate of the th sample, is the sum of the predicted values before the th tree, is the th sample's feature vector, is the predicted value of the th tree.
[0026] Furthermore, in step S2, the abnormal sorting constructs a quantitative abnormal sorting model based on the SHAP algorithm, calculates the marginal contribution of each feature factor to the line loss rate of the substation area, and quantifies the influence weight of each feature factor on the line loss rate size, specifically as follows:
[0027] ;
[0028] wherein, represents the predicted value of the line loss in the substation area, is a linear function of the SHAP value, indicating mapping the SHAP value to the predicted result of the model, is the baseline value, representing the predicted result of the model without the influence of features, represents the SHAP value of the i-th feature, indicating the contribution of this feature to the predicted result of the model, , indicating how many features are included in the decision path of the sample among all n features; for a certain sample, if feature k is not in its decision path, then the SHAP value of the corresponding feature is 0, that is , indicating that this feature k will not attribute to the sample and has no contribution to the final predicted value;
[0029] The calculation formula for the contribution value of the substation area feature factor is as follows:
[0030]
[0031] wherein, is the contribution of the th feature factor to the prediction of the line loss in the substation area, which is the weighted sum and summation over all possible combinations of feature values; is the number of feature factors, is a subset of the feature factors applied by the model, is the vector of the feature values of all instances, is the feature factor not included in the is the prediction of the union features of the set and the feature factor , is the prediction of the feature values in the set, which is marginalized over the features not included in the set, is the combination number, representing the combination number of selecting ∣S∣ features from n−1 features.
[0032] Furthermore, in step S2, the sorting result summary obtains the importance degree of each abnormal cause in each substation area class according to all the calculated abnormal causes in the substation area, and standardizes all the important abnormalities in the substation area to obtain the abnormal weight. The specific formula is as follows:
[0033] ;
[0034] wherein, Indicates the importance degree of the abnormal cause i Indicates the total number of the e-th type of distribution transformers Indicates the number of times the abnormal cause i is important in the e-th type of distribution transformers
[0035] Further, in step S3, the Bayesian network model is an indeterministic causal association model. The influence degree of the anomaly X on the anomaly Y under the Bayesian network is defined as the probability of Y when X occurs divided by the probability of Y when X does not occur. The formula is as follows:
[0036] ;
[0037] The Bayesian network model can learn different combinations of the anomaly X on the anomaly Y under different anomaly combinations group(a), and the occurrence conditions of other anomalies others(o) outside the anomaly group group(a) are unknown
[0038] Further, in step S4, the abnormal user impact quantification model analyzes the abnormal value of the historical data of the user, replaces the data exceeding 1.5 times the normal range with the upper and lower bounds of the normal range; uses the exponential smoothing algorithm to estimate the power consumption of the user when the anomaly occurs as the power loss reduction space of the user; for users that cannot converge, uses the average value of the same-period power consumption as the power loss reduction space of the user
[0039] The abnormal distribution transformer impact quantification model obtains all the abnormal user impact power data of the distribution transformers with unqualified line losses according to the output result of the abnormal user impact quantification; calculates the impact quantification of the terminal and the distribution transformer by referring to the abnormal user impact quantification rule, and obtains all the abnormal terminal impact power data and the distribution transformer impact power data of the distribution transformers with unqualified line losses; obtains all the abnormal impact power of the distribution transformers with unqualified line losses; obtains the power loss reduction space of the distribution transformers with unqualified line losses for one distribution transformer and one index
[0040] The present invention also provides a line loss intelligent diagnosis device based on comprehensive analysis by a machine learning method, adopting the above-mentioned line loss intelligent diagnosis method based on comprehensive analysis by a machine learning method. The line loss intelligent diagnosis device includes:
[0041] The unqualified characteristic classification module of the distribution transformer classifies the daily line loss rate type by dividing the daily line loss rate according to the daily line loss rate result of the distribution transformer and the theoretical value of one distribution transformer and one index, and calculates the volatility of the distribution transformer according to the daily line loss result of the distribution transformer, so as to classify the abnormal distribution transformers
[0042] The main cause analysis module of the unqualified line loss obtains the abnormal causes of the unqualified line loss of the distribution transformer through the integrated learning model and the SHAP algorithm quantification model
[0043] The main cause analysis module of the abnormal association analyzes the relationship between the abnormal causes through the Bayesian network model, so as to locate the main cause of the abnormal association
[0044] An abnormal quantization impact analysis module, through an abnormal user impact quantization model and an abnormal substation area impact quantization model, is used to locate abnormal users and abnormal substation areas for governance.
[0045] The present invention can achieve the following technical effects:
[0046] By constructing models such as substation area unqualified characteristic classification, main cause analysis of line loss unqualified, main cause analysis of abnormal association, and abnormal quantization impact analysis, applying machine learning models such as the SHAP algorithm analysis model and the Bayesian network model, and combining the experience of business experts, the present invention forms a set of abnormal quantization evaluation and optimal governance solutions for substation area line loss, which can efficiently focus on and solve key abnormal problems on the premise of limited resources, so as to achieve the goals of reducing line loss and improving efficiency. Description of the Drawings
[0047] Figure 1 It is a schematic flowchart of a line loss intelligent diagnosis method based on comprehensive analysis by a machine learning method according to an embodiment of the present invention. Detailed Embodiment
[0048] In the following, embodiments of the present invention will be described with reference to the drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.
[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and do not constitute a limitation to the present invention.
[0050] An embodiment of the present invention provides a line loss intelligent diagnosis method based on comprehensive analysis by a machine learning method. The process is as Figure 1 shown, and includes the following steps:
[0051] 1. Construct a substation area unqualified characteristic classification model.
[0052] Since the abnormalities in the substation area have their own characteristics, a classification modeling method is adopted to reduce the model complexity and the model training calculation amount, and at the same time, the application efficiency of the model is improved according to the characteristics of different abnormal substation areas. Therefore, a substation area unqualified characteristic classification model is constructed based on line loss volatility and daily line loss rate, and each abnormal substation area is classified. The substation area abnormalities are divided into 15 categories, specifically as follows:
[0053] 11. Divide the daily line loss rate types according to the daily line loss rate result of the substation area and the theoretical value of one index per substation area, specifically including:
[0054] 111) Ultra-high loss: The daily line loss rate ≥ threshold upper limit + 20%;
[0055] 112) High loss: Threshold upper limit ≤ daily line loss rate < threshold upper limit + 20%;
[0056] 113) Small negative loss: -1% ≤ daily line loss rate < 0%;
[0057] 114) Negative loss: Daily line loss rate < -1%;
[0058] 115) Unable to calculate.
[0059] 12. Calculate the volatility of the transformer substation area based on the daily line loss results of the transformer substation area, specifically including:
[0060] 121) Stable: The number of days with abnormal line loss within a week ≥ 3 and the abnormal dispersion coefficient in the recent 60 days < 3;
[0061] 122) Fluctuating: 1 ≤ the number of days with abnormal line loss within a week < 3 and the abnormal dispersion coefficient in the recent 60 days ≥ 3;
[0062] 123) Sudden: When the daily line loss is abnormal and it is not a stable or fluctuating transformer substation area.
[0063] The formula for calculating the abnormal dispersion coefficient is:
[0064] ,
[0065] where, represents the line loss rate of the i-th day of the transformer substation area, N represents the data period required for analyzing the volatility of the transformer substation area, which is defaulted to 60 days, and the higher the abnormal dispersion coefficient the greater the abnormal fluctuation of the daily line loss of the transformer substation area.
[0066] 2. Build an analysis model for the main reasons of unqualified line loss.
[0067] Build an analysis model for the main reasons of unqualified line loss, which is used to locate the main abnormal reasons for unqualified line loss in the transformer substation area, including model training, abnormal sorting, and sorting result summary, specifically as follows:
[0068] 21. Model training.
[0069] Based on the unqualified characteristic classification model of the transformer substation area and the abnormal factor library, for each type of transformer substation area, the method of combining Pearson correlation coefficient and decision tree regression is used to screen abnormal factors, improve the efficiency and stability of model training, and the integrated learning XGBoost algorithm is used to train the model to realize the line loss prediction of abnormal transformer substation areas.
[0070] Among them, the Pearson correlation coefficient is a method that can help understand the relationship between features and the dependent variable. This method measures the linear correlation between variables, and the value range of the result is [-1, 1]. An obvious defect of this method is that it is only sensitive to linear relationships. If the relationship is non-linear, even if two variables have a one-to-one correspondence, the correlation may be close to 0. Decision tree regression is a tree-based method that can reflect the non-linear relationship between features and the dependent variable and is good at modeling non-linear relationships. Each node in the decision tree is a condition about a certain feature, aiming to divide the data set into two according to different dependent variables to determine the division node. For a decision tree, it is possible to calculate how much the impurity is reduced for each feature and use the reduced impurity as the value of feature selection.
[0071] The integrated learning algorithm is used to achieve the prediction of the line loss in the substation area. The XGBoost (Extreme Gradient Boosting) extreme gradient boosting algorithm is an algorithm that efficiently improves the gradient boosting tree (GBDT). It uses the boosting integration idea, including an additive model (the strong learner is composed of a series of weak learners linearly added) and a forward distribution algorithm (the next learner is trained based on the previous learner).
[0072] Define the training set as a data set containing m features and n. , the predicted output of the training model by the integrated learning XGBoost algorithm is expressed as ;
[0073] Among them, represents the overall prediction function of the model, which is composed of the sum of the predicted outputs of all K trees. represents the regression tree, K is the number of regression trees. represents the feature vector of an input sample, and the output predicted value is the sum of the predicted values of K regression trees;
[0074] The objective function of XGBoost adds a regularization term , that is, the sum of the regularization values of all regression trees. The objective function of the th tree is:
[0075] ;
[0076] Among them, is the L1 regularization coefficient of the number of leaf nodes, is the number of leaf nodes of the tree, is the L2 regularization coefficient of the leaf node weights, is the weight of the jth leaf node of the is the actual line loss rate of the th sample, is the sum of the predicted values before the th tree, is the feature vector of the th sample, is the predicted value of the th tree.
[0077] The loss function of the XGBoost algorithm adopts second-order Taylor expansion and adds a regularization term, featuring high accuracy and being less prone to overfitting.
[0078] 22. Abnormal sorting.
[0079] Based on the SHAP algorithm, a quantitative abnormal sorting model is constructed to calculate the marginal contribution of each abnormal feature factor to the line loss rate of the transformer substation area, and to quantify the influence weight of each abnormal feature factor on the magnitude of the line loss rate. The core idea of SHAP is based on the distribution method of Shapley values, mainly referring to that what is obtained should match one's own contribution. Through this idea, we can quantify the marginal contribution of each feature factor to the line loss prediction. Since Shapley satisfies marginality, symmetry, effectiveness, and additivity, it has a solid theoretical foundation, ensuring the fair distribution of the prediction results among the abnormal feature factors. The specific process is as follows:
[0080] 221. Input the feature factors of each predicted data and the line loss rate value obtained from model training and prediction.
[0081] 222. According to the SHAP algorithm, construct an additive feature attribution mathematical expression to calculate the baseline of the entire model when the feature factors do not participate in the calculation (the average predicted value, the bias term of the model itself, which is the sum of the expected values of all tree root nodes).
[0082] The additive feature attribution mathematical expression is as follows:
[0083] ;
[0084] where represents the predicted value of the line loss of the transformer substation area, is a linear function of the SHAP value, indicating mapping the SHAP value to the prediction result of the model, is the baseline value, representing the model prediction result without feature influence, represents the SHAP value of the i-th feature, indicating the contribution of this feature to the model prediction result, , indicating how many features among all n features are included in the decision path of this sample; for a certain sample, if feature k is not in its decision path, then the SHAP value of the corresponding feature is 0, that is It means that the feature k will not contribute to the sample and has no contribution to the final predicted value.
[0085] 223. According to the calculation formula of the contribution value of the substation area feature factor, calculate the contribution value of the feature factor respectively. According to the data distribution and decision tree structure, distribute the difference between the predicted value and the average predicted value fairly among the feature values of each substation area to complete the quantification work of each feature index.
[0086] The calculation formula of the contribution value of the substation area feature factor is as follows:
[0087]
[0088] Among them, is the contribution of the th feature factor to the line loss prediction of the substation area, which is the weighted sum of all possible combinations of feature values; is the number of feature factors, is the subset of feature factors applied by the model, is the vector of feature values of all instances, is the feature factor not included in the set, is the prediction of the union feature of the set and the feature factor , is the prediction of the feature values in the set , which is marginalized on the features not included in the set , is the combination number, representing the combination number of selecting ∣S∣ features from n−1 features.
[0089] After determining the contribution value of the abnormal feature factor to the line loss rate through the above method, sort and output the contribution degree ranking of various abnormalities. Take the features with contribution degree greater than a certain threshold as the main causes of line loss governance abnormalities.
[0090] 23. Summary of sorting results.
[0091] According to all the main causes of abnormalities in each substation area calculated, obtain the importance degree of each abnormality in each substation area category according to the substation area classification. Finally, it is necessary to standardize all important abnormalities under the substation area to obtain the abnormality weight. The specific formula is as follows:
[0092] ;
[0093] Among them, represents the importance degree of the abnormal cause i, represents the total number of substations in the e-th category, represents the number of times the abnormal cause i is important in the e-th category of substations.
[0094] 3. Build an analysis model for the main causes of abnormal associations.
[0095] By building an association rule model between abnormal types, analyze the "causal" relationship between abnormalities. Since the causes of major abnormalities are diverse and much information cannot be recorded, such as the amount of electricity stolen, it is impossible to truly determine the causal relationship between abnormalities. Therefore, an algorithm is used to analyze the association relationship between abnormalities and score the importance of each abnormal cause to help maintenance personnel determine the main cause.
[0096] Specifically, a Bayesian network model is constructed by combining abnormal group data with expert opinions. Analyze the relationship between the causes of electricity consumption abnormalities of each user online and give quantitative analysis results such as the probability of abnormality occurrence and the improvement degree. Different results will be analyzed for different abnormal group models. For the analysis of the main causes of abnormal associations, a Bayesian network model is constructed by combining abnormal group data with expert opinions. Analyze the relationship between the causes of abnormalities in each substation area or user online and give quantitative analysis results such as the probability of abnormality occurrence and the improvement degree. Different results will be analyzed for different abnormal group models.
[0097] The Bayesian network, also known as the belief network, is a probabilistic graphical model. A graphical model can be constructed from data or expert opinions. The model usually has application forms such as predictive analysis, abnormal data detection, and filling in missing data values. Compared with other machine learning models, the Bayesian network is good at giving prediction results in the case of missing information conditions. Mobile phone signal stability technology and gene identification, which are closely related to life, are all applications of the Bayesian network.
[0098] Formally, the Bayesian network belongs to a directed acyclic graph, and the nodes represent random variables. The directed edges between the nodes represent conditional dependence relationships, pointing from the parent node to the child node. Each node is associated with a probability function. The input of the probability function is a set of specific values of the random variables represented by the parent nodes of this node, and the output is the probability value of the random variable represented by the current node.
[0099] ;
[0100] The graph structure of the network is calculated by combining the A* algorithm with the d-separation algorithm, and then the conditional probability distribution on each node is obtained by the counting method.
[0101] Bayesian network is also an uncertainty causal association model, which can learn and reason under the condition of known limited, incomplete and uncertain information. Therefore, it is widely used in fields such as fault diagnosis and maintenance decision-making. The main usage method is information completion. For example, when any one node or several nodes are known, the probability of another certain node occurring. To adapt to the analysis of abnormal main causes, the influence degree of abnormal X on abnormal Y under the Bayesian network is defined as the probability of Y when X occurs divided by the probability of Y when X does not occur. The formula is as follows:
[0102] ;
[0103] The network can learn different combinations of abnormal X on abnormal Y under different abnormal combinations group(a). And the occurrence of other abnormalities others(o) outside the abnormal group group(a) is unknown.
[0104] 4. Build an abnormal quantitative impact analysis model.
[0105] Abnormal quantitative impact analysis includes an abnormal user impact quantification model and an abnormal substation area impact quantification model to locate the main abnormal users and main abnormal substation areas for treatment.
[0106] 41. Abnormal user impact quantification model. With the help of the main cause impact weight coefficient of line loss unqualified main cause analysis, the classification of unqualified characteristics of the substation area, and the power loss reduction space analysis of abnormal users, etc., the power loss reduction space of each abnormal user is calculated by multiplying two variables. The power loss reduction space of users is analyzed for the substation areas with abnormal line loss and abnormal users on a daily basis with the user as the granularity. Specifically, calculate the metering abnormal loss power as the power loss reduction space of the metering abnormal electric meter. Use the time series algorithm to calculate the power of the collected abnormal electric meter as the power loss reduction space of the collection abnormality. The specific analysis process is as follows:
[0107] 411. Analyze the abnormal values of the historical data of each user in the past two months. Replace the data exceeding 1.5 times the normal range with the upper and lower bounds of the normal range.
[0108] 412. Use the exponential smoothing algorithm to estimate the electricity consumption of users when the abnormality occurs as the power loss reduction space of users by using the user historical data.
[0109] 413. For users that cannot converge, use the average value of the same period electricity consumption in each of the past 8 weeks as the power loss reduction space of users.
[0110] 42. Abnormal Substation Area Impact Quantification Model. By means of user / substation area abnormal impact quantification analysis and models such as one index per substation area, the electricity consumption that can be reduced due to user / substation area abnormalities and the target value of substation area line loss are output. The impact of substations with unqualified line loss is quantified through these two values, and finally, the main abnormal treatment substations are located according to the size of the reducible loss space. The specific analysis process is as follows:
[0111] 421. According to the output results of abnormal user impact quantification, obtain all the electricity consumption data A of abnormal users in substations with unqualified line loss.
[0112] 422. Refer to the abnormal user impact quantification rules to calculate the impact quantification of terminals and substations, and obtain all the electricity consumption data B of abnormal terminals and the electricity consumption data C of the substation area in substations with unqualified line loss.
[0113] 423. Obtain all the abnormal impact electricity consumption in substations with unqualified line loss: A + B + C.
[0114] 424. Obtain the reducible loss space of substations with unqualified line loss with one index per substation area.
[0115] By constructing models such as substation area unqualified characteristic classification, main cause analysis of unqualified line loss, main cause analysis of abnormal association, and abnormal quantification impact analysis, applying machine learning models such as SHAP algorithm analysis model and Bayesian network model, and combining with the experience of business experts, a set of substation area line loss abnormal quantification evaluation and optimal treatment plan is formed. The present invention can efficiently focus on and solve key abnormal problems on the premise of limited resources to achieve the goals of line loss reduction and efficiency improvement.
[0116] The present invention also provides a line loss intelligent diagnosis device based on comprehensive analysis by machine learning methods, adopting the above-mentioned line loss intelligent diagnosis method based on comprehensive analysis by machine learning methods, including: a substation area unqualified characteristic classification module, which is used to classify abnormal substations by dividing the daily line loss rate type according to the daily line loss rate result of the substation area and the theoretical value of one index per substation area, and calculating the volatility of the substation area according to the daily line loss result of the substation area; a main cause analysis module of unqualified line loss, which is used to obtain the abnormal causes of unqualified line loss in the substation area through an ensemble learning model and a SHAP algorithm quantification model; an abnormal association main cause analysis module, which is used to locate the main cause of abnormal association by analyzing the relationship between abnormal causes through a Bayesian network model; and an abnormal quantification impact analysis module, which is used to locate the abnormal users and abnormal substations to be treated through an abnormal user impact quantification model and an abnormal substation area impact quantification model.
[0117] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0118] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
[0119] The above specific implementation manners of the present invention do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A line loss intelligent diagnosis method based on comprehensive analysis of machine learning methods, characterized in that: The steps include: S1. Construct a classification model for substation unqualified characteristics based on line loss volatility and daily line loss rate, including classifying daily line loss rate types according to the daily line loss rate results of the substation and the theoretical value of an indicator of the substation, and calculating the substation volatility according to the daily line loss results of the substation, which is used to classify abnormal substations; S2. Construct a main cause analysis model for line loss failure, including model training, abnormal sorting, and sorting result summary, to obtain the abnormal causes of line loss failure in the substation area; The model training is based on the substation unqualified characteristic classification model and abnormal factor library. For each type of substation, a method combining Pearson correlation coefficient and decision tree regression is used to screen abnormal factors to improve the efficiency and stability of model training. The integrated learning XGBoost algorithm is used to train the model to realize line loss prediction in abnormal substations. The abnormal sorting constructs a quantitative abnormal sorting model based on the SHAP algorithm, calculates the marginal contribution of each characteristic factor to the line loss rate of the substation area, and quantifies the influence weight of each characteristic factor on the line loss rate; The ranking result is summarized based on all calculated abnormal causes of the substations, and the importance of each abnormal cause in each substation class is obtained according to the substation classification, and all important abnormalities under the substation are standardized to obtain abnormal weights; S3. Construct an abnormal correlation main cause analysis model, and analyze the relationship between abnormal causes by constructing a Bayesian network model to locate the main abnormal cause; S4. Construct an abnormal quantitative impact analysis model, including an abnormal user impact quantitative model and an abnormal area impact quantitative model, to locate the main abnormal users and main abnormal areas for governance; The abnormal user impact quantification model analyzes the abnormal values of the user's historical data, and replaces the data that exceeds the normal range by 1.5 times with the upper and lower limits of the normal range; using the user's historical data, an exponential smoothing algorithm is used to estimate the user's power consumption when the abnormality occurs as the user's loss reduction space; for users who cannot converge, the average power consumption in the same period is used as the user's loss reduction space; The abnormal substation impact quantification model obtains the data on the power consumption affected by all abnormal users in the substation with unqualified line loss according to the quantification output results of the abnormal user impact; calculates the impact quantification of terminals and substations with reference to the quantification rules of the abnormal user impact, and obtains the data on the power consumption affected by all abnormal terminals and the power consumption affected by the substation in the substation with unqualified line loss; obtains the power consumption affected by all abnormalities in the substation with unqualified line loss; and obtains the loss reduction space for the substation with unqualified line loss in one indicator.
2. The line loss intelligent diagnosis method based on comprehensive analysis by machine learning method according to claim 1 is characterized in that: In step S1, the daily line loss rate type is divided according to the daily line loss rate result of the area and the theoretical value of an index of an area, including ultra-high loss, high loss, small negative loss, negative loss and uncalculated; Ultra-high loss: daily line loss rate ≥ upper threshold + 20%, high loss: upper threshold ≤ daily line loss rate < upper threshold + 20%, small negative loss: -1% ≤ daily line loss rate < 0%, negative loss: daily line loss rate < -1%; Calculate the volatility of the area based on the daily line loss results of the area, including stability, fluctuation and suddenness; Stable: abnormal line loss days in a week ≥ 3 and abnormal dispersion coefficient in the past 60 days < 3, Fluctuation: 1≤ abnormal line loss days in a week < 3 and abnormal dispersion coefficient in the past 60 days ≥ 3, Sudden: abnormal line loss on the day and non-stable and fluctuating areas; The calculation formula of the abnormal dispersion coefficient is: Among them, θ i It represents the line loss rate of the i-th day in the substation, N represents the data period required to analyze the volatility of the substation, and the default value is 60 days. The higher the abnormal dispersion coefficient σ, the greater the abnormal fluctuation of the daily line loss in the substation.
3. The line loss intelligent diagnosis method based on comprehensive analysis by machine learning method according to claim 2 is characterized in that: In step S2, the model training is specifically as follows: The training set is defined as a data set containing m samples and n features, |D| = {(x1, y1), (x2, y2), ..., (x m ,y m )}, the prediction output of the ensemble learning XGBoost algorithm training model is expressed as Among them, φ(x i ) represents the overall prediction function of the model, which is composed of the sum of the prediction outputs of all K trees, f k represents the regression tree, K is the number of regression trees, x i Represents the feature vector of a sample of input and outputs the predicted value Add the predicted values of K regression trees; XGBoost's objective function adds a regularization term based on the GBDT loss function. That is, the sum of the regularization values of all regression trees, the objective function of the tth tree is: Where γ is the L1 regularization coefficient of the number of leaf nodes, J is the number of leaf nodes in the tree, λ is the L2 regularization coefficient of the leaf node weight, and w tj is the weight of the jth leaf node of the tth tree; y i is the actual line loss rate of the ith sample, f t-1 (x i ) is the sum of the predicted values before the t-th tree, x i is the feature vector of the i-th sample, f t (x i ) is the predicted value of the tth tree.
4. The line loss intelligent diagnosis method based on comprehensive analysis by machine learning method according to claim 1, characterized in that: In step S3, the Bayesian network model is an indeterminate causal association model, and the degree of influence of abnormality X on abnormality Y under the Bayesian network is defined as the probability of Y when X occurs compared to the probability of Y when X does not occur, and the formula is as follows: The Bayesian network model learns different combinations of anomaly X to anomaly Y under different anomaly combinations group(a), and the occurrence of other anomalies others(o) outside the anomaly group(a) is unknown.
5. A line loss intelligent diagnosis device based on comprehensive analysis by machine learning method, adopting the line loss intelligent diagnosis method based on comprehensive analysis by machine learning method as claimed in any one of claims 1 to 4, characterized in that: The device comprises: The unqualified characteristics classification module of the substation area is used to classify the abnormal substation area by dividing the daily line loss rate type according to the daily line loss rate results of the substation area and the theoretical value of an indicator of the substation area, and calculating the substation area volatility according to the daily line loss results of the substation area; The main cause analysis module of line loss failure uses an integrated learning model and SHAP algorithm quantification model to obtain the abnormal causes of line loss failure in the substation area; The abnormal correlation main cause analysis module uses the Bayesian network model to analyze the relationship between abnormal causes and locate the main cause of abnormal correlation; The abnormal quantitative impact analysis module is used to locate abnormal users and abnormal substations for management through the abnormal user impact quantitative model and the abnormal substation impact quantitative model.
Citation Information
Patent Citations
Transformer area line loss abnormity auxiliary diagnosis method
CN111781463A
Line loss main cause analysis method and related equipment thereof
CN116245505A
Method for determining main characteristic factors influencing line loss rate of low-voltage transformer area
CN118132914A