Power transmission line pollution flashover prediction method and system based on meteorological parameters and random forest algorithm

Through the pollution flash prediction method based on meteorological parameters and random forest algorithm, the problems of low precision of pollution flash prediction and insufficient real-time performance in the existing technology are solved, and more accurate and real-time pollution flash risk prediction is achieved, operating risks and operation and maintenance costs are reduced, and the intelligent development of the power grid is promoted.

CN120217174APending Publication Date: 2025-06-27QUANZHOU POWER SUPPLY COMPANY OF STATE GRID FUJIAN ELECTRIC POWER +1
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510229471.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology has low prediction accuracy in pollution flash prediction, cannot fully consider the coupling effect of multi-dimensional factors, is insufficient sensitivity to meteorological conditions and dynamic changes of pollutants, lacks real-time and intelligence, and is difficult to meet the requirements of modern power grids to respond quickly to faults.

Method used

The pollution prediction method of transmission line based on meteorological parameters and random forest algorithm is adopted to collect meteorological data in real time, combine historical pollution data, and establish a pollution prediction model. The meteorological data is extracted and modeled through the random forest algorithm, a classification tree is constructed and the prediction results of multiple decision trees are combined to generate pollution risk probability prediction.

Benefits of technology

It significantly improves the accuracy and real-time nature of the pollution flash prediction, reduces the operating risks of transmission lines caused by pollution flash failure, optimizes the allocation of maintenance resources, reduces operation and maintenance costs, and promotes the intelligent development of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217174A_ABST
    Figure CN120217174A_ABST
Patent Text Reader

Abstract

The invention provides a power transmission line pollution flashover prediction method and system based on meteorological parameters and a random forest algorithm, and the method comprises the steps: collecting meteorological data around a power transmission line in real time, and building a pollution flashover prediction model in combination with historical pollution flashover data; and performing feature extraction and modeling on the acquired data by adopting a random forest algorithm, constructing a classification tree, integrating prediction results of a plurality of decision trees, and generating pollution flashover risk probability prediction through majority voting. The model is deployed on a cloud or a local server, receives meteorological data input in real time, dynamically outputs pollution flashover probability and risk level, and performs visual display through an upper computer and a cloud platform. And when the predicted risk exceeds a set threshold value, the system triggers an early warning mechanism and sends early warning information to operation and maintenance personnel to assist in formulating preventive maintenance measures. According to the method, the accuracy and the real-time performance of pollution flashover prediction are effectively improved, the operation risk caused by the pollution flashover fault of the power transmission line is remarkably reduced, and the method has relatively high application value and popularization prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent operation and maintenance of power systems, and particularly to a method and system for predicting flashover of transmission lines based on meteorological parameters and random forest algorithm. Background Art

[0002] Flashover is a serious environment-related fault in power systems, and its occurrence poses a great threat to the safe and stable operation of transmission lines. The flashover phenomenon usually occurs after the surface of the insulators of transmission lines is contaminated, and a conductive water film is formed under high humidity or rain and fog conditions, resulting in insulator flashover or even breakdown discharge. In recent years, with the rapid development of industrialization and urbanization, the concentration of atmospheric pollutants has increased. Especially in heavy industrial areas, the concentration of suspended particles such as PM2.5 and PM10 has increased significantly, greatly increasing the risk of flashover of transmission lines. At the same time, the change of climate conditions, such as frequent rain and snow weather, atmospheric pressure fluctuations and extreme humidity conditions, also has a complex impact on the flashover risk. At present, traditional flashover monitoring and prediction technologies are mainly based on empirical formulas or statistical analysis. These methods have the following deficiencies. For example, the prediction accuracy is not high. Traditional methods cannot comprehensively consider the coupling effect of multi-dimensional factors and are not sensitive enough to the dynamic changes of meteorological conditions and pollutants. The real-time performance and intelligence are insufficient. Most existing systems are difficult to realize real-time data acquisition, analysis and early warning, and fail to meet the requirements of modern power grids for rapid response to faults. Lack of support from multi-source data. A single environmental data source, such as only focusing on humidity or pollutant concentration, is difficult to comprehensively reflect the complex mechanism leading to flashover. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method and system for predicting flashover of transmission lines based on meteorological parameters and random forest algorithm, which effectively improves the accuracy and real-time performance of flashover prediction, and significantly reduces the operation risk of transmission lines caused by flashover faults.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions: A method for predicting flashover of transmission lines based on meteorological parameters and random forest algorithm, which collects meteorological data around the transmission line in real time, combines historical flashover data, and establishes a flashover prediction model; uses the random forest algorithm to extract features and model the collected meteorological data, constructs classification trees and synthesizes the prediction results of multiple decision trees, and generates a flashover risk probability prediction through majority voting; the model is deployed on a cloud or local server, receives meteorological data input in real time, dynamically outputs the flashover probability and risk level, and is displayed through a host computer and a cloud platform; when the predicted risk exceeds the set threshold, the system triggers an early warning mechanism and sends an early warning message.

[0005] In a preferred embodiment, the meteorological data includes temperature, humidity, PM2.5 and PM10 concentrations, number of precipitation days and precipitation amount, and atmospheric pressure; meteorological data and pollution flash occurrence records over a period of time are obtained through a database, and the features and labels are aligned.

[0006] In a preferred embodiment, the formula for a random forest: Assume there are N decision trees, and the prediction result of each tree is T i (x), and the final classification result y is determined by majority voting:

[0007] y = vote{T1(x), T2(x), … T N (x)}

[0008] where T i (x) is the prediction result of the i-th tree.

[0009] In a preferred embodiment, each decision tree is generated in the following steps:

[0010] Generate a sub-dataset using Bootstrap sampling from the training data; and at each splitting node of the decision tree, randomly select m candidate features, where m < M and M is the total number of features, calculate the best splitting point of the candidate features, and select the splitting feature based on the Gini coefficient:

[0011]

[0012] where p k represents the proportion of samples of the k-th class in the current node. The smaller the Gini coefficient, the lower the impurity of the node; therefore, by calculating the Gini coefficients of each feature and constructing the decision tree according to the principle of splitting the one with the larger Gini coefficient first;

[0013] When an input sample passes through the decision tree, it will reach a certain leaf node all the way from the root node according to the splitting rules; the sample distribution information stored in the leaf node, that is, the number of samples of different classes that fall into this node in the training data; based on this distribution information, the decision tree calculates the probability that the sample belongs to a certain class; then for a single decision tree, given an input sample x, the pollution flash probability P i for it under the i-th decision tree is:

[0014]

[0015] where, U p and U t are the number of pollution flash events and the total number of samples in the leaf node that the sample x finally reaches in the decision path.

[0016] In a preferred embodiment, for a new sample x, each decision tree outputs the probability P i, based on the random forest model, the random forest takes the average of the probability values of all trees as the final output. The formula for finally predicting the flashover probability is:

[0017]

[0018] where P(x) is the predicted flashover probability, and its value range is [0, 1]. P i is the flashover probability predicted by the i-th tree, and N is the total number of decision trees in the forest.

[0019] In a preferred embodiment, the architecture of model deployment can be divided into three main parts: data acquisition layer, processing and prediction layer, and application layer.

[0020] In a preferred embodiment, the data acquisition layer includes real-time meteorological data acquisition; using meteorological stations, environmental monitoring devices or accessing third-party meteorological service interface APIs to collect relevant meteorological parameters in real time; historical meteorological and flashover data storage stores long-term accumulated meteorological data and flashover event records of transmission lines through cloud or local databases for model training and subsequent optimization.

[0021] In a preferred embodiment, the data processing and prediction layer includes a model running environment, a data preprocessing module, and a prediction module; the model running environment deploys a random forest prediction model on a cloud server or a local server. Using the trained model, it converts the real-time input meteorological data into a flashover probability output; the data preprocessing module performs operations such as cleaning, filling missing values, and filtering outliers on the collected meteorological data to ensure data quality and unify the data format to meet the model input requirements; the prediction module predicts the flashover probability in real time through the random forest model; the output results include: flashover probability and flashover risk level.

[0022] In a preferred embodiment, the application layer includes the display of prediction results and alarm and operation and maintenance decision support;

[0023] Display of prediction results: On the upper computer or cloud platform interface, the flashover probability, risk level, and meteorological data trend chart are displayed in real time; Alarm and operation and maintenance decision support: When the flashover probability exceeds the set threshold, the alarm mechanism is triggered to send a warning message to the operation and maintenance personnel.

[0024] The present invention also provides a transmission line flashover prediction system based on meteorological parameters and the random forest algorithm, which adopts the above-mentioned transmission line flashover prediction method based on meteorological parameters and the random forest algorithm, and includes a meteorological module, a cloud / local database, an upper computer, and a model training module.

[0025] Compared with the prior art, the present invention has the following beneficial effects: 1. Improve the operation safety of transmission lines. By accurately predicting the possibility of flashover, preventive measures such as cleaning insulators and adjusting the operation mode of the power grid can be taken in advance, reducing the risk of line faults and ensuring the safe and stable operation of the transmission system. 2. Reduce maintenance costs. Traditional operation and maintenance of transmission lines require periodic manual inspections and insulator cleaning, which is inefficient and costly. Through accurate flashover prediction, the allocation of maintenance resources can be optimized, concentrating cleaning and maintenance work in high-risk areas, thereby reducing the overall operation and maintenance costs. 3. Promote the intelligent development of the power grid. The present invention combines meteorological data collection, cloud big data analysis and artificial intelligence algorithms, conforming to the current trend of the intelligent and digital development of the power grid, providing new technical support for the construction of a smart grid. 4. Cope with complex climate and environmental challenges. With the intensification of global climate change, the frequent occurrence of extreme weather (such as heavy precipitation and high humidity) further exacerbates the flashover risk. By introducing meteorological parameters, this method can adapt to complex environmental changes and improve the applicability and robustness of prediction. 5. The innovation of combining theory with practice. Traditional flashover research focuses more on static characteristic analysis and laboratory simulation, while the present invention is based on a large amount of historical data in the real environment, combined with the feature selection and modeling capabilities of machine learning, providing a new path from theory to practice. Through the non-linear modeling ability of the random forest algorithm, the complex relationship between meteorological parameters, pollutant concentration and flashover risk can be comprehensively captured. In summary, the present invention has important theoretical value and practical significance. It can not only significantly reduce the flashover risk of transmission lines, but also provide a new technical solution for the efficient operation and reliable operation of smart grids, with broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a block diagram of the method of the preferred embodiment of the present invention;

[0027] Figure 2 It is a schematic diagram of the random forest of the preferred embodiment of the present invention;

[0028] Figure 3 It is the random forest algorithm flow of the preferred embodiment of the present invention;

[0029] Figure 4 It is the overall flowchart of the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The present invention will be further described below in conjunction with the drawings and embodiments.

[0031] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0032] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0033] A transmission line pollution flashover prediction system based on meteorological parameters and random forest algorithm, refer to Figures 1-4 , including:

[0034] (1) Meteorological module

[0035] The meteorological module is one of the core components of the present invention and is used to collect real-time meteorological environment parameters around the transmission line. This module monitors key meteorological factors such as temperature, humidity, PM2.5, PM10, atmospheric pressure, precipitation, etc. through a variety of sensor devices, and transmits the collected data to the cloud / local database to support the analysis and training of the pollution flashover prediction model. The meteorological module includes the following hardware devices and functional sub-modules:

[0036] Temperature and humidity sensor: Used to collect real-time ambient temperature and humidity data around the transmission line. These data are important meteorological parameters affecting the occurrence of pollution flashover and can reflect the conditions for the formation of a water film on the surface of the line insulator.

[0037] Particulate matter monitor (PM2.5, PM10): Used to monitor the concentration of fine particulate matter and coarse particulate matter in the air. Particulate matter is an important source of surface contamination of insulators, and its concentration directly affects the pollution flashover risk.

[0038] Atmospheric pressure sensor: Used to detect the change in atmospheric pressure in the area where the transmission line is located. Atmospheric pressure has a certain influence on the discharge electric field strength of the insulator.

[0039] Precipitation and precipitation days monitoring equipment: Records the precipitation and precipitation frequency in the area through devices such as rain gauges. These data can reflect the possibility of forming a conductive water film on the surface of the insulator.

[0040] Data acquisition terminal: Integrates the data of each sensor and transmits the data to the cloud / local database through 4G communication.

[0041] (2) Cloud / local database

[0042] The cloud / local database module is the core data storage and management system of the present invention, which is used to centrally store, manage, and process historical flashover data and real-time collected meteorological data. This module realizes the dynamic update, cross-platform sharing, and efficient invocation of data, providing a solid data foundation for the training, update, and real-time prediction of the flashover prediction model.

[0043] The cloud database is used for the storage and cross-regional access of large-scale historical data, while the local database is responsible for the rapid storage and processing of real-time data. The two achieve efficient synchronization through 4G communication technology. The database supports standardized interfaces, providing instant invocation services for the training and prediction modules of the random forest prediction model. In addition, the database also has functions such as data encryption, permission management, and automatic backup to ensure the security and reliability of data transmission and storage.

[0044] By dynamically integrating historical and real-time data, the cloud / local database provides high-quality input data for the flashover prediction model, significantly improving the accuracy and reliability of the prediction system, and at the same time providing strong data support for the operation and maintenance decision-making of transmission lines.

[0045] (III) Host Computer and Model Training

[0046] The host computer is the core control and interaction platform in the present invention, responsible for the overall management of the system, data processing, model training, and display of prediction results. As a key node connecting the cloud / local database, the meteorological module, and the flashover prediction module, the host computer realizes the unified scheduling and real-time monitoring of each functional module of the system through a combination of software and hardware.

[0047] The host computer calls the random forest algorithm for offline training of the flashover prediction model. During the model training process, the host computer extracts key features and generates a model by combining historical flashover data and meteorological parameters. As new data is continuously entered, the host computer can regularly update the model to adapt to the latest meteorological and operating environments. Based on real-time meteorological data, the host computer loads the trained random forest model for online prediction. The prediction results are analyzed by the warning module of the host computer. If a high flashover risk is found, the system will send out corresponding warning signals and send them to the terminals of operation and maintenance personnel through the network.

[0048] A transmission line flashover prediction model is constructed using the random forest model. Through in-depth correlation analysis of historical flashover data and meteorological parameters, the model can identify the relationship between meteorological conditions and flashover events, and achieve accurate prediction of flashover risks.

[0049] 1. Basic Steps of Model Construction (1) Data Preparation

[0050] Input Features (Meteorological Parameters):

[0051] Temperature, humidity, PM2.5 and PM10 concentrations, number of precipitation days and precipitation amount, atmospheric pressure.

[0052] Target variable (flashover label):

[0053] Classification label for flashover events, defined as: = 1: Flashover occurred, = 0: No flashover occurred.

[0054] Source of historical data:

[0055] Historical meteorological parameters and corresponding flashover records are the core data sources for training the model. Obtain meteorological data and flashover occurrence records within a certain period through the database, and align features and labels.

[0056] Division of training set and test set: Divide according to 70% for the training set and 30% for the test set

[0057] (2) Data preprocessing

[0058] Filling of missing values: For missing values in meteorological data, use mean filling or interpolation method to fill.

[0059] Removal of outliers: Remove extreme values (such as negative values of human input errors) for features such as PM2.5, temperature and humidity.

[0060] Normalization processing: Normalize all input meteorological parameters to the same range ([0,1]):

[0061]

[0062] where, x min and x max are the minimum and maximum values of this feature respectively.

[0063] (3) Construction of random forest model

[0064] Principle of random forest algorithm:

[0065] Random forest is an ensemble learning algorithm composed of multiple decision trees. Its main ideas include: randomly extracting several sub-samples from the data set (Bootstrap sampling), each decision tree randomly selects some features for splitting to generate a classification tree, and integrating the prediction results of all trees, and determining the final classification through majority voting.

[0066] Formula of random forest: Assume there are N decision trees, and the prediction result of each tree is T i (x), the final classification result y is determined through majority voting:

[0067] y = vote{T1(x), T2(x), … T N (x)}

[0068] where T i (x) is the prediction result of the i-th tree

[0069] Hyperparameter setting

[0070] Number of trees: For example, 100 trees, which will be dynamically adjusted during the subsequent training of the pollution flashover prediction model.

[0071] Maximum depth: Limit the depth of the tree to prevent overfitting.

[0072] Number of splitting features: The number of features randomly selected each time for splitting. When used for pollution flashover prediction, we randomly select 5 features.

[0073] 2. Model input parameters

[0074] In the model of the present invention, the input parameters are the following 7 key meteorological factors, as shown in Table 1, which are used to comprehensively describe the meteorological conditions for the occurrence of pollution flashover:

[0075] Table 1:

[0076]

[0077] 3. Decision tree construction

[0078] Each decision tree is generated in the following steps:

[0079] Generate a sub-dataset (containing about 63% of the original samples) from the training data using Bootstrap sampling. And at each splitting node of the tree, randomly select m candidate features (m < M is the total number of features), calculate the best splitting point of the candidate features, and select the splitting feature based on the Gini Impurity:

[0080]

[0081] where p k represents the proportion of samples of the k-th class in the current node. The smaller the Gini coefficient, the lower the impurity of the node. Therefore, by calculating the Gini coefficients of each feature and following the principle of splitting the one with the larger Gini coefficient first, the decision tree is constructed.

[0082] When an input sample passes through the decision tree, it will reach a certain leaf node all the way from the root node according to the splitting rules. The leaf node stores the sample distribution information (i.e., the number of samples of different classes) that falls into this node in the training data. Based on this information, the decision tree can calculate the probability that the sample belongs to a certain class (such as the occurrence of pollution flashover y = 1). Then for a single decision tree, given an input sample x, its pollution flashover probability P i is:

[0083]

[0084] Among them, U p and U t are the number of flashover events and the total number of samples in the leaf node finally reached by the decision path of the sample x.

[0085] 4. Flashover model output

[0086] For the new sample x, each tree outputs the probability P i that it belongs to the flashover event (flashover occurs). The random forest takes the average of the probability values of all trees as the final output. Based on the random forest model, the formula for finally predicting the flashover probability is:

[0087]

[0088] where P(x) is the predicted flashover probability, and its value range is [0, 1]. P i is the flashover probability predicted by the i-th tree, and N is the total number of decision trees in the forest.

[0089] 5. Model performance optimization and parameter tuning

[0090] Optimize the performance of the random forest and improve the flashover probability prediction effect through the following means:

[0091] Number of trees N: The larger N is, the stronger the model stability, but the training time increases. Usually, N = 100 - 500 is taken.

[0092] Maximum tree depth d: Control the complexity of each tree to avoid overfitting. Usually, d ∈ [10, 30] is set.

[0093] Number of splitting features m: Generally set to or log2(M), to balance accuracy and speed.

[0094] Minimum number of samples for splitting: The minimum number of samples at each node can control the splitting depth and prevent overfitting.

[0095] Subsequent optimization and verification:

[0096] Model integration

[0097] Model integration can be combined with other models (such as logistic regression, gradient boosting tree) to improve the overall prediction performance.

[0098] Model verification

[0099] Conduct independent verification on the test set to ensure that the model is not overfitted to the training set and the validation set. Use the real meteorological data stream input to simulate the actual operation scenario and evaluate the dynamic prediction performance of the model.

[0100] 6. Model deployment

[0101] The architecture of model deployment can be divided into three main parts: data collection layer, processing and prediction layer, and application layer.

[0102] (1) Data collection layer

[0103] Real-time meteorological data collection:

[0104] Use meteorological stations, environmental monitoring equipment or access to third-party meteorological service interfaces (APIs) to collect relevant meteorological parameters in real time, including: temperature, humidity, PM2.5, PM10, number of precipitation days, precipitation amount, atmospheric pressure, etc. Ensure the stability of the collection equipment, and the data collection frequency can be set to once per hour or once per minute.

[0105] Historical meteorological and pollution flashover data storage:

[0106] Store the long-term accumulated meteorological data and pollution flashover event records of transmission lines through cloud or local databases for model training and subsequent optimization.

[0107] (2) Data processing and prediction layer

[0108] Model operating environment:

[0109] The random forest prediction model deployed on the cloud server or local server uses the trained model to convert the real-time input meteorological data into the output of pollution flashover probability.

[0110] Data preprocessing module:

[0111] Perform operations such as cleaning, filling missing values, and filtering outliers on the collected meteorological data to ensure data quality, unify the data format, and make it meet the model input requirements.

[0112] Prediction module:

[0113] Real-time predict the pollution flashover probability through the random forest model. The output results include: pollution flashover probability (a value between 0 and 1).

[0114] Pollution flashover risk level (low, medium, high), divided according to the set threshold.

[0115] (3) Application layer

[0116] Display of prediction results:

[0117] On the upper computer or cloud platform interface, display the pollution flashover probability, risk level, and meteorological data trend chart in real time. Provide simple and intuitive visualization charts, such as: time series curve (change of pollution flashover probability over time), map distribution (pollution flashover risk distribution in different transmission line areas).

[0118] Alarm and operation and maintenance decision support:

[0119] When the pollution flashover probability exceeds the set threshold, the alarm mechanism is triggered to send warning messages such as text messages, emails or platform pop-ups to the operation and maintenance personnel.

[0120] The present invention has successfully realized the real-time monitoring and accurate prediction of pollution flashover events. By integrating a variety of meteorological data and historical pollution flashover records, and using the non-linear modeling ability of the random forest algorithm, this method can comprehensively analyze the complex relationship between meteorological parameters and pollution flashover risks, thereby generating quantitative pollution flashover probabilities and risk levels. The application of the present invention can effectively reduce the transmission line failures caused by pollution flashovers, improve the stability and reliability of the transmission system, and has broad application prospects and promotion value. At the same time, this method also has certain reference significance in the field of predicting other environment-related power grid failures, providing technical support for the development of smart grids. In the future, this method can be combined with artificial intelligence technology to further improve the model performance and be extended to a wider range of power grid intelligent operation and maintenance scenarios, contributing to the intelligent development of the power industry.

Claims

1. A method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm, characterized in that: Real-time collection of meteorological data around the transmission line, combined with historical pollution flashover data, to establish a pollution flashover prediction model; using the random forest algorithm, feature extraction and modeling of the collected meteorological data, building a classification tree and integrating the prediction results of multiple decision trees, generating a pollution flashover risk probability prediction through majority voting; The model is deployed on the cloud or local server, receives meteorological data input in real time, dynamically outputs pollution flashover probability and risk level, and displays them through the host computer and cloud platform; when the predicted risk exceeds the set threshold, the system triggers the early warning mechanism and sends a warning message.

2. A method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 1, characterized in that: The meteorological data include temperature, humidity, PM2.5 and PM10 concentrations, precipitation days and precipitation, and atmospheric pressure. The meteorological data and pollution flashover occurrence records over a period of time are obtained through a database, and the features and labels are aligned.

3. The method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 2 is characterized in that: The formula of random forest: Assume there are N decision trees, and the prediction result of each tree is T i (x), the final classification result y is determined by majority voting: y=vote{T1(x),T2(x),…T N (x)} Where T i (x) is the prediction result of the i-th tree.

4. The method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 1 is characterized in that: Each decision tree is generated in the following steps: Use Bootstrap sampling to generate a sub-dataset from the training data; and randomly select m candidate features at each split node of the decision tree, where m < M, M is the total number of features, calculate the best split point for the candidate features, and select the split feature based on the Gini coefficient: where p k It indicates the proportion of samples of the kth class in the current node. The smaller the Gini coefficient, the lower the impurity of the node. Therefore, the decision tree is constructed by calculating the Gini coefficient of each feature and following the principle of preferential splitting based on the large Gini coefficient. When an input sample passes through a decision tree, it will follow the splitting rule from the root node to a leaf node; the leaf node stores the sample distribution information that falls into the node in the training data, that is, the number of samples of different categories; based on this distribution information, the decision tree calculates the probability that the sample belongs to a certain category; then for a single decision tree, given an input sample x, its probability of contamination flashover P under the i-th decision tree i for: Among them, U p and U t is the number of pollution flashover events and the total number of samples in the leaf node that sample x finally reaches in the decision path.

5. A method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 4, characterized in that: For a new sample x, each decision tree outputs the probability P that it belongs to a pollution flash event. i , random forest takes the average probability value of all trees as the final output. Based on the random forest model, the final formula for predicting the probability of pollution flashover is: Where P(x) is the predicted probability of pollution flashover, ranging from [0,1], P i is the pollution flashover probability predicted by the i-th tree, and N is the total number of decision trees in the forest.

6. The method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 1 is characterized in that: The architecture of model deployment can be divided into three main parts: data collection layer, processing and prediction layer, and application layer.

7. A method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 6, characterized in that: The data collection layer includes real-time meteorological data collection; using meteorological stations, environmental monitoring equipment or accessing third-party meteorological service interface APIs to collect relevant meteorological parameters in real time; historical meteorological and pollution flashover data storage uses the cloud or local database to store long-term accumulated meteorological data and pollution flashover event records of transmission lines for model training and subsequent optimization.

8. The method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 6 is characterized in that: The data processing and prediction layer includes the model operation environment, data preprocessing module and prediction module. The model operation environment is deployed on the random forest prediction model of the cloud server or local server. The trained model is used to convert the real-time input meteorological data into the pollution flashover probability output. The data preprocessing module cleans the collected meteorological data, fills in missing values, filters out outliers, and other operations to ensure data quality and unify the data format to meet the model input requirements; the prediction module predicts the probability of pollution flashover in real time through the random forest model; the output results include: pollution flashover probability and pollution flashover risk level.

9. The method for predicting power line pollution flashover based on meteorological parameters and random forest algorithm according to claim 6, characterized in that: The application layer includes the display of prediction results as well as alarm and operation and maintenance decision support; Display of prediction results: On the host computer or cloud platform interface, the pollution flashover probability, risk level and meteorological data trend chart are displayed in real time; Alarm and operation and maintenance decision support: When the pollution flashover probability exceeds the set threshold, the alarm mechanism is triggered to send early warning information to the operation and maintenance personnel.

10. A power transmission line pollution flashover prediction system based on meteorological parameters and random forest algorithm, characterized in that A power transmission line pollution flashover prediction method based on meteorological parameters and random forest algorithm as described in any one of claims 1 to 9 is adopted, including a meteorological module, a cloud / local database, a host computer and a model training module.

Citation Information

Cited By

  • Multi-dimensional sensing panoramic monitoring method and system for power transmission line

    CN120707972A

  • A multi-dimensional sensing panoramic monitoring method and system for power transmission lines

    CN120707972B

  • Power equipment operation and maintenance data acquisition method and system

    CN121415560A

  • Ground wire state monitoring and early warning system and method based on random forest algorithm

    CN121461601A

  • Grounding wire state monitoring and early warning system and method based on random forest algorithm

    CN121461601B