Insulator Pollution Risk Classification and Maintenance Decision-Making Methods and Systems

CN122573128APending Publication Date: 2026-08-14SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]本发明提供了一种绝缘子污秽风险分级与维护决策方法,解决了传统方法难以将连续风险预测结果转化为可执行决策,且缺乏对湿润工况、运行态势、资源约束等多因素的系统性优化的技术问题

Benefits of technology

[0051]通过整合上游连续污秽风险指数与微气象环境数据,利用含湿润激发耦合项的概率映射函数将抽象指数转化为具有明确物理意义的闪络失效概率,解决了传统风险度量缺乏直接行动含义的技术瓶颈。其中湿润耦合机制能够精准捕捉高积污与高湿润的协同效应,提升了极端工况下的风险识别准确性。通过离散化状态等级和状态回退机制的建立,为后续动态分级与多目标优化提供了统一量化基础,改善了传统运维中因度量不统一导致的漏报与误报问题,为绝缘子污秽风险的精准化、智能化管理奠定了技术基础。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122573128A_ABST
    Figure CN122573128A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for insulator pollution risk classification and maintenance decision-making, belonging to the field of condition-based maintenance of power equipment. The method acquires pollution risk index and micrometeorological data, calculates flashover failure probability using a probability mapping function containing a moisture excitation coupling term; constructs a decision function based on operational status, dynamically optimizes thresholds using Bayesian criteria to achieve risk classification; constructs a multi-objective optimization model, and uses an improved deep reinforcement learning algorithm to solve for Pareto optimal maintenance scheduling; finally, it executes the target scheduling according to risk preferences, forming a closed-loop update mechanism. This invention transforms the abstract continuous pollution risk index into a real-time flashover probability with clear physical meaning, solving the problem of traditional risk measurement lacking direct action guidance; its moisture coupling mechanism effectively captures the synergistic effect of pollution and moisture, improves the accuracy of extreme condition identification, and improves false alarms and missed alarms through quantitative state basis, providing technical support for insulator operation and maintenance management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of condition-based maintenance of power equipment, specifically to a method and system for classifying pollution risks of insulators and making maintenance decisions. Background Technology

[0002] Outdoor insulators are exposed to dust, salt spray, coal ash, and industrial emissions for extended periods, causing soluble salts and dust particles to gradually accumulate on their surfaces. In dry conditions, the resistivity of the contaminant layer is relatively high, limiting the decrease in insulation strength. However, in humid weather conditions such as fog, dew, or drizzle, the contaminant layer absorbs moisture and deliquesces, forming an electrolyte film. This significantly increases surface conductivity, leading to increased leakage current and potentially inducing arcing in dry areas, which may ultimately develop into a pollution flashover accident. Pollution flashover accidents are sudden and multi-point, potentially causing severe consequences such as restricted power transmission from generating units, increased risks associated with switching power supply to plant auxiliary power, and widespread power outages.

[0003] To prevent flashover, existing operation and maintenance models mainly include the following two categories: The first is periodic maintenance, which involves cleaning or applying anti-flashover coating to insulators at fixed intervals. This model ignores the differences in pollution accumulation under different locations, wind directions, and emission conditions, easily leading to over-maintenance of low-risk equipment and under-maintenance of high-risk equipment. The second is condition-based maintenance based on static thresholds, triggered by leakage current or a single alarm threshold. However, signals such as leakage current are highly sensitive to humidity disturbances, and static thresholds cannot distinguish between "increased risk due to high pollution" and "signal fluctuations due to instantaneous high humidity," easily leading to high false alarm rates, alarm fatigue, and missed alarms.

[0004] Existing technologies propose a method for assessing and predicting insulator pollution risk, which can output a continuous pollution risk index to characterize risk trends. This type of prediction technology solves the problem of difficulty in accurately measuring risk, but a gap remains where the prediction results cannot be directly translated into actionable maintenance work orders. This is mainly manifested in the following ways: the continuous index lacks direct actionable meaning; the risk classification threshold cannot adapt to operational conditions; the maintenance strategy lacks a systematic trade-off between cost and risk; and there is a lack of fully public algorithms that incorporate engineering constraints such as resources and weather conditions into the solution.

[0005] Therefore, there is a need for an operation and maintenance decision-making method that can take into account the upstream dynamic diagram prediction results and form a closed loop in terms of wetting mechanism correction, dynamic classification and multi-objective scheduling. Summary of the Invention

[0006] This invention provides a method for insulator pollution risk classification and maintenance decision-making, which solves the technical problems of traditional methods that make it difficult to transform continuous risk prediction results into actionable decisions and lack systematic optimization of multiple factors such as wet conditions, operating status, and resource constraints.

[0007] This invention is achieved through the following technical solution:

[0008] In a first aspect, this application provides a method for insulator pollution risk classification and maintenance decision-making, the method comprising the following steps:

[0009] Obtain the continuous pollution risk index output by the upstream risk prediction model, collect micrometeorological environmental data, and calculate the real-time flashover failure probability.

[0010] The system acquires power grid operation status data, constructs a decision cost function, optimizes the dynamic classification threshold based on the Bayesian criterion, compares the real-time flashover failure probability with the dynamic classification threshold to classify the risk, and outputs the risk classification result.

[0011] Define a set of operational actions, including waiting, based on the status rollback mechanism and risk classification results, and establish a set of constraints.

[0012] Based on risk status, dynamic classification thresholds, and constraint sets, a multi-objective optimization model is established with the goal of minimizing total maintenance cost and system cumulative expected risk. An improved actor-critic deep reinforcement learning algorithm is used for rolling optimization solution, and Pareto optimal maintenance schedule set is output under the dynamic weight tendency of each objective.

[0013] Based on the actual risk preference, the target schedule is selected from the Pareto optimal operation and maintenance schedule set, and the status rollback results and rewards after the execution are fed back to the policy network for closed-loop iterative updates.

[0014] A further optimization scheme is to use a probability mapping function containing a wetting excitation coupling term to calculate the real-time flashover failure probability, as shown in the following equation:

[0015] ;

[0016] In the formula, This is the linear prediction value of the logistic regression model. ; Let be the flashover failure probability of target insulator device i at time t. As a continuous pollution risk index, For regression coefficients, This is to characterize the wetting-induced coupling term that represents the simultaneous existence of high contamination and high humidity.

[0017] A further optimization is that the effective wetting factor... The construction logic satisfies:

[0018] ;

[0019] In the formula: Normalized effective wetting factor; The threshold for determining condensation; To effectively flush the threshold; Relative humidity; For dew point difference; Rainfall intensity; The wetting intensity coefficient under condensation / fog conditions; The wetting intensity coefficient under drizzle conditions; This serves as the baseline coefficient under drying conditions. Humidity sensitivity coefficient; This represents the drying sensitivity coefficient.

[0020] A further optimization scheme involves obtaining the dynamic hierarchical threshold based on the Bayesian decision criterion, specifically including the following steps:

[0021] The threshold decision expected cost function is constructed as follows:

[0022] ;

[0023] In the formula, Threshold The expected cost at time t, The cost of underreporting as operational status changes, To pay the price for false alarms This represents the probability of missed reports. This represents the false alarm probability.

[0024] Solving in the threshold candidate set makes smallest As a dynamic grading threshold.

[0025] A further optimization scheme is that the constraint set specifically includes:

[0026] Resource capacity constraints are used to ensure that the total amount of resources required for operation and maintenance actions performed per unit of time does not exceed the upper limit of available resources;

[0027] Forward-looking meteorological correlation constraints are used to dynamically determine whether to initiate long-cycle operation and maintenance actions based on the operation preparation cycle and future weather forecasts. If the forecast shows that there is a risk of interruption during the operation window, the option to wait or take short-cycle actions will be forced.

[0028] Meteorological operation window constraints are used to prohibit outdoor maintenance operations during periods of strong winds, thunderstorms, or when safety procedures are not met, based on meteorological forecast data.

[0029] The mutual exclusion of target insulator equipment and the minimum safe working interval constraint are used to prevent the same target insulator equipment from performing multiple actions in parallel at the same time period, and to ensure that the minimum time interval is met between two actions of the same target insulator equipment.

[0030] A further optimization scheme is that the multi-objective optimization model takes the total operation and maintenance cost as the objective. and system cumulative expected risk target The objective function is as follows:

[0031] ;

[0032] ;

[0033] In the formula, For the comprehensive cost function, For conditional flashover failure probability, For the first Weighting of consequences and losses for each target insulator device The target insulator equipment's historical action sequence; For time t, the first The operation and maintenance actions taken for each target insulator device, where T is the preset rolling window length and N is the number of target insulator devices.

[0034] A further optimization scheme is proposed, in which the improved actor-critic deep reinforcement learning algorithm includes the following mechanisms:

[0035] The state space construction mechanism uses stacked state vectors to integrate the quantized current risk state, weather forecast features of future time windows, and manually set preference weight vectors.

[0036] The reward function construction and numerical scaling mechanism define a comprehensive reward function that includes state fallback rewards, cost consumption, and constraint penalty terms, and introduces a numerical scaling factor to solve the gradient vanishing problem; the constraint handling mechanism uses penalty terms to handle weather and resource constraints, generating strong negative penalty signals for violations; and,

[0037] The action space processing mechanism outputs the action probability distribution through the actor network and uses the critic network to evaluate the state value.

[0038] A further optimized solution is that the closed-loop iterative update includes:

[0039] After the target insulator equipment has undergone cleaning, sweeping or coating operations, the contamination status is reset or reduced in the risk assessment at the next moment.

[0040] When the target insulator equipment performs actions that do not require processing, the risk accumulates naturally according to the upstream predicted trend.

[0041] The status after maintenance is reported back to the policy network for closed-loop iterative updates.

[0042] Secondly, this application provides an insulator pollution risk classification and maintenance decision-making system, characterized in that it includes:

[0043] The data acquisition module is used to acquire the continuous pollution risk index output by the upstream risk prediction model and collect micrometeorological environmental data.

[0044] The probability calculation module communicates with the data acquisition module and is used to calculate the real-time flashover failure probability using a probability mapping function containing a wet excitation coupling term, and quantify and map it into a discrete risk state level. At the same time, it constructs a state rollback mechanism triggered by operation and maintenance actions.

[0045] The dynamic grading module communicates with the probability calculation module and is used to construct a decision cost function based on the real-time flashover failure probability output by the probability calculation module. It then optimizes the dynamic grading threshold according to the Bayesian decision criterion and performs risk grading to obtain the risk grading result.

[0046] The strategy construction module communicates with the probability calculation module and is used to define a set of operation and maintenance actions, including waiting, based on the state rollback mechanism and risk classification results, and to establish a set of constraints.

[0047] The optimization solution module communicates with the probability calculation module, dynamic grading module, and strategy construction module. It is used to establish a multi-objective optimization model based on risk status, dynamic grading threshold, and constraint set. It adopts an improved deep reinforcement learning algorithm for rolling optimization and solves the problem. Under the dynamic weight tendency of each objective, it outputs a Pareto optimal operation and maintenance schedule set.

[0048] The decision execution module communicates with the optimization solution module and is used to select the execution target schedule based on risk preference, and feed back the state rollback results and rewards to the policy network to achieve closed-loop iterative updates.

[0049] Thirdly, this application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores an insulator pollution risk classification and maintenance decision program, wherein when the insulator pollution risk classification and maintenance decision program is executed by a processor, it implements the steps of the insulator pollution risk classification and maintenance decision method as described above.

[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0051] By integrating upstream continuous pollution risk index and micrometeorological environment data, and utilizing a probability mapping function containing a wetting excitation coupling term, the abstract index is transformed into a flashover failure probability with clear physical meaning, overcoming the technical bottleneck of traditional risk measurement lacking direct actionable implications. The wetting coupling mechanism accurately captures the synergistic effect of high pollution accumulation and high humidity, improving the accuracy of risk identification under extreme operating conditions. The establishment of discretized state levels and a state rollback mechanism provides a unified quantitative basis for subsequent dynamic classification and multi-objective optimization, improving the problems of missed and false alarms caused by inconsistent measurements in traditional operation and maintenance, and laying the technical foundation for precise and intelligent management of insulator pollution risks. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0053] In the attached diagram:

[0054] Figure 1 A flowchart of the insulator pollution risk classification and maintenance decision-making method provided in the embodiments of this application;

[0055] Figure 2 A schematic diagram illustrating the principle of how the dynamic grading threshold changes with the operational status and real-time risk index in the embodiments of this application;

[0056] Figure 3 A schematic diagram of the risk and cost Pareto optimal operation and maintenance schedule set provided by the multi-objective solution output in the embodiments of this application;

[0057] Figure 4 This application provides a schematic diagram of a maintenance schedule that takes into account weather windows and resource constraints. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0059] First, some of the technical terms used in this application will be explained to help those skilled in the art understand this application.

[0060] Actor Network;

[0061] Critic Network;

[0062] Reward Feedback.

[0063] Firstly, such as Figure 1 As shown, this application provides a method for insulator pollution risk classification and maintenance decision-making, which specifically includes the following steps:

[0064] Step S1: Obtain the continuous pollution risk index output by the upstream risk prediction model, collect micro-meteorological environmental data and calculate the real-time flashover failure probability; wherein, the continuous pollution risk index refers to the quantitative value of insulator pollution risk that changes continuously over time and is output by the upstream risk prediction model, which is used to dynamically assess the level of pollution accumulation and the resulting flashover probability.

[0065] Step S2: Obtain the power grid's operational status data, construct a decision cost function, optimize the dynamic classification threshold based on the Bayesian criterion, compare the real-time flashover failure probability with the dynamic classification threshold to perform risk classification, and output the risk classification result.

[0066] Step S3: Define a set of operation and maintenance actions, including waiting, based on the status rollback mechanism and risk classification results, and establish a set of constraints;

[0067] Step S4: Based on the risk status, dynamic classification threshold and constraint set, establish a multi-objective optimization model with the goal of minimizing the total operation and maintenance cost and the cumulative expected risk of the system. Use the improved actor-critic deep reinforcement learning algorithm to perform rolling optimization and solve the problem. Output the Pareto optimal operation and maintenance schedule set under the dynamic weight tendency of each objective.

[0068] Step S5: Select the target schedule to be executed from the Pareto optimal operation and maintenance schedule set according to the actual risk preference, and feed back the status rollback result and reward after the execution action to the policy network for closed-loop iterative update.

[0069] This embodiment integrates upstream continuous pollution risk index and micrometeorological environment data, and uses a probability mapping function containing a humid excitation coupling term to calculate the physically meaningful real-time flashover failure probability, improving the accuracy of risk identification under humid extreme conditions. Based on the operational status, it dynamically constructs a cost function for missed and false alarms, and adaptively optimizes thresholds using Bayesian decision criteria to achieve intelligent grading and suppress false alarms. It introduces a "waiting" strategy and establishes a set of constraints such as resource capacity, meteorological windows, and mutual exclusion of target insulator equipment to ensure the engineering feasibility of the operation and maintenance strategy. It establishes a multi-objective optimization model for risk and cost, and uses an improved deep reinforcement learning algorithm for rolling optimization, outputting a Pareto-optimal operation and maintenance schedule set under the dynamic weight tendency of each objective, supporting systematic trade-offs of different preferences. Finally, it achieves closed-loop iteration through feedback of execution results, forming an adaptive operation and maintenance closed loop from risk perception to decision execution.

[0070] In one embodiment, step S1: obtaining the continuous pollution risk index output by the upstream risk prediction model, collecting micrometeorological environmental data, and calculating the real-time flashover failure probability specifically includes the following steps:

[0071] Step S11: Receive the normalized continuous pollution risk index of the target insulator equipment within a preset time window, output by the upstream prediction model, denoted as:

[0072] Equation (1)

[0073] In the formula: i represents the target insulator or insulator string number; t represents the discrete time index; This represents the normalized continuous pollution risk index of target insulator device i at time t; where target insulator device refers to the target insulator or insulator string number.

[0074] Step S12: Collect micro-meteorological environment data corresponding to the location of the target insulator equipment, including relative humidity. Rainfall intensity Dew point difference The dew point difference is the difference between the dry bulb temperature and the dew point temperature.

[0075] Step S13: Construct a piecewise nonlinear effective wetting factor using micro-meteorological environmental data to characterize the surge in conductivity of the pollution layer under specific meteorological conditions. This follows the physical mechanism of "drizzle promotes flashover and heavy rainfall cleans insulators":

[0076] Equation (2)

[0077] The above formula defines the normalized effective wetting factor. The calculation method is a typical piecewise function, and its specific form depends on the real-time meteorological conditions:

[0078] In areas with no rain but extremely high humidity ( Under condensation or fog conditions, its value is determined by an exponential function of relative humidity RH(t), simulating the slow increase in conductivity caused by high humidity and moisture absorption.

[0079] When it rains but the intensity does not reach the washing threshold ( When the value is proportional to the rainfall intensity R(t), it simulates the surge effect of drizzle on the conductivity of the pollutant layer;

[0080] When the rainfall intensity is high enough ( When the factor is set to zero, it represents the scouring and washing effect of heavy rainfall.

[0081] In the absence of rain and with dry air ( Under the condition of ), its value decreases exponentially as the dew point difference increases, simulating the inhibition of moisture by drying.

[0082] In the formula: Normalized effective wetting factor; The threshold for determining condensation; To effectively flush the threshold; The wetting intensity coefficient under condensation / fog conditions; The wetting intensity coefficient under drizzle conditions; This is the baseline coefficient under drying conditions, and its value is relatively small. This is the humidity sensitivity coefficient, used to adjust the growth rate of the exponential function; The dew point sensitivity coefficient, its physical meaning is: when the dew point difference... Greater than At this point, the air is not saturated. As the dew point difference further increases, the air becomes drier, and the humidification factor decreases exponentially, approaching 0.

[0083] Step S14: Combining the fouling risk index and the effective wetting factor, calculate the flashover failure probability using a probability mapping function containing a coupling excitation term, as shown in the following formula:

[0084] Equation (3)

[0085] In the formula: Let be the probability of flashover failure of target insulator device i at time t, and its value range be [value range missing]. ; This is a linear prediction value from a logistic regression model, derived from the insulator pollution risk index. Effective moisturizing factors Probability mapping function of the relationship between them Calculated; where, The intercept term constant; The regression coefficient of the pollution index; The regression coefficient of the wetting factor; These are the coefficients of the coupling excitation term; The physical meaning of the humidification-induced coupling term is that the product term increases significantly only when "high contamination" and "high humidity" coexist, leading to an exponential increase in the probability of flashover failure, thereby accurately capturing the suddenness of flashover accidents.

[0086] Step S15: Based on the historical dataset, where flashover event samples are labeled as label 1 (positive samples) and normal operation samples are labeled as label 0 (negative samples), the regression coefficients of the fitted probability mapping model are calculated. When historical pollution flashover event samples are lacking, artificial pollution test data are used to initialize the regression coefficients, and the parameters are updated through incremental calibration. During operation, incremental calibration is triggered at a preset cycle to output the updated regression coefficients, ensuring the accuracy of probability mapping.

[0087] For each sample in the historical dataset, define the likelihood function. for:

[0088] Equation (4)

[0089] In the formula, It is the label (0 or 1) of sample j. The feature vector of the corresponding sample (including) , (and coupling terms), where N is the number of samples;

[0090] The regression coefficients of the probability mapping model are solved by minimizing the negative log-likelihood loss function, as shown in the following equation:

[0091] Equation (5)

[0092] The regression coefficients are iteratively updated using optimization algorithms (such as gradient descent or the Newton-Raphson method) until convergence. When historical accident samples are lacking, the coefficients are initialized using artificial contamination test data, for example, by obtaining label data through simulating high contamination and high humidity conditions, and then the above fitting is performed.

[0093] In one embodiment, step S2: acquiring power grid operation status data, constructing a decision cost function and optimizing the dynamic classification threshold according to the Bayesian criterion, comparing the real-time flashover failure probability with the dynamic classification threshold to perform risk classification, and outputting the risk classification result, specifically includes the following steps:

[0094] Step S21: Obtain the power grid operation status data from the power grid dispatch coefficient, denoted as... , which is an importance coefficient that changes over time, including operating mode, power supply reliability requirements, load importance, target insulator equipment importance or equivalent safety margin index, and becomes the situational operation risk weighting factor;

[0095] Step S22: Construct the costs of missed detections and false positives; further, this includes the following steps:

[0096] Step S221: The cost of underreporting changes with the operational status. Preferably, when... When the situation is indicated as high-risk (such as a critical power supply guarantee period), a higher penalty for underreporting should be set. More specifically, based on importance coefficient Dynamically calculate the cost of underreporting As shown in the following formula:

[0097] Equation (6)

[0098] Step S222: Cost of False Alarms It can be set to a constant or vary with maintenance resources and power outage costs. It can be set to the loss caused by false positives or the labor cost of a single invalid inspection.

[0099] Step S23: Based on the cost of underreporting Cost of misreporting For candidate thresholds The expected cost function based on the Bayesian decision criterion is constructed as follows:

[0100] Equation (7)

[0101] In the formula: Candidate threshold The expected cost at time t; This represents the false negative probability, i.e., the probability of an actual fault being less than the predicted probability. The probability of; This refers to the false alarm probability, which is the probability that the actual error is normal but the predicted probability is greater than 1. The probability of;

[0102] Step S24: Obtain the conditional probability distribution based on historical prediction error statistics or online rolling statistics. and And based on this, calculate at a given threshold The probability of missed reports and false alarm probability Output the probability evaluation results;

[0103] Step S25: Utilize the expected cost function Based on the probability evaluation results, a numerical optimization algorithm is applied to the threshold candidate set Θ to solve for the dynamic hierarchical threshold that minimizes the expected cost, as shown in the following equation:

[0104] Equation (8)

[0105] In the formula: The dynamic grading threshold at time t;

[0106] The optimal threshold As the operating status The principle of adaptive change is as follows: Figure 2 As shown, this demonstrates the dynamic characteristics of improving sensitivity during critical power supply periods and balancing the costs of false alarms and missed alarms during routine periods.

[0107] Step S26: Utilize Mapping For discrete risk levels; more specifically, compare the real-time flashover failure probability. and dynamic grading threshold Obtain discrete risk classification results ; Specifically, if If it is high risk, it is marked as high risk; otherwise, it is low risk.

[0108] In one embodiment, step S3: defining a set of operation and maintenance actions, introducing waiting as an active strategy, and establishing a set of constraints, specifically includes the following steps:

[0109] Step S31: Input typical operation types from the historical operation and maintenance action library. After processing and definition, construct an operation and maintenance action set containing five specific actions:

[0110] Equation (9)

[0111] in: For waiting actions; For rinsing with electrified water; Cleaning during power outage; For anti-flashover coating; To replace the insulator;

[0112] Simultaneously, a state rollback mechanism is established; specifically, each specific maintenance action 'a' (such as cleaning or coating) is mapped to a discrete rollback factor. This factor acts on the current pollution risk index of the target insulator equipment. This determines its state at the next moment. Its mathematical expression is:

[0113] Equation (10)

[0114] Among them, the function The effect of the action is defined; for example, after performing the cleaning action, It may be reset to a lower value or reduced proportionally;

[0115] Updated Filthy Status Combined with the effective moisturizing factor in the next moment The values ​​are input together into the risk probability mapping function g(⋅) to calculate the new flashover failure probability after the action is performed:

[0116] Equation (11)

[0117] This demonstrates that actions directly affect the risk level of the target insulator equipment by changing its state.

[0118] Compare the changes in risk probability of the target insulator before and after the action is executed, and calculate the change in expected system risk caused by the action by combining the failure consequence loss weight L(i) of the target insulator. Summing over all target insulators i yields the global risk increment:

[0119] Equation (12)

[0120] Ultimately, this risk increment with incremental operation and maintenance costs Both are included in the instant reward function This serves as a direct signal for evaluating the quality of an action.

[0121] Step S32: The core of establishing forward-looking meteorological correlation constraints lies in introducing a "forward-looking lock-in" mechanism to avoid sunk costs caused by weather disruptions during operations. This constraint is based on the preparation cycle of meteorological forecast data and operational actions. Make judgments, for example, large-scale maintenance work requires preparation of personnel, materials and power outage plans a week in advance.

[0122] Specifically, let the current decision time be t. If the weather forecast indicates a future planned operation window (from...) If there is intermittent weather such as thunderstorms, a three-tiered logic judgment process will be triggered:

[0123] Sunk cost determination: Specifically, the system identifies that if the operation is started as planned at time t, all the preparatory costs invested in the early stage (such as manpower allocation, material procurement, planning, etc.) will not generate the expected benefits due to the subsequent operation being forcibly interrupted by weather, and thus all become sunk costs.

[0124] Execution constraint determination: Based on the above determination, the system automatically prohibits the initiation of any operation requiring a long preparation period (≥ 10 ... ) maintenance actions.

[0125] Decision adjustment output: In the decision-making process of the optimization model, the available actions are restricted to "wait" or "short-cycle action", thus forming a weather window prohibition rule in the scheduling scheme.

[0126] This mechanism establishes a strong forward-looking correlation between meteorological conditions and operational execution, enabling proactive global risk avoidance based on forecasts. It can effectively reduce plan interruptions and resource waste caused by weather uncertainties, and significantly reduce ineffective operation and maintenance costs.

[0127] Step S33: Establish resource capacity constraints; specifically, establish resource data and available limits for personnel, vehicles, tools, etc., and set the resource usage per unit time not to exceed the limit to prevent over-allocation.

[0128] Step S34: Establish a meteorological operation window constraint and risk accumulation mechanism; specifically, establish a hard operation window constraint based on real-time weather forecasts and safety procedures. When it is determined that the current or recent weather period is a period of strong winds, thunderstorms or other weather conditions that do not meet the safety conditions for outdoor operations, a hard prohibition signal is automatically generated and output to physically prevent the execution of any outdoor operation and maintenance actions.

[0129] When the aforementioned hard constraints force the current action to be "wait," the pollution risk state of the target insulator will not be reduced. Instead, it will evolve naturally based on the upstream prediction model in step S1. This mechanism ensures that the "under-maintenance" state caused by the current window closing will be propagated to the next time step, forcing the policy network to weigh the higher maintenance costs after risk escalation in future feasible windows, thereby achieving global optimum.

[0130] Step S35: Establish mutual exclusion and minimum safe operating interval constraints for target insulator equipment; specifically, establish constraints such as prohibiting the parallel operation of the same target insulator equipment or related target insulator equipment at the same time period, and requiring a minimum interval between two operations on the same target insulator equipment; further:

[0131] Regarding mutual exclusion constraints for target insulator equipment, it is strictly prohibited for multiple maintenance actions to be performed concurrently on the same target insulator equipment within the same time period. This restriction takes into account both the physical characteristics of the target insulator equipment itself and the topological relationships between related target insulator equipment. When the system detects that a target insulator equipment has been scheduled to perform a certain task, it will automatically block scheduling requests for other concurrent actions on that target insulator equipment, thereby avoiding resource conflicts and operational risks.

[0132] Regarding the minimum safe working interval constraint, a time buffer requirement is set for different actions performed consecutively on the same target insulator equipment. This constraint requires that after the target insulator equipment completes the previous operation, a preset minimum time interval must elapse before the next operation can begin. The setting of this time interval takes into account multiple factors such as target insulator equipment restoration, safety inspection, and operation preparation, leaving necessary safety margins for subsequent operations.

[0133] The final set of constraints includes resource capacity constraints, forward-looking meteorological correlation constraints, meteorological operation window constraints, and mutual exclusion and minimum safe operation interval constraints for target insulator equipment. .

[0134] In one embodiment, step S4: based on risk status Dynamic grading threshold and constraint set A multi-objective optimization model is established with the goal of minimizing total maintenance costs and the system's cumulative expected risk. An improved actor-critic deep reinforcement learning algorithm is used for rolling optimization. Under the dynamic weight bias of each objective, a Pareto-optimal maintenance schedule set is output. The specific steps include:

[0135] Step S41: Establish the rolling optimization time domain; specifically, set the preset rolling window length as T and the number of target insulator devices as N;

[0136] Step S42: Total Operation and Maintenance Cost Target The goal is to minimize the total cost of all target insulator equipment within the rolling window, including direct material costs, labor costs, and power outage losses. Its mathematical expression is:

[0137] Equation (13)

[0138] In the formula: This is a comprehensive cost function that includes direct material costs, labor costs, and power outage losses; For time t, the first Maintenance actions taken for a target insulator device, such as cleaning and waiting;

[0139] S43: Establish the system's cumulative expected risk objective function; specifically, the system's cumulative expected risk objective function... Minimizing the cumulative value of the probability of flashover failure of the target insulator equipment and the weight of the resulting loss reflects the risk prevention and control requirements:

[0140] Equation (14)

[0141] In the formula: To accumulate expected risk targets for the system; The historical action sequence of the target insulator device i; The consequence loss weight for target insulator device i;

[0142] S44: Establish state reset or reduction update rules; dynamically update the pollution status of target insulator equipment based on the execution effect of maintenance actions.

[0143] When effective actions such as cleaning and coating are performed, the state of contamination is reset or reduced proportionally, such as... , where k is the reduction factor;

[0144] When the "wait" option is selected or no action is required, the state evolves naturally through the upstream prediction model. This mechanism ensures that risk assessment and action execution effects are synchronized in real time.

[0145] S45: Construct the feasible region; apply the constraint set Ω defined in step S3 (including constraints such as resource capacity, weather window, and mutual exclusion of target insulator equipment) to the decision variables. By using mathematical programming methods to limit the range of action values, we can ensure that all generated scheduling schemes meet the requirements of project executability. For example, resource constraints can be represented as:

[0146] Equation (15)

[0147] Where R(⋅) is the resource demand function, Let t be the resource limit at time t.

[0148] S46: Solved using an improved actor-critic deep reinforcement learning algorithm, the core mechanism of which includes the following sub-steps:

[0149] S4A: To handle the temporal dependence of the environment and its correlation with long-term meteorological conditions, a stacked state vector containing current observations and future forecasts is constructed to capture the temporal dependence and long-term meteorological correlation, denoted as:

[0150] Equation (16)

[0151] In the formula: This is used to quantify the risk level of each target insulator device at the current moment, in order to solve the gradient vanishing problem caused by the original probability value being too small; For the future The meteorological forecast feature sequence within the time window is used to give the model the forward-looking ability to perceive the risk of future weather disruptions such as thunderstorms, thereby avoiding sunk costs; This represents the current remaining resource status characteristics; A vector of risk and cost preference weights, set by the user, is used as a conditional input, enabling a single-policy network to output different operational preferences.

[0152] S4B: Constructing an actor-critic network architecture; specifically, establishing an Actor network as a policy network and a Critic network as a value network, wherein:

[0153] The input to the Actor network is a stacked state. The output layer is the dimension corresponding to the set of operation and maintenance actions. The original score tensor ,in, Size of the action set;

[0154] The Critic network input is in a stacked state. Output the value assessment of the current state. It is used to assist in the convergence of the policy network.

[0155] S4C: Defines the rules for generating additive action masks, which ensure the physical feasibility of decisions through strict integration of engineering constraints. Specifically, it generates a mask vector based on the aforementioned set of constraints. For operational actions that are currently feasible, their corresponding positions in the mask vector are set to zero; while for all explicitly prohibited actions, such as scheduling work during a closed weather window, their corresponding positions are set to a very large negative constant. This mask vector is then... The original action preferences output by the policy network By directly adding them, we get the masked output. When this result is used to calculate the final action probability distribution through the Softmax activation function, the probability value of the masked action will approach zero, thus being completely prohibited from selection at the algorithm level;

[0156] S4D: Yes Normalize it to transform it into a probability distribution of each action being selected. And determine the action to be performed through sampling or a greedy strategy. Specifically, it includes the following steps:

[0157] The value of allowed actions for each device is normalized using the Softmax function, calculated as follows:

[0158] Equation (17)

[0159] in, Indicates the current state The summation operation is performed only on action b within this set, ensuring that the probability of a prohibited action is zero.

[0160] Action Value Matrix From the original action value With mask matrix Adding them together, we get:

[0161] Equation (18)

[0162] Mask matrix The assignment rules are as follows, and their purpose is to "mask" actions that violate constraints before calculation:

[0163] Equation (19)

[0164] In the formula, To determine an operational action at decision time t. Whether a function is allowed; where It is a variable used to iterate through set A.

[0165] The decision-making rule for the actual action selection is as follows:

[0166] During the training phase, a random sampling strategy is adopted, denoted as... That is, actions are randomly selected based on probability distribution π to encourage the algorithm to explore more possibilities. In the application phase, a greedy strategy is adopted, denoted as... That is, directly select the action with the highest probability at present to ensure the stability and optimality of the decision; among which, It is based on strategy The specific action sampled or selected from action set A;

[0167] S4E: Design a comprehensive instant reward function to quantify the short-term benefits of actions and the penalties for violating constraints.

[0168] Equation (20)

[0169] In the formula: For instant rewards; This is the numerical scaling factor. This represents an increase in risk. This represents an increase in costs; and These are the preference weights for risk and cost, respectively; This is a constraint violation function, whose value is a very large constant much larger than the system's maximum possible cost. It is used when an action violates a weather window or preparation period constraint. It takes effect, providing strong negative feedback; These are constraint violation penalties, used to suppress actions that result in infeasible schedules;

[0170] Ultimately, you receive a comprehensive, instant reward. By quantifying the trade-off between "risk reduction" and "cost consumption" for each action and imposing constraints and penalties, the reinforcement learning agent is precisely guided to search for and learn strategies in rolling optimization, aiming to minimize the total operating cost and the cumulative risk of the system in the long term.

[0171] S4F: Defines the discrete time step of the simulation environment. It lasts 24 hours. One training round covers a complete sludge accumulation cycle. At each time step... Rewards based on environmental feedback and the state at the next moment The network parameters are updated using the policy gradient algorithm;

[0172] S4G: Weights for different preferences during the training phase Sampling is performed to enable the policy network to learn the mapping between "state-preference-action"; in the inference stage, a set of discrete preference vectors are input to generate multiple scheduling schemes, and the dominated schemes are eliminated to obtain the Pareto optimal operation and maintenance schedule set; the Pareto front diagram of the risk and cost trade-off obtained from this multi-objective optimization solution is shown in the figure below. Figure 3 As shown, each point represents a feasible operation and maintenance strategy, which decision-makers can choose according to their preferences.

[0173] S4H: Uses the inferred Pareto optimal operation and maintenance schedule for the current rolling window execution, generating, for example... Figure 4 The Gantt chart-style work order recommendations shown intuitively illustrate the specific operational and maintenance action sequence after considering weather windows and resource constraints. Scheduling results are fed back to the policy network in real time, supporting online adaptive updates.

[0174] This embodiment integrates multi-objective optimization and improved deep reinforcement learning to achieve a dynamic trade-off between risk and cost under strict constraints. Its output Pareto-optimal operation and maintenance schedule set provides executable decision-making solutions for different operation and maintenance preferences, and continuous optimization through closed-loop iteration ensures the system's adaptability and engineering practicality.

[0175] In one embodiment, step S5: selecting the target schedule from the Pareto optimal operation and maintenance schedule set according to the actual risk preference, and feeding back the state rollback result and reward after the execution action to the policy network for closed-loop iterative update, specifically includes the following steps:

[0176] S51: Based on the Pareto optimal operation and maintenance schedule set and the risk preference weight vector W, select the target schedule according to preference (such as cost priority or risk priority) and output the target execution schedule;

[0177] S52: Input target schedule and on-site execution data, execute actions such as cleaning and waiting, and collect status rollback results, such as risk reset, and reward feedback including risk increment and cost increment. Output the rollback result and reward signal;

[0178] S53: Utilizing rewards and feedback and the next state The actor-critic network parameters are updated using the policy gradient algorithm, and the optimized network parameters are output.

[0179] S54: Input new network parameters and real-time data, re-input the feedback results to steps S1 and S4, perform the next round of risk assessment and scheduling solution, output continuously optimized risk assessment and scheduling, and realize an adaptive closed-loop system.

[0180] Secondly, this application provides an insulator pollution risk classification and maintenance decision system, the system comprising: a data acquisition module, used to acquire the continuous pollution risk index output by the upstream risk prediction model and collect micro-meteorological environment data;

[0181] The probability calculation module communicates with the data acquisition module and is used to calculate the real-time flashover failure probability using a probability mapping function containing a wet excitation coupling term, and quantify and map it into a discrete risk state level. At the same time, it constructs a state rollback mechanism triggered by operation and maintenance actions.

[0182] The dynamic grading module communicates with the probability calculation module and is used to construct a decision cost function based on the real-time flashover failure probability output by the probability calculation module. It then optimizes the dynamic grading threshold according to the Bayesian decision criterion and performs risk grading to obtain the risk grading result.

[0183] The strategy construction module communicates with the probability calculation module and is used to define a set of operation and maintenance actions, including waiting, based on the state rollback mechanism and risk classification results, and to establish a set of constraints.

[0184] The optimization solution module communicates with the probability calculation module, dynamic grading module, and strategy construction module. It is used to establish a multi-objective optimization model based on risk status, dynamic grading threshold, and constraint set. It adopts an improved deep reinforcement learning algorithm for rolling optimization and solves the problem. Under the dynamic weight tendency of each objective, it outputs a Pareto optimal operation and maintenance schedule set.

[0185] The decision execution module communicates with the optimization solution module and is used to select the execution target schedule based on risk preference, and feed back the state rollback results and rewards to the policy network to achieve closed-loop iterative updates.

[0186] The functions of each module in the above-mentioned insulator pollution risk classification and maintenance decision system correspond to the steps in the above-mentioned insulator pollution risk classification and maintenance decision method embodiment, and their functions and implementation processes will not be described in detail here.

[0187] Thirdly, embodiments of this application provide an insulator pollution risk classification and maintenance decision-making device, which can be a personal computer (PC), laptop computer, server, or other device with data processing capabilities.

[0188] In this embodiment, the insulator pollution risk classification and maintenance decision-making device may include a processor, a memory, a communication interface, and a communication bus.

[0189] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0190] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting devices within the insulator pollution risk assessment and maintenance decision-making equipment, as well as interfaces used for interconnecting the insulator pollution risk assessment and maintenance decision-making equipment with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.

[0191] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0192] The processor can be a general-purpose processor, which can call the insulator pollution risk classification and maintenance decision program stored in the memory and execute the insulator pollution risk classification and maintenance decision method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the insulator pollution risk classification and maintenance decision program is called can refer to the various embodiments of the insulator pollution risk classification and maintenance decision method of this application, which will not be repeated here.

[0193] Fourthly, embodiments of this application also provide a readable storage medium.

[0194] The present application has a storage medium storing an insulator pollution risk classification and maintenance decision program, wherein when the insulator pollution risk classification and maintenance decision program is executed by a processor, the steps of the insulator pollution risk classification and maintenance decision method as described above are implemented.

[0195] The method implemented when the insulator pollution risk classification and maintenance decision-making procedure is executed can be referred to in various embodiments of the insulator pollution risk classification and maintenance decision-making method of this application, and will not be repeated here.

[0196] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for classifying insulator pollution risk and making maintenance decisions, characterized in that, The method includes the following steps: Obtain the continuous pollution risk index output by the upstream risk prediction model, collect micrometeorological environmental data, and calculate the real-time flashover failure probability. The system acquires power grid operation status data, constructs a decision cost function, optimizes the dynamic classification threshold based on the Bayesian criterion, compares the real-time flashover failure probability with the dynamic classification threshold to classify the risk, and outputs the risk classification result. Define a set of operational actions, including waiting, based on the status rollback mechanism and risk classification results, and establish a set of constraints. Based on risk status, dynamic classification thresholds, and constraint sets, a multi-objective optimization model is established with the goal of minimizing total maintenance cost and system cumulative expected risk. An improved actor-critic deep reinforcement learning algorithm is used for rolling optimization solution, and Pareto optimal maintenance schedule set is output under the dynamic weight tendency of each objective. Based on the actual risk preference, the target schedule is selected from the Pareto optimal operation and maintenance schedule set, and the status rollback results and rewards after the execution are fed back to the policy network for closed-loop iterative updates.

2. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The real-time flashover failure probability is calculated using a probability mapping function containing a wetting excitation coupling term, as shown in the following equation: ; In the formula, This is the linear prediction value of the logistic regression model. ; Let be the flashover failure probability of target insulator device i at time t. As a continuous pollution risk index, For regression coefficients, This is to characterize the wetting-induced coupling term that represents the simultaneous existence of high contamination and high humidity.

3. The insulator pollution risk classification and maintenance decision-making method according to claim 2, characterized in that, The effective wetting factor The construction logic satisfies: ; In the formula: Normalized effective wetting factor; The threshold for determining condensation; To effectively flush the threshold; Relative humidity; For dew point difference; Rainfall intensity; The wetting intensity coefficient under condensation / fog conditions; The wetting intensity coefficient under drizzle conditions; This serves as the baseline coefficient under drying conditions. Humidity sensitivity coefficient; This represents the drying sensitivity coefficient.

4. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The process of obtaining the dynamic classification threshold based on the Bayesian decision criterion specifically includes the following steps: The threshold decision expected cost function is constructed as follows: ; In the formula, Threshold The expected cost at time t, The cost of underreporting as operational status changes, To pay the price for false alarms This represents the probability of missed reports. This represents the false alarm probability. Solving in the threshold candidate set makes smallest As a dynamic grading threshold.

5. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The set of constraints specifically includes: Resource capacity constraints are used to ensure that the total amount of resources required for operation and maintenance actions performed per unit of time does not exceed the upper limit of available resources; Forward-looking meteorological correlation constraints are used to dynamically determine whether to initiate long-cycle operation and maintenance actions based on the operation preparation cycle and future weather forecasts. If the forecast shows that there is a risk of interruption during the operation window, the option to wait or take short-cycle actions will be forced. Meteorological operation window constraints are used to prohibit outdoor maintenance operations during periods of strong winds, thunderstorms, or when safety procedures are not met, based on meteorological forecast data. The mutual exclusion of target insulator equipment and the minimum safe working interval constraint are used to prevent the same target insulator equipment from performing multiple actions in parallel at the same time period, and to ensure that the minimum time interval is met between two actions of the same target insulator equipment.

6. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The objective function of the multi-objective optimization model includes the total maintenance cost objective. and system cumulative expected risk target As shown in the following formula: ; ; In the formula, For the comprehensive cost function, For conditional flashover failure probability, For the first Weighting of consequences and losses for each target insulator device The target insulator equipment's historical action sequence; For time t, the first The operation and maintenance actions taken for each target insulator device, where T is the preset rolling window length and N is the number of target insulator devices.

7. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The improved actor-critic deep reinforcement learning algorithm includes the following mechanisms: The state space construction mechanism uses stacked state vectors to integrate the quantized current risk state, weather forecast features of future time windows, and manually set preference weight vectors. The reward function construction and numerical scaling mechanism define a comprehensive reward function that includes state fallback rewards, cost consumption, and constraint penalty terms, and introduces a numerical scaling factor to solve the gradient vanishing problem; the constraint handling mechanism uses penalty terms to handle weather and resource constraints, generating strong negative penalty signals for violations; and, The action space processing mechanism outputs the action probability distribution through the actor network and uses the critic network to evaluate the state value.

8. The insulator pollution risk classification and maintenance decision-making method according to claim 1, characterized in that, The closed-loop iterative update includes: After the target insulator equipment has undergone cleaning, sweeping or coating operations, the contamination status is reset or reduced in the risk assessment at the next moment. When the target insulator equipment performs actions that do not require processing, the risk accumulates naturally according to the upstream predicted trend. The status after maintenance is reported back to the policy network for closed-loop iterative updates.

9. A pollution risk classification and maintenance decision-making system for insulators, characterized in that, include: The data acquisition module is used to acquire the continuous pollution risk index output by the upstream risk prediction model and collect micrometeorological environmental data. The probability calculation module communicates with the data acquisition module and is used to calculate the real-time flashover failure probability using a probability mapping function containing a wet excitation coupling term, and quantify and map it into a discrete risk state level. At the same time, it constructs a state rollback mechanism triggered by operation and maintenance actions. The dynamic grading module communicates with the probability calculation module and is used to construct a decision cost function based on the real-time flashover failure probability output by the probability calculation module. It then optimizes the dynamic grading threshold according to the Bayesian decision criterion and performs risk grading to obtain the risk grading result. The strategy construction module communicates with the probability calculation module and is used to define a set of operation and maintenance actions, including waiting, based on the state rollback mechanism and risk classification results, and to establish a set of constraints. The optimization solution module communicates with the probability calculation module, dynamic grading module, and strategy construction module. It is used to establish a multi-objective optimization model based on risk status, dynamic grading threshold, and constraint set. It adopts an improved deep reinforcement learning algorithm for rolling optimization and solves the problem. Under the dynamic weight tendency of each objective, it outputs a Pareto optimal operation and maintenance schedule set. The decision execution module communicates with the optimization solution module and is used to select the execution target schedule based on risk preference, and feed back the state rollback results and rewards to the policy network to achieve closed-loop iterative updates.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an insulator pollution risk classification and maintenance decision program, wherein when the insulator pollution risk classification and maintenance decision program is executed by a processor, it implements the steps of the insulator pollution risk classification and maintenance decision method as described in any one of claims 1 to 8.