A method and system for reducing the closed-loop cycle management of power transmission line external damage hidden dangers from discovery to disposal

CN122736373APending Publication Date: 2026-09-11STATE GRID SHANDONG ELECTRIC POWER CO PINGDU POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610620862.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

输电线路外力破坏隐患具有突发性强、分布范围广、隐蔽性高的显著特点,给线路运维带来了极大挑战

Benefits of technology

[0044]To achieve a controllable and manageable closed-loop cycle for hazard handling, differentiated target closed-loop cycle benchmarks are set, and handling time limits are dynamically generated and continuously optimized. Combined with process prediction and deviation identification, the closed-loop cycle is transformed from passive statistics to active control, effectively solving the problems of uncontrollable and inconsistent handling time in traditional management and control, and ensuring that hazard handling is completed on time and efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736373A_ABST
    Figure CN122736373A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of power system operation and maintenance, and particularly relates to a closed-loop cycle management and control method and system for reducing the hidden danger of external damage to power transmission lines from discovery to disposal. The method includes setting differentiated target closed-loop cycle benchmarks, quantifying comprehensive risks of hidden dangers and initially allocating resources, dynamically generating and rolling optimization of disposal time limits, predicting disposal deviations, optimizing resource allocation, carrying out quality acceptance and parameter iterative optimization. The system includes multiple modules such as data acquisition, cycle benchmark setting, risk quantification, etc., and realizes full-process coverage of data support, process management and control, and parameter optimization. The present application can realize active regulation of hidden danger disposal time, optimize resource allocation, and form a continuously optimized management and control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system operation and maintenance technology, and in particular relates to a closed-loop cycle management method and system for reducing the potential for external damage to transmission lines from discovery to handling. Background Technology

[0002] Transmission lines are a core component of the power system, and their safe and stable operation is directly related to the reliability of power supply. External force damage to transmission lines is characterized by its suddenness, wide distribution, and high degree of concealment, posing a significant challenge to line operation and maintenance.

[0003] In traditional management and control models, the process of "notification-monitoring-acceptance" is often linear after a hazard is discovered, with each step operating independently and lacking systematic control over the overall duration of hazard handling. Currently, mainstream management and control methods in the industry include fixed-time notification, manual monitoring, post-event statistics, and electronic work order records. These methods all use the closed-loop cycle as a post-event statistical indicator after hazard elimination, only calculating the time after handling is completed, and cannot provide real-time intervention and control during the handling process. This leads to inconsistent hazard handling times, a lack of unified and differentiated control standards, and a mismatch between handling resource investment and hazard risk levels. High-risk hazards are not handled promptly, while low-risk hazards suffer from wasted resources. Furthermore, the management and control process lacks an effective feedback and optimization mechanism, making it difficult to continuously improve handling strategies based on historical data, and failing to meet the needs of lean operation and maintenance of transmission lines. Summary of the Invention

[0004] Purpose of the invention

[0005] The present invention aims to establish a management and control mechanism with the closed-loop cycle as the core control variable, transforming the closed-loop cycle from a post-event statistical indicator into a process control target, thereby enabling proactive regulation of the time required to handle potential hazards, optimizing resource allocation, and forming a continuously optimized management and control system.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A closed-loop management method for reducing the potential for external damage to power transmission lines from detection to handling includes:

[0008] Cycle benchmark setting: Based on line voltage level, regional risk characteristics and historical handling data, establish differentiated target closed-loop cycle benchmark values;

[0009] Risk quantification and initial resource allocation: Construct a multi-dimensional comprehensive risk scoring model for hidden dangers, integrate voltage level, construction type, construction duration, mechanical risk and time series characteristics to calculate a comprehensive risk score, and determine the initial resource matching degree based on the score and the frequency of handling similar hidden dangers in the past.

[0010] Dynamic time limit generation: Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial processing time limit is determined through a nonlinear mapping relationship, and an environmental disturbance correction factor is introduced for dynamic adjustment;

[0011] Process prediction and deviation identification: During the handling of potential hazards, the time elapsed is monitored in real time, and the Kalman filter algorithm is used to predict the remaining handling time to calculate the estimated closed-loop cycle, thereby identifying cycle control deviations;

[0012] Time limit rolling optimization: Construct a fuzzy adaptive control mechanism to dynamically adjust the subsequent processing time limit based on the periodic control deviation and its changing trend;

[0013] Resource optimization and allocation: Establish a multi-objective resource optimization and allocation model with the optimization objectives of disposal cost, risk reduction rate and cycle achievement rate. Solve the optimal resource allocation scheme through genetic algorithm, and implement a hierarchical resource allocation strategy based on the cycle control deviation range.

[0014] Quality Acceptance Evaluation: Construct a multi-dimensional closed-loop quality scoring system to comprehensively evaluate the effectiveness of hazard handling. When the score reaches the qualified threshold, the acceptance is deemed qualified; otherwise, it is returned for re-handling.

[0015] Parameter optimization: Based on a deep reinforcement learning framework, with the goal of minimizing the closed-loop cycle deviation, the risk score weight, resource matching influence coefficient, time limit correction coefficient, and target closed-loop cycle benchmark value are optimized online iteratively.

[0016] The periodic reference setting specifically includes:

[0017] First, three core influencing factors are selected: voltage level, regional risk level, and historical handling efficiency. The quantitative standards for each factor are then defined. Next, based on the basic closed-loop cycle, the weight coefficients and quantitative values ​​of each factor are combined to calculate the differentiated target closed-loop cycle benchmark value through weighted summation. Finally, the target closed-loop cycle benchmark value is calibrated quarterly using the latest historical handling data to ensure that the benchmark value is timely and accurate.

[0018] The aforementioned risk quantification and initial resource allocation specifically include:

[0019] A risk assessment index system was constructed, comprising five core dimensions: voltage level, construction type, construction duration, mechanical risk, and time series characteristics. The grading and quantification standards for each dimension were clearly defined. A multi-dimensional comprehensive risk scoring model for potential hazards was then constructed. The weights of each dimension were determined using the analytic hierarchy process (AHP), and a weighted summation method was used to calculate the comprehensive risk score. Subsequently, the initial resource matching degree was calculated by combining the comprehensive risk score with the average monthly handling frequency of similar hazards over the past 12 months through linear fitting. Finally, based on the initial resource matching degree, three resource allocation levels—low, medium, and high—were defined, and corresponding maintenance personnel, equipment, and materials were allocated to ensure that resource allocation matched the risk level of the potential hazard.

[0020] The dynamic time limit generation specifically includes:

[0021] Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial response time limit is calculated using an exponential nonlinear mapping relationship. Subsequently, an environmental disturbance correction factor is introduced, which integrates three types of environmental factors: weather, traffic, and terrain. By quantifying each environmental factor and setting environmental impact weight coefficients, a calculation formula for the environmental disturbance correction factor is constructed. The initial response time limit is then multiplied by the environmental disturbance correction factor to obtain the final dynamic response time limit. Finally, upper and lower limits are set to constrain the time limit range and avoid insufficient response due to excessively short time limits or increased risk due to excessively long time limits.

[0022] The process prediction and deviation identification specifically include:

[0023] During the hazard handling process, the used handling time is collected in real time at a fixed frequency. The remaining handling time is used as the state variable to construct the Kalman filter state equation and observation equation. The remaining handling time is predicted through a two-step iterative operation of prediction and update. The predicted remaining time is added to the used time to obtain the estimated closed-loop cycle. The cycle control deviation is obtained by calculating the difference between the estimated cycle and the dynamic handling time limit. Based on the preset deviation threshold range, the cycle control deviation is divided into three categories: advance deviation, normal deviation, and overtime deviation, thereby realizing real-time deviation identification of the handling cycle.

[0024] The time-limited rolling optimization specifically includes:

[0025] First, the periodic control deviation and its changing trend are defined as fuzzy control inputs, and the time limit adjustment is defined as the output. The fuzzy linguistic variables and quantization ranges of each quantity are then clarified. Next, fuzzy control rules are formulated based on operational experience. The Mamdani fuzzy inference method is used in conjunction with triangular membership functions for inference. The specific time limit adjustment is obtained by defuzzification using the centroid method. Finally, a new handling time limit is calculated based on the time limit adjustment to ensure that it meets the preset time limit boundary constraints. Rolling optimization is performed every 2 hours to dynamically adjust the subsequent handling time limit.

[0026] The aforementioned resource optimization and allocation specifically includes:

[0027] A multi-objective optimization model was constructed, with disposal cost, risk reduction rate, and cycle achievement rate as optimization objectives. This model was transformed into a single-objective function using a weighted sum method, and three types of constraints were set: total resource volume, disposal capacity, and non-negativity. Subsequently, a genetic algorithm was used to solve the optimization model. By initializing the population, setting the fitness function, performing genetic operations, and iteratively terminating, the optimal resource allocation scheme was obtained. Finally, based on the cycle control deviation range, a tiered resource allocation strategy was implemented: early deviations reduced resource allocation to lower costs, normal deviations maintained optimal allocation, and overdue deviations increased resource allocation to accelerate progress, ensuring precise matching of resource allocation with disposal needs and control objectives.

[0028] The quality acceptance evaluation specifically includes:

[0029] First, a multi-dimensional quality evaluation system is constructed, comprising four core dimensions: compliance of handling, degree of hazard elimination, equipment status restoration, and document completeness. The scoring criteria for each dimension are clearly defined. Then, a weighted summation method is used to calculate the comprehensive quality score by combining the preset weights of each dimension. Finally, a pass / fail threshold is set. If the score is greater than the pass / fail threshold, the acceptance is deemed qualified and the hazard handling is completed. If the score is less than the pass / fail threshold, the acceptance is deemed unqualified, and the handling is repeated until the acceptance is qualified, ensuring that the hazard handling effect meets the requirements.

[0030] The parameter optimization specifically includes:

[0031] A deep reinforcement learning model (DQN) is used to construct an interaction model between the agent and the hazard management environment. A reward function is designed with minimizing the closed-loop cycle deviation as the optimization objective, and a corresponding neural network structure with input, hidden, and output layers is built. Samples are sampled through agent-environment interaction and stored in an experience replay pool. After meeting the sample size threshold, gradient descent is used to update the network parameters. The target network is periodically synchronized to ensure training stability. The iteration terminates when the mean closed-loop cycle deviation reaches the target and remains stable. The optimized risk score weights, resource matching influence coefficients, time limit correction coefficients, and the target closed-loop cycle baseline value are output. This iterative optimization process is repeated quarterly, applying the optimized parameters to subsequent management processes to continuously improve the management accuracy of the closed-loop cycle.

[0032] A closed-loop management system for reducing the risk of external damage to power transmission lines from detection to handling includes:

[0033] The data acquisition module is used to collect data on transmission line voltage levels, regional risk characteristics, historical hazard handling data, on-site construction information of hazards, environmental parameters and handling process data, providing data support for each module;

[0034] The cycle benchmark setting module, connected to the data acquisition module, is used to select three core factors: voltage level, regional risk level, and historical handling efficiency, clarify quantitative standards, calculate differentiated target closed-loop cycle benchmark values ​​through weighted summation, and calibrate them periodically.

[0035] The risk quantification and initial resource allocation module, connected to the data acquisition module and the periodic benchmark setting module, is used to construct a multi-dimensional risk assessment indicator system, calculate the comprehensive risk score of hidden dangers, determine the initial resource matching degree by combining the historical handling frequency, and complete the initial resource allocation.

[0036] The dynamic time limit generation module is connected to the cycle benchmark setting module and the risk quantification and resource initial allocation module. It is used to calculate the initial treatment time limit through exponential nonlinear mapping, introduce environmental disturbance correction factors for dynamic adjustment, and set time limit boundary constraints.

[0037] The process prediction and deviation identification module is connected to the data acquisition module and the dynamic time limit generation module. It is used to monitor the time elapsed in real time, predict the remaining time through the Kalman filter algorithm, calculate the estimated closed-loop cycle and identify the cycle control deviation.

[0038] The rolling optimization module is connected to the process prediction and deviation identification module to build a fuzzy adaptive control mechanism. Based on the periodic control deviation and its changing trend, the time limit adjustment amount is obtained through fuzzy inference and defuzzification to realize the rolling optimization of the disposal time limit.

[0039] The resource optimization and allocation module is connected to the process prediction and deviation identification module and the risk quantification and initial resource allocation module. It is used to build a multi-objective resource optimization model, solve the optimal allocation scheme through a genetic algorithm, and perform hierarchical resource allocation based on the deviation interval.

[0040] The quality acceptance and evaluation module is connected to the resource optimization and allocation module. It is used to build a multi-dimensional quality evaluation system, calculate the comprehensive quality score, determine the acceptance result based on the pass threshold, and trigger the reprocessing process if it fails to pass.

[0041] The parameter optimization module is connected to the aforementioned modules. It uses DQN deep reinforcement learning to build an interactive model, iterates and optimizes various control parameters online, and updates and applies them to the control process every quarter.

[0042] The data storage module is used to store all data collected, calculated, and optimized by each module, ensuring data traceability and providing support for subsequent management and parameter optimization.

[0043] The present invention has the following beneficial effects:

[0044] To achieve a controllable and manageable closed-loop cycle for hazard handling, differentiated target closed-loop cycle benchmarks are set, and handling time limits are dynamically generated and continuously optimized. Combined with process prediction and deviation identification, the closed-loop cycle is transformed from passive statistics to active control, effectively solving the problems of uncontrollable and inconsistent handling time in traditional management and control, and ensuring that hazard handling is completed on time and efficiently.

[0045] Optimize the rationality of resource allocation by accurately matching initial resources through a multi-dimensional risk quantification model, combined with a multi-objective resource optimization and allocation model and a hierarchical allocation strategy, to achieve a precise correspondence between resource input and the level of hidden danger risk, effectively avoiding the phenomenon of untimely handling of high-risk hidden dangers and waste of resources for low-risk hidden dangers, and reducing handling costs.

[0046] A continuous feedback and optimization control system is constructed. Based on a deep reinforcement learning framework, control parameters are iteratively optimized online. Combined with a quality acceptance and evaluation mechanism, the effectiveness of the response is guaranteed, forming a closed loop of "setting-execution-feedback-optimization" to continuously improve the scientificity and adaptability of control strategies. Attached Figure Description

[0047] Figure 1 This is a flowchart of a closed-loop cycle management method for reducing the potential for external force damage to power transmission lines, from detection to handling, as proposed in this invention. Detailed Implementation

[0048] The present invention will be further described in detail below with reference to specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Example 1:

[0050] A closed-loop management method for reducing the potential for external damage to power transmission lines from detection to handling includes:

[0051] Cycle benchmark setting: Based on line voltage level, regional risk characteristics and historical handling data, establish differentiated target closed-loop cycle benchmark values;

[0052] Risk quantification and initial resource allocation: Construct a multi-dimensional comprehensive risk scoring model for hidden dangers, integrate voltage level, construction type, construction duration, mechanical risk and time series characteristics to calculate a comprehensive risk score, and determine the initial resource matching degree based on the score and the frequency of handling similar hidden dangers in the past.

[0053] Dynamic time limit generation: Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial processing time limit is determined through a nonlinear mapping relationship, and an environmental disturbance correction factor is introduced for dynamic adjustment;

[0054] Process prediction and deviation identification: During the handling of potential hazards, the time elapsed is monitored in real time, and the Kalman filter algorithm is used to predict the remaining handling time to calculate the estimated closed-loop cycle, thereby identifying cycle control deviations;

[0055] Time limit rolling optimization: Construct a fuzzy adaptive control mechanism to dynamically adjust the subsequent processing time limit based on the periodic control deviation and its changing trend;

[0056] Resource optimization and allocation: Establish a multi-objective resource optimization and allocation model with the optimization objectives of disposal cost, risk reduction rate and cycle achievement rate. Solve the optimal resource allocation scheme through genetic algorithm, and implement a hierarchical resource allocation strategy based on the cycle control deviation range.

[0057] Quality Acceptance Evaluation: Construct a multi-dimensional closed-loop quality scoring system to comprehensively evaluate the effectiveness of hazard handling. When the score reaches the qualified threshold, the acceptance is deemed qualified; otherwise, it is returned for re-handling.

[0058] Parameter optimization: Based on a deep reinforcement learning framework, with the goal of minimizing the closed-loop cycle deviation, the risk score weight, resource matching influence coefficient, time limit correction coefficient, and target closed-loop cycle benchmark value are optimized online iteratively.

[0059] The periodic reference setting specifically includes:

[0060] First, three core influencing factors are selected: voltage level, regional risk level, and historical handling efficiency. The quantitative standards for each factor are then defined. Next, based on the basic closed-loop cycle, the weight coefficients and quantitative values ​​of each factor are combined to calculate the differentiated target closed-loop cycle benchmark value through weighted summation. Finally, the target closed-loop cycle benchmark value is calibrated quarterly using the latest historical handling data to ensure that the benchmark value is timely and accurate.

[0061] The aforementioned risk quantification and initial resource allocation specifically include:

[0062] A risk assessment index system was constructed, comprising five core dimensions: voltage level, construction type, construction duration, mechanical risk, and time series characteristics. The grading and quantification standards for each dimension were clearly defined. A multi-dimensional comprehensive risk scoring model for potential hazards was then constructed. The weights of each dimension were determined using the analytic hierarchy process (AHP), and a weighted summation method was used to calculate the comprehensive risk score. Subsequently, the initial resource matching degree was calculated by combining the comprehensive risk score with the average monthly handling frequency of similar hazards over the past 12 months through linear fitting. Finally, based on the initial resource matching degree, three resource allocation levels—low, medium, and high—were defined, and corresponding maintenance personnel, equipment, and materials were allocated to ensure that resource allocation matched the risk level of the potential hazard.

[0063] The dynamic time limit generation specifically includes:

[0064] Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial response time limit is calculated using an exponential nonlinear mapping relationship. Subsequently, an environmental disturbance correction factor is introduced, which integrates three types of environmental factors: weather, traffic, and terrain. By quantifying each environmental factor and setting environmental impact weight coefficients, a calculation formula for the environmental disturbance correction factor is constructed. The initial response time limit is then multiplied by the environmental disturbance correction factor to obtain the final dynamic response time limit. Finally, upper and lower limits are set to constrain the time limit range and avoid insufficient response due to excessively short time limits or increased risk due to excessively long time limits.

[0065] The process prediction and deviation identification specifically include:

[0066] During the hazard handling process, the used handling time is collected in real time at a fixed frequency. The remaining handling time is used as the state variable to construct the Kalman filter state equation and observation equation. The remaining handling time is predicted through a two-step iterative operation of prediction and update. The predicted remaining time is added to the used time to obtain the estimated closed-loop cycle. The cycle control deviation is obtained by calculating the difference between the estimated cycle and the dynamic handling time limit. Based on the preset deviation threshold range, the cycle control deviation is divided into three categories: advance deviation, normal deviation, and overtime deviation, thereby realizing real-time deviation identification of the handling cycle.

[0067] The time-limited rolling optimization specifically includes:

[0068] First, the periodic control deviation and its changing trend are defined as fuzzy control inputs, and the time limit adjustment is defined as the output. The fuzzy linguistic variables and quantization ranges of each quantity are then clarified. Next, fuzzy control rules are formulated based on operational experience. The Mamdani fuzzy inference method is used in conjunction with triangular membership functions for inference. The specific time limit adjustment is obtained by defuzzification using the centroid method. Finally, a new handling time limit is calculated based on the time limit adjustment to ensure that it meets the preset time limit boundary constraints. Rolling optimization is performed every 2 hours to dynamically adjust the subsequent handling time limit.

[0069] The aforementioned resource optimization and allocation specifically includes:

[0070] A multi-objective optimization model was constructed, with disposal cost, risk reduction rate, and cycle achievement rate as optimization objectives. This model was transformed into a single-objective function using a weighted sum method, and three types of constraints were set: total resource volume, disposal capacity, and non-negativity. Subsequently, a genetic algorithm was used to solve the optimization model. By initializing the population, setting the fitness function, performing genetic operations, and iteratively terminating, the optimal resource allocation scheme was obtained. Finally, based on the cycle control deviation range, a tiered resource allocation strategy was implemented: early deviations reduced resource allocation to lower costs, normal deviations maintained optimal allocation, and overdue deviations increased resource allocation to accelerate progress, ensuring precise matching of resource allocation with disposal needs and control objectives.

[0071] The quality acceptance evaluation specifically includes:

[0072] First, a multi-dimensional quality evaluation system is constructed, comprising four core dimensions: compliance of handling, degree of hazard elimination, equipment status restoration, and document completeness. The scoring criteria for each dimension are clearly defined. Then, a weighted summation method is used to calculate the comprehensive quality score by combining the preset weights of each dimension. Finally, a pass / fail threshold is set. If the score is greater than the pass / fail threshold, the acceptance is deemed qualified and the hazard handling is completed. If the score is less than the pass / fail threshold, the acceptance is deemed unqualified, and the handling is repeated until the acceptance is qualified, ensuring that the hazard handling effect meets the requirements.

[0073] The parameter optimization specifically includes:

[0074] A deep reinforcement learning model (DQN) is used to construct an interaction model between the agent and the hazard management environment. A reward function is designed with minimizing the closed-loop cycle deviation as the optimization objective, and a corresponding neural network structure with input, hidden, and output layers is built. Samples are sampled through agent-environment interaction and stored in an experience replay pool. After meeting the sample size threshold, gradient descent is used to update the network parameters. The target network is periodically synchronized to ensure training stability. The iteration terminates when the mean closed-loop cycle deviation reaches the target and remains stable. The optimized risk score weights, resource matching influence coefficients, time limit correction coefficients, and the target closed-loop cycle baseline value are output. This iterative optimization process is repeated quarterly, applying the optimized parameters to subsequent management processes to continuously improve the management accuracy of the closed-loop cycle.

[0075] Example 2:

[0076] A closed-loop management system for reducing the risk of external damage to power transmission lines from detection to handling includes:

[0077] The data acquisition module is used to collect data on transmission line voltage levels, regional risk characteristics, historical hazard handling data, on-site construction information of hazards, environmental parameters and handling process data, providing data support for each module;

[0078] The cycle benchmark setting module, connected to the data acquisition module, is used to select three core factors: voltage level, regional risk level, and historical handling efficiency, clarify quantitative standards, calculate differentiated target closed-loop cycle benchmark values ​​through weighted summation, and calibrate them periodically.

[0079] The risk quantification and initial resource allocation module, connected to the data acquisition module and the periodic benchmark setting module, is used to construct a multi-dimensional risk assessment indicator system, calculate the comprehensive risk score of hidden dangers, determine the initial resource matching degree by combining the historical handling frequency, and complete the initial resource allocation.

[0080] The dynamic time limit generation module is connected to the cycle benchmark setting module and the risk quantification and resource initial allocation module. It is used to calculate the initial treatment time limit through exponential nonlinear mapping, introduce environmental disturbance correction factors for dynamic adjustment, and set time limit boundary constraints.

[0081] The process prediction and deviation identification module is connected to the data acquisition module and the dynamic time limit generation module. It is used to monitor the time elapsed in real time, predict the remaining time through the Kalman filter algorithm, calculate the estimated closed-loop cycle and identify the cycle control deviation.

[0082] The rolling optimization module is connected to the process prediction and deviation identification module to build a fuzzy adaptive control mechanism. Based on the periodic control deviation and its changing trend, the time limit adjustment amount is obtained through fuzzy inference and defuzzification to realize the rolling optimization of the disposal time limit.

[0083] The resource optimization and allocation module is connected to the process prediction and deviation identification module and the risk quantification and initial resource allocation module. It is used to build a multi-objective resource optimization model, solve the optimal allocation scheme through a genetic algorithm, and perform hierarchical resource allocation based on the deviation interval.

[0084] The quality acceptance and evaluation module is connected to the resource optimization and allocation module. It is used to build a multi-dimensional quality evaluation system, calculate the comprehensive quality score, determine the acceptance result based on the pass threshold, and trigger the reprocessing process if it fails to pass.

[0085] The parameter optimization module is connected to the aforementioned modules. It uses DQN deep reinforcement learning to build an interactive model, iterates and optimizes various control parameters online, and updates and applies them to the control process every quarter.

[0086] The data storage module is used to store all data collected, calculated, and optimized by each module, ensuring data traceability and providing support for subsequent management and parameter optimization.

[0087] Application example:

[0088] A closed-loop management method for reducing the potential for external damage to power transmission lines from detection to handling, comprising the following steps:

[0089] Step 1: Setting the Periodic Reference

[0090] Based on the differences in voltage levels of transmission lines, the risk characteristics of the area (such as population density, construction activity, and weather conditions), and historical data on hazard handling, a differentiated target closed-loop cycle benchmark value is established to provide a basis for subsequent time-limited management. The specific implementation is as follows:

[0091] 1.1 Determining the benchmark impact factors: Three types of core impact factors were selected, namely voltage level factors. Regional risk level factors Historical disposal efficiency factor ;

[0092] 1.2 Constructing the benchmark value calculation model: A weighted summation model is adopted, with the target closed-loop period benchmark value... The calculation formula is:

[0093]

[0094] in:

[0095] : Basic closed-loop cycle (unit: h), set according to industry standards and line operation and maintenance experience, with a value range of [24, 72] and a default value of 48h;

[0096] Voltage level weighting coefficient, with a value range of [0.8, 1.5]. The higher the voltage level, the greater the weight (1.5 for 500kV, 1.2 for 220kV, 1.0 for 110kV, and 0.8 for 35kV).

[0097] Voltage level quantification values: 1.0 for 500kV, 0.8 for 220kV, 0.6 for 110kV, and 0.4 for 35kV;

[0098] Regional risk level weighting coefficient, with a value range of [0.9, 1.6]. The higher the risk, the greater the weight.

[0099] Regional risk quantification value is divided into 5 levels (1-5) based on regional construction activity, population density, and frequency of meteorological disasters, with corresponding quantification values ​​of 1.0-1.8. The higher the level, the larger the quantification value.

[0100] Historical processing efficiency weighting coefficient, with a value range of [0.7, 1.1]. The higher the historical processing efficiency, the smaller the weight.

[0101] Historical processing efficiency quantification value, calculated using the following formula: The value range is [0.8, 1.2].

[0102] 1.3 Benchmark Calibration: The benchmark value is calibrated quarterly based on the latest historical disposal data. Calibration is performed to ensure the timeliness and accuracy of the reference values.

[0103] Step 2: Risk Quantification and Initial Resource Allocation

[0104] A multi-dimensional comprehensive risk scoring model for potential hazards is constructed to quantify the risk level of potential hazards. Combined with historical handling data, the initial resource matching degree is determined to achieve preliminary rational allocation of resources. The specific steps are as follows:

[0105] 2.1 Constructing a multi-dimensional risk assessment index system: Five core assessment dimensions were selected, namely voltage level ( ), construction type ( ), construction time ( Mechanical risks ), time series features ( );

[0106] 2.2 Indicator Quantification: Each dimension is quantified in a hierarchical manner, with a quantification range of [0, 10]. The specific hierarchical standards are as follows:

[0107] voltage level : 500kV=10, 220kV=8, 110kV=6, 35kV=4;

[0108] Construction type Large machinery construction (such as tower cranes and excavators) = 10, medium machinery construction = 7, manual construction = 4, temporary work = 2;

[0109] Construction time :>72h=10, 24-72h=7, 12-24h=4, <12h=2;

[0110] Mechanical risks Distance between machinery and line <5m=10, 5-10m=7, 10-15m=4, >15m=2;

[0111] Time series features Peak load periods (such as summer and winter peak electricity consumption) = 10, weekday daytime = 7, weekday nighttime = 4, holidays = 2.

[0112] 2.3 Comprehensive Risk Scoring Model: The Analytic Hierarchy Process (AHP) is used to determine the weights of each dimension. (satisfy Comprehensive risk score The calculation formula is:

[0113]

[0114] The default values ​​for the weights are: , , , , It can be adjusted according to the actual operation and maintenance scenario.

[0115] 2.4 Initial resource matching degree calculation: combined with comprehensive risk score and frequency of handling similar hidden dangers in the past The initial resource matching degree is calculated using linear fitting. The formula is:

[0116]

[0117] Parameter definition:

[0118] Risk score impact coefficient, with a value range of [0.05, 0.1] and a default value of 0.08;

[0119] Historical frequency influence coefficient, with a value range of [0.1, 0.2] and a default value of 0.15;

[0120] Average monthly frequency of handling similar hazards in the past 12 months (unit: times / month), with a value range of [1, 20].

[0121] Initial resource matching degree, with a value range of [0.5, 1.5]. The larger the size, the more resources are required.

[0122] 2.5 Initial Resource Configuration: Based on Three resource allocation levels (low, medium, and high) are defined, with corresponding allocations of different numbers of maintenance personnel, equipment, and materials to ensure that resources are matched with risks.

[0123] Step 3: Dynamic Time Limit Generation

[0124] Based on the target closed-loop period benchmark value determined in step 1 and the initial resource matching degree determined in step 2 The initial treatment time limit is determined through a nonlinear mapping relationship, and an environmental disturbance correction factor is introduced for dynamic adjustment to ensure the rationality and adaptability of the time limit. The specific implementation is as follows:

[0125] 3.1 Initial treatment time limit calculation: An exponential nonlinear mapping model is adopted to calculate the initial treatment time limit. and Mapped to initial processing time limit The formula is:

[0126]

[0127] Parameter definition: These are nonlinear mapping coefficients, with a value range of [0.3, 0.6] and a default value of 0.45. The larger the value, the more significant the impact of resource matching degree on the time limit.

[0128] 3.2 Construction of Environmental Disturbance Correction Factors: Introducing Environmental Disturbance Correction Factors Taking into account three environmental factors—weather, traffic, and topography—the formula is:

[0129]

[0130] Parameter definition:

[0131] Environmental impact weighting coefficient, with a value range of [0.05, 0.15] and a default value of 0.1;

[0132] Weather influencing factors: Sunny = 0, Cloudy = 0.1, Light rain = 0.2, Heavy rain / blizzard / typhoon = 0.5;

[0133] Traffic impact factors: Smooth traffic = 0, Congested traffic = 0.2, Road interruption = 0.4;

[0134] Topographical influence factor: Plains = 0, Hills = 0.1, Mountainous areas = 0.3.

[0135] 3.3 Determination of Dynamic Handling Time Limits: The initial handling time limit will be determined... Environmental disturbance correction factor Multiply by each product to obtain the final dynamic processing time limit. The formula is:

[0136]

[0137] 3.4 Time Limit Boundary Constraints: Set upper and lower limits for the time limit, with the lower limit being... The upper limit is This ensures that the timeframe is neither too short, leading to insufficient handling, nor too long, leading to increased risks.

[0138] Step 4: Process Prediction and Deviation Identification

[0139] During the hazard mitigation process, the time elapsed for mitigation is monitored in real time. The Kalman filter algorithm is used to predict the remaining mitigation time, calculate the estimated closed-loop cycle, and then identify cycle control deviations. This provides a basis for subsequent time limit optimization and resource allocation. The specific steps are as follows:

[0140] 4.1 Real-time monitoring: The operation and maintenance management system collects the time elapsed for handling potential hazards in real time. The data collection frequency is once every 1 hour.

[0141] 4.2 Kalman Filter Algorithm Construction: Based on Remaining Processing Time Using state variables, we construct the state equation and observation equation for Kalman filtering to achieve accurate prediction of the remaining time.

[0142] Equations of state:

[0143] Observation equation:

[0144] Parameter definition:

[0145] : The state vector at time k, here (Remaining processing time at time k);

[0146] : State transition matrix, here a 1×1 matrix, (Considering efficiency fluctuations during the disposal process);

[0147] : Control input at time k, which is the processing efficiency adjustment amount, with a value range of [-0.5, 0.5];

[0148] : Control matrix, here a 1×1 matrix, ;

[0149] Process noise follows a normal distribution. , This represents the process noise covariance, with a default value of 0.01.

[0150] : The observation at time k, i.e. the estimated remaining processing time at time k;

[0151] : Observation matrix, here a 1×1 matrix, ;

[0152] Observational noise follows a normal distribution. , To observe the noise covariance, the default value is 0.02.

[0153] 4.3 Kalman Filter Iterative Process:

[0154] Prediction step: (State prediction); (Covariance prediction);

[0155] Update steps: (Kalman gain); (Status update); (Covariance update), where It is an identity matrix.

[0156] 4.4 Predicting the closed-loop period and identifying deviations:

[0157] Estimated closed-loop cycle The calculation formula is:

[0158] Periodic control deviation The calculation formula is:

[0159] Deviation judgment criteria: To anticipate deviations, This is within the normal deviation range. This is due to timeout deviation.

[0160] Step 5: Time Limit Rolling Optimization

[0161] Construct a fuzzy adaptive control mechanism based on the periodic control deviation identified in step 4. and its changing trends The subsequent processing time limit is dynamically adjusted to ensure that the closed-loop cycle remains stable within the target range. This is specifically achieved as follows:

[0162] 5.1 Fuzzy Control Input / Output Definitions:

[0163] Input 1: Periodic control deviation The fuzzy linguistic variables are "negative large (NB), negative small (NS), zero (ZO), positive small (PS), positive large (PB)", and the quantization range is [-6, 6].

[0164] Input Quantity 2: Deviation Change Trend (Unit: h / h), which is the ratio of the difference between two adjacent deviations to the time interval. The fuzzy linguistic variables are "negative large (NB), negative small (NS), zero (ZO), positive small (PS), positive large (PB)", and the quantization range is [-3,3].

[0165] Output: Time Limit Adjustment (Unit: h), the fuzzy linguistic variables are “negative large (NB), negative small (NS), zero (ZO), positive small (PS), positive large (PB)”, and the quantization range is [-10, 10].

[0166] 5.2 Fuzzy Rule Formulation: Based on operational experience, a 5×5 fuzzy control rule table was formulated (some rules are shown below):

[0167] like For NB, If it is NB, then For NB (significantly shortened time limit);

[0168] like For ZO, is ZO, then Set to ZO (no time limit adjustment);

[0169] like For PB, If it is PB, then For PB (significantly extended time limit).

[0170] 5.3 Fuzzy Inference and Defuzzification: The Mamdani fuzzy inference method is adopted, combined with triangular membership functions for fuzzy inference, and defuzzification is performed using the centroid method to obtain the specific time limit adjustment amount. The formula is:

[0171]

[0172] in, Let be the membership degree of the i-th fuzzy set. Let n be the center value of the i-th fuzzy set, and n be the number of fuzzy sets.

[0173] 5.4 Determination of time limits after rolling optimization: New processing time limits The calculation formula is: At the same time, the time limit boundary constraints of step 3.4 are satisfied, and the rolling optimization is performed once every 2 hours.

[0174] Step 6: Optimize and allocate resources

[0175] A multi-objective resource optimization and allocation model is established, with the optimization objectives being disposal cost, risk reduction rate, and cycle achievement rate. An evolutionary algorithm is used to solve for the optimal resource allocation scheme, and a tiered resource allocation strategy is implemented in conjunction with the cycle control deviation range. The specific steps are as follows:

[0176] 6.1 Multi-objective optimization objective function:

[0177] Construct three objective functions, and use a weighted sum method to transform the multi-objective optimization into a single objective function. The overall objective function is... for:

[0178]

[0179] Definitions of objective functions and parameters:

[0180] disposal costs The calculation formula is: ,in The average operation and maintenance cost per person (RMB / hour) For the number of maintenance personnel, The cost of equipment usage (RMB / hour) For the number of devices, Cost of materials consumed (yuan). This refers to the amount of materials consumed.

[0181] Risk reduction rate The calculation formula is: ,in Initial risk assessment for potential hazards. The current risk score ranges from [0, 1].

[0182] Cycle achievement rate The calculation formula is: The value range is [0, 1];

[0183] , , Target weight, satisfying Default value: , , .

[0184] 6.2 Constraints:

[0185] Total resource constraints: , , ,in , , These represent the maximum number of available maintenance personnel, equipment, and materials, respectively.

[0186] Constraints on handling capacity: ,in per capita processing efficiency For the processing efficiency of a single device. The minimum efficiency required for handling potential hazards;

[0187] Nonnegativity constraint: , , And all are integers.

[0188] 6.3 Genetic Algorithm Solution: The optimal resource allocation scheme is solved using a genetic algorithm. The steps are as follows:

[0189] Initialize the population: The population size is 50, the chromosome encoding uses real number encoding, and each chromosome corresponds to a set of resource allocation schemes (N, M, K).

[0190] Fitness function: (Objective function) The smaller the size, the higher the adaptability.

[0191] Genetic operations: The selection operator uses roulette wheel selection, the crossover operator uses single-point crossover (crossover probability 0.8), and the mutation operator uses Gaussian mutation (mutation probability 0.05).

[0192] Iteration Termination: When the number of iterations reaches 100 or the fitness tends to stabilize, the iteration terminates and the optimal resource allocation scheme is output.

[0193] 6.4 Tiered resource allocation strategy: Controlling deviations based on cycles Different resource allocation strategies are implemented within the specified intervals:

[0194] Lead time deviation ( ): Appropriately reduce resource allocation (reduce personnel or equipment by 10%-20%) to lower disposal costs;

[0195] Normal deviation ( ): Maintain the optimal resource allocation scheme unchanged;

[0196] Timeout deviation ( Increase resource allocation (increase personnel or equipment by 20%-30%) to speed up the processing and reduce the risk of delays.

[0197] Step 7: Quality Acceptance and Evaluation

[0198] A multi-dimensional closed-loop quality scoring system is established to comprehensively evaluate the effectiveness of hazard mitigation, ensuring that the quality of mitigation meets standards. If the quality is not up to standard, the hazard is returned for re-treatment. The specific steps are as follows:

[0199] 7.1 Multi-dimensional quality evaluation index system: Four core evaluation dimensions are selected, namely, compliance of handling ( ), degree of elimination of hidden dangers ( ), Equipment status recovery ( Document integrity Each dimension is scored out of 100.

[0200] 7.2 Scoring criteria for each dimension:

[0201] Compliance of handling Strict adherence to operation and maintenance procedures earns 80-100 points; minor violations earn 60-79 points; and serious violations earn 0-59 points.

[0202] Degree of hazard elimination Complete elimination of hazards earns 80-100 points; partial elimination of hazards (remaining risk score ≤ 3) earns 60-79 points; and no elimination of hazards (remaining risk score > 3) earns 0-59 points.

[0203] Equipment status recovery : 80-100 points for equipment fully restored to normal, 60-79 points for equipment basically restored to normal (with minor abnormalities), and 0-59 points for equipment unable to operate normally;

[0204] Document integrity Complete and standardized documents such as handling records and test reports will receive 80-100 points; documents that are basically complete but have a few non-standard aspects will receive 60-79 points; and documents that are missing or seriously non-standard will receive 0-59 points.

[0205] 7.3 Calculation of Overall Quality Score: A weighted summation model is used to calculate the overall quality score. The calculation formula is:

[0206]

[0207] Parameter definition: , , , ,satisfy .

[0208] 7.4 Acceptance Criteria: Set Acceptance Thresholds points, if If the acceptance is deemed satisfactory, the closed loop ends; If the acceptance is deemed unqualified, return to step 6 for resource reallocation and disposal until acceptance is qualified.

[0209] Step 8: Parameter Optimization

[0210] Based on the Deep Reinforcement Learning (DRL) framework, with the goal of minimizing the closed-loop cycle deviation, the risk score weight, resource matching impact coefficient, time limit correction coefficient, and target closed-loop cycle benchmark value are iteratively optimized online to continuously improve the accuracy of closed-loop cycle control. The specific implementation is as follows:

[0211] 8.1 Construction of Deep Reinforcement Learning Framework: The DQN (Deep Q-Network) algorithm is used to construct an "agent-environment" interaction model, in which:

[0212] Environment: Scenario for handling potential hazards of external damage to power transmission lines; state space. Including periodic control deviation Deviation change trend Comprehensive risk score Initial resource matching degree ;

[0213] Intelligent agent: Parameter optimization subject, action space Including adjusting the risk score weights Resource matching impact coefficient Time Limit Correction Factor Target closed-loop cycle benchmark value The adjustment range;

[0214] Reward function: The reward function aims to minimize the closed-loop periodic deviation. The calculation formula is: ,in This is the cost penalty coefficient, with a value range of [0.01, 0.05] and a default value of 0.03. The larger the value, the better the parameter adjustment effect.

[0215] 8.2 Network Structure Design:

[0216] Input layer: 4-dimensional (corresponding to 4 state variables in the state space);

[0217] Hidden layers: 2 fully connected layers, with 32 neurons in the first layer and 16 neurons in the second layer, using ReLU as the activation function;

[0218] Output layer: The dimension is the size of the action space, and it outputs the Q value (action value) of each action.

[0219] 8.3 Online Iterative Optimization Process:

[0220] Initialization: Initialize DQN network parameters, experience replay pool (capacity 10000), and target network parameters (consistent with DQN network parameters);

[0221] Interactive sampling: The agent samples the state of the environment. Select Action (Adjust parameters) After executing the action, a new state is obtained. and rewards , experience Store in the experience replay pool;

[0222] Network update: When the number of experiences in the experience replay pool reaches the threshold (1000), a batch of experiences (batch size 32) is randomly sampled, and the target Q value is calculated. ,in The discount factor is set to 0.95, and the DQN network parameters are updated using gradient descent.

[0223] Target network update: Every 100 iterations, the DQN network parameters are copied to the target network to ensure the stability of the target Q value;

[0224] Termination condition: When the average value of the closed-loop periodic deviation... If the iteration continues for 100 closed-loop cycles, the iteration is terminated and the optimized parameter values ​​are output; otherwise, the iteration optimization continues.

[0225] 8.4 Parameter Update: Apply the optimized parameters to the closed-loop management of subsequent hazard handling. Repeat steps 8.1-8.3 every quarter to achieve continuous parameter optimization.

[0226] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A closed-loop cycle management method for reducing the potential for external force damage to transmission lines from detection to handling, characterized in that: include: Cycle benchmark setting: Based on line voltage level, regional risk characteristics and historical handling data, establish differentiated target closed-loop cycle benchmark values; Risk quantification and initial resource allocation: Construct a multi-dimensional comprehensive risk scoring model for hidden dangers, integrate voltage level, construction type, construction duration, mechanical risk and time series characteristics to calculate a comprehensive risk score, and determine the initial resource matching degree based on the score and the frequency of handling similar hidden dangers in the past. Dynamic time limit generation: Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial processing time limit is determined through a nonlinear mapping relationship, and an environmental disturbance correction factor is introduced for dynamic adjustment; Process prediction and deviation identification: During the handling of potential hazards, the time elapsed is monitored in real time, and the Kalman filter algorithm is used to predict the remaining handling time to calculate the estimated closed-loop cycle, thereby identifying cycle control deviations; Time limit rolling optimization: Construct a fuzzy adaptive control mechanism to dynamically adjust the subsequent processing time limit based on the periodic control deviation and its changing trend; Resource optimization and allocation: Establish a multi-objective resource optimization and allocation model with the optimization objectives of disposal cost, risk reduction rate and cycle achievement rate. Solve the optimal resource allocation scheme through genetic algorithm, and implement a hierarchical resource allocation strategy based on the cycle control deviation range. Quality Acceptance Evaluation: Construct a multi-dimensional closed-loop quality scoring system to comprehensively evaluate the effectiveness of hazard handling. When the score reaches the qualified threshold, the acceptance is deemed qualified; otherwise, it is returned for re-handling. Parameter optimization: Based on a deep reinforcement learning framework, with the goal of minimizing the closed-loop cycle deviation, the risk score weight, resource matching influence coefficient, time limit correction coefficient, and target closed-loop cycle benchmark value are optimized online iteratively.

2. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The periodic reference setting specifically includes: First, three core influencing factors are selected: voltage level, regional risk level, and historical handling efficiency. The quantitative standards for each factor are then defined. Next, based on the basic closed-loop cycle, the weight coefficients and quantitative values ​​of each factor are combined to calculate the differentiated target closed-loop cycle benchmark value through weighted summation. Finally, the target closed-loop cycle benchmark value is calibrated quarterly using the latest historical handling data to ensure that the benchmark value is timely and accurate.

3. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The aforementioned risk quantification and initial resource allocation specifically include: A risk assessment index system was constructed, comprising five core dimensions: voltage level, construction type, construction duration, mechanical risk, and time series characteristics. The grading and quantification standards for each dimension were also clarified. A multi-dimensional comprehensive risk scoring model for hidden dangers was then constructed. The weights of each dimension were determined by the analytic hierarchy process, and a weighted summation method was used to calculate the comprehensive risk score for hidden dangers. Subsequently, by combining the comprehensive risk score with the average monthly handling frequency of similar hazards over the past 12 months, the initial resource matching degree was calculated through linear fitting. Finally, based on the initial resource matching degree, three resource configuration levels—low, medium, and high—we allocated corresponding maintenance personnel, equipment, and materials to ensure that resource configuration matches the hazard risk level.

4. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The dynamic time limit generation specifically includes: Based on the target closed-loop cycle benchmark value and the initial resource matching degree, the initial response time limit is calculated using an exponential nonlinear mapping relationship. Subsequently, an environmental disturbance correction factor is introduced, which integrates three types of environmental factors: weather, traffic, and terrain. By quantifying each environmental factor and setting environmental impact weight coefficients, a calculation formula for the environmental disturbance correction factor is constructed. The initial response time limit is then multiplied by the environmental disturbance correction factor to obtain the final dynamic response time limit. Finally, upper and lower limits are set to constrain the time limit range and avoid insufficient response due to excessively short time limits or increased risk due to excessively long time limits.

5. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The process prediction and deviation identification specifically include: During the hazard handling process, the used handling time is collected in real time at a fixed frequency. The remaining handling time is used as the state variable to construct the Kalman filter state equation and observation equation. The remaining handling time is predicted through a two-step iterative operation of prediction and update. The predicted remaining time is added to the used time to obtain the estimated closed-loop cycle. The cycle control deviation is obtained by calculating the difference between the estimated cycle and the dynamic handling time limit. Based on the preset deviation threshold range, the cycle control deviation is divided into three categories: advance deviation, normal deviation, and overtime deviation, thereby realizing real-time deviation identification of the handling cycle.

6. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The time-limited rolling optimization specifically includes: First, the periodic control deviation and its changing trend are defined as fuzzy control inputs, and the time limit adjustment is defined as the output. The fuzzy linguistic variables and quantization ranges of each quantity are then clarified. Next, fuzzy control rules are formulated based on operational experience. The Mamdani fuzzy inference method is used in conjunction with triangular membership functions for inference. The specific time limit adjustment is obtained by defuzzification using the centroid method. Finally, a new handling time limit is calculated based on the time limit adjustment to ensure that it meets the preset time limit boundary constraints. Rolling optimization is performed every 2 hours to dynamically adjust the subsequent handling time limit.

7. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The aforementioned resource optimization and allocation specifically includes: A multi-objective optimization model was constructed, with disposal cost, risk reduction rate, and cycle achievement rate as optimization objectives. This model was transformed into a single-objective function using a weighted sum method, and three types of constraints were set: total resource volume, disposal capacity, and non-negativity. Subsequently, a genetic algorithm was used to solve the optimization model. By initializing the population, setting the fitness function, performing genetic operations, and iteratively terminating, the optimal resource allocation scheme was obtained. Finally, based on the cycle control deviation range, a tiered resource allocation strategy was implemented: early deviations reduced resource allocation to lower costs, normal deviations maintained optimal allocation, and overdue deviations increased resource allocation to accelerate progress, ensuring precise matching of resource allocation with disposal needs and control objectives.

8. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The quality acceptance evaluation specifically includes: First, a multi-dimensional quality evaluation system is constructed, which includes four core dimensions: compliance of handling, degree of elimination of hidden dangers, restoration of equipment status, and integrity of documentation. The scoring criteria for each dimension are clearly defined. Then, a weighted summation method is used to calculate the comprehensive quality score by combining the preset weights of each dimension. Finally, a pass threshold is set. If the score is greater than the pass threshold, the acceptance is deemed qualified and the handling of hidden dangers is completed. If the score is less than the passing threshold, the acceptance is deemed unqualified, and the process is repeated until the acceptance is qualified, ensuring that the hazard handling effect meets the requirements.

9. The closed-loop cycle management method for reducing the potential for external force damage to transmission lines from discovery to handling, as described in claim 1, is characterized in that: The parameter optimization specifically includes: A deep reinforcement learning model (DQN) is used to construct an interaction model between the agent and the hazard management environment. A reward function is designed with minimizing the closed-loop cycle deviation as the optimization objective, and a corresponding neural network structure with input, hidden, and output layers is built. Samples are sampled through agent-environment interaction and stored in an experience replay pool. After meeting the sample size threshold, gradient descent is used to update the network parameters. The target network is periodically synchronized to ensure training stability. The iteration terminates when the mean closed-loop cycle deviation reaches the target and remains stable. The optimized risk score weights, resource matching influence coefficients, time limit correction coefficients, and the target closed-loop cycle baseline value are output. This iterative optimization process is repeated quarterly, applying the optimized parameters to subsequent management processes to continuously improve the management accuracy of the closed-loop cycle.

10. A closed-loop cycle management system for reducing the potential for external force damage to transmission lines from detection to handling, characterized in that: include: The data acquisition module is used to collect data on transmission line voltage levels, regional risk characteristics, historical hazard handling data, on-site construction information of hazards, environmental parameters and handling process data, providing data support for each module; The cycle benchmark setting module, connected to the data acquisition module, is used to select three core factors: voltage level, regional risk level, and historical handling efficiency, clarify quantitative standards, calculate differentiated target closed-loop cycle benchmark values ​​through weighted summation, and calibrate them periodically. The risk quantification and initial resource allocation module, connected to the data acquisition module and the periodic benchmark setting module, is used to construct a multi-dimensional risk assessment indicator system, calculate the comprehensive risk score of hidden dangers, determine the initial resource matching degree by combining the historical handling frequency, and complete the initial resource allocation. The dynamic time limit generation module is connected to the cycle benchmark setting module and the risk quantification and resource initial allocation module. It is used to calculate the initial treatment time limit through exponential nonlinear mapping, introduce environmental disturbance correction factors for dynamic adjustment, and set time limit boundary constraints. The process prediction and deviation identification module is connected to the data acquisition module and the dynamic time limit generation module. It is used to monitor the time elapsed in real time, predict the remaining time through the Kalman filter algorithm, calculate the estimated closed-loop cycle and identify the cycle control deviation. The rolling optimization module is connected to the process prediction and deviation identification module to build a fuzzy adaptive control mechanism. Based on the periodic control deviation and its changing trend, the time limit adjustment amount is obtained through fuzzy inference and defuzzification to realize the rolling optimization of the disposal time limit. The resource optimization and allocation module is connected to the process prediction and deviation identification module and the risk quantification and initial resource allocation module. It is used to build a multi-objective resource optimization model, solve the optimal allocation scheme through a genetic algorithm, and perform hierarchical resource allocation based on the deviation interval. The quality acceptance and evaluation module is connected to the resource optimization and allocation module. It is used to build a multi-dimensional quality evaluation system, calculate the comprehensive quality score, determine the acceptance result based on the pass threshold, and trigger the reprocessing process if it fails to pass. The parameter optimization module is connected to the aforementioned modules. It uses DQN deep reinforcement learning to build an interactive model, iterates and optimizes various control parameters online, and updates and applies them to the control process every quarter. The data storage module is used to store all data collected, calculated, and optimized by each module, ensuring data traceability and providing support for subsequent management and parameter optimization.