Lightning arrester risk level dynamic adjustment control method based on reinforcement learning

By combining reinforcement learning methods with multi-source state perception and an improved FEDformer model, the problem of multi-factor coupling in surge arrester risk assessment is solved, enabling dynamic adjustment and forward-looking prediction of surge arrester risk levels, thereby improving the operational safety and intelligence level of surge arresters.

CN121765646APending Publication Date: 2026-03-31HEILONGJIANG ELECTRIC POWER SCIENCE RESEARCH INSTITUTE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing methods for monitoring and assessing the operation of surge arresters rely on single or limited indicators, which are insufficient to reflect the nonlinear and time-varying characteristics under the combined effects of multiple factors. They lack forward-looking predictive capabilities, leading to delayed or misjudged risk level adjustments, and the adjustment strategies lack systematic optimization.

Method used

A reinforcement learning-based approach is adopted to predict risk potential energy through multi-source state perception, phase-coupled risk modeling, and an improved FEDformer model. Combined with a time-return look-ahead safety barrier and a three-layer reinforcement learning adjustment strategy, the risk level of the surge arrester is dynamically adjusted.

Benefits of technology

It improves the accuracy and foresight of surge arrester risk assessment, adaptively optimizes adjustment strategies, ensures control safety and timely response, and is suitable for long-term safety management in complex operating environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765646A_ABST
    Figure CN121765646A_ABST
Patent Text Reader

Abstract

The invention discloses a lightning arrester risk level dynamic adjustment control method based on reinforcement learning, and the method comprises the steps: collecting multi-source operation state data of a lightning arrester, and carrying out the preprocessing of the data, and generating a standard data set; analyzing phase characteristics in a multi-band frequency domain, embedding a manifold space, and calculating risk potential energy; the risk potential energy and the manifold are used as input, and the FEDform is improved to predict the multi-step risk potential energy; constructing a safety barrier based on a prediction result, and performing reverse mapping to generate a safety action constraint; executing a three-layer reinforcement learning strategy under constraint, and generating candidate action output; and fusing the verification candidate actions, generating an adjustment instruction, and issuing the adjustment instruction to the control device. According to the invention, by fusing the phase coupling risk modeling, the prospective risk prediction and the reinforcement learning safety decision mechanism, the prospective, dynamic and safe controllable adjustment of the risk level of the lightning arrester is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system equipment operation control technology, and in particular to a method for dynamic adjustment and control of surge arrester risk level based on reinforcement learning. Background Technology

[0002] Surge arresters, as crucial overvoltage protection devices in power systems, are widely used in transmission lines, substations, and distribution networks. Their operational status directly impacts the safety and reliability of the power system. Current technologies typically rely on single or limited indicators such as leakage current, temperature, or lightning strike counts for operational monitoring and risk assessment of surge arresters. The arrester's condition is judged by setting fixed thresholds or empirical rules, and alarm or maintenance measures are taken accordingly. Traditional methods are simple in structure and easy to implement in engineering, and were widely used in the early operation and maintenance of power equipment.

[0003] However, with the increasingly complex operating environment of power systems, surge arresters exhibit significant nonlinear and time-varying risk evolution under the combined effects of lightning strikes, electrical harmonics, ambient temperature variations, and long-term aging. Existing methods based on static thresholds or single-frequency analysis struggle to reflect the phase relationships and coupling effects across different frequency bands, failing to accurately characterize the comprehensive changes in the internal electrical-thermal state of surge arresters. Traditional methods often employ post-hoc judgment or short-term trend analysis, lacking the ability to proactively predict future operational risks, which can easily lead to delayed or misjudged risk levels.

[0004] Existing surge arrester regulation strategies are mostly driven by manual experience or simple logic control, lacking a systematic optimization decision-making mechanism, making it difficult to achieve dynamic adjustment of risk levels while ensuring safety. Even with the introduction of some intelligent algorithms, related technologies mostly focus on state recognition or lifespan assessment, failing to form an effective closed loop with actual control actions, and even less able to constrain the safety of control actions over multiple future operating cycles. How to achieve accurate modeling, forward prediction, and safe and controllable dynamic adjustment of surge arrester risks based on multi-source information fusion remains a key problem that needs to be solved in existing technologies.

[0005] Therefore, how to provide a method for dynamic adjustment and control of the risk level of surge arresters based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a dynamic adjustment and control method for the risk level of surge arresters based on reinforcement learning. This invention achieves intelligent dynamic adjustment of the risk level of surge arresters by integrating multi-source operational state perception, phase-coupled risk modeling, forward-looking risk prediction, and safety constraint control. This invention utilizes a phase-coupled risk manifold to perform refined modeling of the operational risk of surge arresters under the coupled effects of multiple factors such as lightning strikes, electrical harmonics, temperature rise, and aging. It combines an improved FEDformer model to perform multi-step prediction of the future evolution trend of risk potential energy. Through a time-retrograde forward-looking safety barrier, future risk constraints are mapped to the current control action space. Based on this, a three-layer reinforcement learning adjustment strategy is employed to generate safe and feasible adjustment commands, thereby completing the dynamic control of the risk level. This effectively improves the accuracy and forward-looking nature of surge arrester risk assessment, achieving adaptive optimization of the adjustment strategy while ensuring control safety. It possesses advantages such as accurate risk identification, timely adjustment response, and strong safety controllability.

[0007] The reinforcement learning-based dynamic adjustment and control method for surge arrester risk level according to embodiments of the present invention includes: Collect multi-source operating status data during the operation of the surge arrester, preprocess the multi-source operating status data, and generate a standard dataset; Multi-band frequency domain analysis is performed on the standard dataset to extract phase features, a set of phase coupling features is constructed and embedded into the risk manifold space to form a phase coupling risk manifold and calculate the risk potential energy index. Using risk potential energy index and phase-coupled risk manifold as inputs, an improved FEDformer model is constructed to predict the evolution trend of surge arrester risk potential energy in multiple future control cycles, and obtain a multi-step risk potential energy prediction sequence. Based on the multi-step risk potential energy prediction sequence, a time-return forward-looking safety barrier is constructed, and the risk potential energy change constraints in multiple future control cycles are mapped inversely to the current control action space to form safe and feasible action constraints at the current moment. Under the constraint of safe and feasible actions, a three-layer reinforcement learning regulation strategy is constructed, including a topological potential energy guidance layer, a return constraint layer, and a dual-track strategy fusion layer using a combination of main track and auxiliary track, to generate candidate risk level actions and candidate continuous parameter actions. The candidate risk level actions and candidate continuous parameter actions are fused and safety verified to generate a risk level dynamic adjustment command and send it to the surge arrester control device. Based on the execution feedback, the phase coupling risk manifold is updated, the FEDformer model is improved, and the reinforcement learning adjustment strategy is applied.

[0008] Optionally, the multi-source operating status data includes surge arrester leakage current signal, multi-band frequency domain components of leakage current, lightning pulse signal, lightning strike count information, temperature data of the surge arrester body or interior, temperature rise rate data, operating parameters reflecting the surge arrester's operating time and aging degree, and environmental parameters characterizing the external operating environment.

[0009] Optionally, the preprocessing of multi-source operating status data includes time alignment, outlier detection and removal, missing data compensation, and amplitude normalization or standardization of the collected multi-source operating status data.

[0010] Optionally, forming the phase-coupled risk manifold and calculating the risk potential energy index includes: A fixed set of multiple frequency bands is set in the standard dataset. The phase of each frequency band is extracted from the leakage current and lightning pulse signal, and time alignment and phase expansion are performed to obtain a continuous and consistent phase sequence. Based on the phase sequence and frequency band pairing relationship, the phase relationship between each frequency band, the short-time stability of each frequency band phase, and the relative hysteresis with the temperature signal phase are calculated, and the phase coupling feature set is constructed by combining the aging parameters in a fixed field order. A phase coupling graph is constructed using a set of phase coupling features. The nodes of the phase coupling graph correspond to each frequency band and temperature. The weights of the edges are jointly determined by the phase relationship quantity and the short-time stability quantity. An aging parameter is introduced to uniformly adjust the weights of all edges. Based on the phase coupling graph, a deterministic embedding is performed while maintaining the consistency of the edge weight order and the main direction to generate a manifold coordinate vector. The manifold coordinate vector and the corresponding phase coupling graph weights are incorporated to form a phase coupling risk manifold, which is jointly defined by frequency domain phase relationship and dual-domain consistency constraints of temperature and aging context. Risk potential energy index is calculated on the phase-coupled risk manifold, which is obtained by weighted aggregation of cross-band phase inconsistency, temperature phase lag, aging level and phase change rate penalty.

[0011] Optionally, obtaining the multi-step risk potential prediction sequence includes: Set a fixed historical window length and prediction step size, and combine the embedding vectors of the risk potential index and the phase-coupled risk manifold into time-series input samples in chronological order, keeping the input dimension consistent. An improved FEDformer model is constructed, which sequentially includes an input embedding layer, a dual-frequency domain cross-attention module, a phase consistency gating module, a residual frequency band aggregation module, and a prediction output head, wherein: The dual-frequency domain cross-attention module separates the input embedding vector into low-frequency and high-frequency channels, calculates the attention weights of each channel, performs a cross-product operation, and reassembles the two cross-features to output a multi-frequency coupled feature sequence. The phase consistency gating module receives multi-band coupled feature sequences and phase difference sequences at the same time step, performs short-time stability assessment on the phase difference, uses the assessment value as a gating factor to multiply into the corresponding feature channel, suppresses channels whose instantaneous phase jitter exceeds a preset threshold, strengthens channels whose phase consistency is within the stable threshold range, and outputs the feature sequence. The residual frequency band aggregation module superimposes the current feature sequence with the historical frequency band features of each layer through residual connection, fuses them according to preset weights to maintain consistent output dimensions, performs weighted fusion on the superposition result according to preset frequency band weights, and outputs the aggregated feature sequence to the prediction output head. In the improved FEDformer model, the dual-frequency domain cross-attention module and the phase consistency gating module are connected in series, and the residual frequency band aggregation module is located at the last stage and is connected in parallel with the prediction output head. Set the training batch size, learning rate, window length and regularization coefficient, and use a joint loss function composed of risk potential error term and short-term trend smoothing term to supervise the training of the improved FEDformer model. After each training cycle, execute an early stopping strategy to avoid overfitting. During the inference phase, the FEDformer model is improved by using online input time series samples of the same dimension, and the multi-step risk potential prediction sequence is obtained through the prediction output head transformation.

[0012] Optionally, the constraints that form the safe and feasible actions at the current moment include: Obtain the current risk potential energy index and multi-step risk potential energy prediction sequence, form a prediction sequence covering multiple future control cycles in chronological order, and record the corresponding time index; Set a return coefficient and step weight associated with the current risk potential for each future control cycle, calculate the upper limit of the risk potential allowed in the cycle, generate a return judgment list arranged in chronological order, and simultaneously establish a frequency domain allocation table to record the incremental contribution of each frequency band in the multi-step risk potential prediction sequence as an independent entry. Based on the comparison of predicted values ​​with corresponding upper limits item by item in the turnaround judgment list, the future multi-step safety judgment results are generated. When an over-limit period occurs, the step-by-step risk budget recovery mechanism is triggered, and the over-limit amount is recovered from the frequency domain allocation table according to the preset priority, forming a turnaround correction instruction. The turnaround decision list, future multi-step safety decision results, and turnaround correction instructions are combined to generate a time-turnaround forward-looking safety barrier, which is given in the form of executable constraints, including: the allowable set and priority of the main track risk level, the value range and restriction direction of the auxiliary track continuous parameters, and the joint feasible combination relationship between the main track and the auxiliary track. The time-reversal forward-looking safety barrier is mapped inversely to the current control action space, forming the safe and feasible action constraints at the current moment.

[0013] Optionally, the action of generating candidate risk levels and candidate continuous parameters includes: At the current moment, obtain the state input consisting of the risk potential energy index and the phase-coupled risk manifold embedding vector, receive the safe and feasible action constraints, and set the discrete value set of the main track risk level action and the value range of the auxiliary track continuous parameter action. The topological potential energy guiding layer generates a guiding direction along the direction of risk potential energy decrease on the phase-coupled risk manifold. It samples the guiding direction according to a preset step size and number of samplings to generate a set of guiding candidate actions. After curvature adaptive step size control and direction calibration, it retains the candidate actions that are consistent with the guiding direction and meet the upper limit of the step size. The backtracking constraint layer will guide the candidate action set to be matched and screened with the safe and feasible action constraints, eliminating candidates that are not in the allowed set and value range. The remaining candidates will be processed in sequence according to the backtracking priority, adjusting the discrete level of the main track and the continuous parameter of the auxiliary track in turn, and completing the constraint satisfaction verification in the order of processing the main track first and then the auxiliary track until the constraint conditions are met or the candidate is confirmed to be unusable. The dual-track strategy fusion layer performs dual-track consistency verification on the retained candidates, matches the main track candidate risk level actions with the auxiliary track candidate continuous parameter actions according to the mapping table, eliminates interlocking and inconsistent combinations, sorts the combinations that pass the verification according to the preset score, and outputs an ordered set of candidate risk level actions and candidate continuous parameter action combinations.

[0014] Optionally, the step of generating a dynamic risk level adjustment command and sending it to the surge arrester control device includes: Receive candidate risk level actions and candidate continuous parameter actions, and construct a set of candidate control action pairs according to a one-to-one correspondence. Based on the safety and feasible action constraints, the set of candidate control action pairs is checked one by one for safety, and control action pairs that do not meet the safety and feasible action constraints are eliminated to obtain a set of safe candidate control action pairs. In the set of candidate control action pairs for safety, each control action pair is comprehensively evaluated based on the trend of risk potential energy change, the adjustment range of control action, and historical execution stability indicators. The control action pair with the best evaluation result is selected as the target control action pair. The target control action is converted into a risk level dynamic adjustment command and sent to the surge arrester control device for execution. At the same time, the operation feedback data after execution is collected. The operation feedback data includes the updated risk potential energy index and the corresponding operation status information. Based on the operational feedback data, the embedding parameters of the phase-coupled risk manifold, the prediction parameters of the improved FEDformer model, and the policy parameters of the reinforcement learning adjustment strategy are updated to form the state input for the next control cycle.

[0015] The beneficial effects of this invention are: This invention introduces a phase-coupled risk manifold modeling approach, mapping multi-source information such as multi-band electrical signals, temperature changes, and aging conditions of surge arresters into a continuous and measurable risk space, thus achieving a refined description of the operational risks of surge arresters. Compared to traditional assessment methods that rely on single indicators or fixed thresholds, this invention can comprehensively characterize the risk evolution features under the coupled effects of multiple factors, effectively reducing risk identification errors caused by fragmented information or single indicators, and improving the accuracy and stability of surge arrester risk assessment.

[0016] This invention combines an improved FEDformer model to perform multi-step forward prediction of risk potential energy. By using a time-retrograde forward safety barrier, it maps the risk change constraints in multiple future control cycles back to the current control decision-making process. This makes risk level adjustment no longer limited to ex-post response, but has the ability to predictively drive advance regulation. It can effectively avoid unreasonable adjustments caused by short-term fluctuations or prediction deviations, and ensure that the adjustment actions always meet safety requirements in future operating cycles, thereby improving the reliability and robustness of the overall control strategy.

[0017] This invention employs a three-layer reinforcement learning dynamic adjustment strategy. Through the synergistic effect of risk potential guidance, forward-looking safety constraints, and dual-track action fusion, it achieves adaptive optimization of risk level and continuous control parameters. By continuously using operational feedback for strategy updates, this invention can continuously optimize the adjustment effect during long-term operation, reducing the need for manual intervention and improving the level of intelligent operation and maintenance. Overall, this invention has significant beneficial effects in terms of risk identification accuracy, adjustment response timeliness, and control safety, and is suitable for the long-term safe operation and management of surge arresters in complex operating environments. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows the dynamic adjustment and control method for the risk level of surge arresters based on reinforcement learning proposed in this invention. Figure 2This is a schematic diagram of the multi-step risk potential energy prediction structure based on the improved FEDformer model for the dynamic adjustment and control method of surge arrester risk level based on reinforcement learning proposed in this invention. Figure 3 This is a schematic diagram illustrating the structure and execution flow of the three-layer reinforcement learning adjustment strategy for the dynamic adjustment and control method for the risk level of surge arresters based on reinforcement learning proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1 , Figure 2 and Figure 3 A reinforcement learning-based dynamic adjustment and control method for surge arrester risk levels includes: Collect multi-source operating status data during the operation of the surge arrester, preprocess the multi-source operating status data, and generate a standard dataset; Multi-band frequency domain analysis is performed on the standard dataset to extract phase features, a set of phase coupling features is constructed and embedded into the risk manifold space to form a phase coupling risk manifold and calculate the risk potential energy index. Using risk potential energy index and phase-coupled risk manifold as inputs, an improved FEDformer model is constructed to predict the evolution trend of surge arrester risk potential energy in multiple future control cycles, and obtain a multi-step risk potential energy prediction sequence. Based on the multi-step risk potential energy prediction sequence, a time-return forward-looking safety barrier is constructed, and the risk potential energy change constraints in multiple future control cycles are mapped inversely to the current control action space to form safe and feasible action constraints at the current moment. Under the constraint of safe and feasible actions, a three-layer reinforcement learning regulation strategy is constructed, including a topological potential energy guidance layer, a return constraint layer, and a dual-track strategy fusion layer using a combination of main track and auxiliary track, to generate candidate risk level actions and candidate continuous parameter actions. The candidate risk level actions and candidate continuous parameter actions are fused and safety verified to generate a risk level dynamic adjustment command and send it to the surge arrester control device. Based on the execution feedback, the phase coupling risk manifold is updated, the FEDformer model is improved, and the reinforcement learning adjustment strategy is applied.

[0021] In this embodiment, the multi-source operating status data includes surge arrester leakage current signal, multi-band frequency domain components of leakage current, lightning pulse signal, lightning strike count information, temperature data of the surge arrester body or interior, temperature rise rate data, operating parameters reflecting the surge arrester's operating time and aging degree, and environmental parameters characterizing the external operating environment.

[0022] In this embodiment, the preprocessing of multi-source operating status data includes time alignment, outlier detection and removal, missing data compensation, and amplitude normalization or standardization of the collected multi-source operating status data.

[0023] In this embodiment, the step of forming a phase-coupled risk manifold and calculating the risk potential energy index includes: A fixed set of multiple frequency bands is set in the standard dataset. The phase of each frequency band is extracted from the leakage current and lightning pulse signal, and time alignment and phase expansion are performed to obtain a continuous and consistent phase sequence. Based on the phase sequence and frequency band pairing relationship, the phase relationship between each frequency band, the short-time stability of each frequency band's phase, and the relative hysteresis with the temperature signal's phase are calculated. Combined with aging parameters arranged in a fixed field order, a phase coupling feature set is constructed. Specifically, the calculation of the phase relationship between each frequency band, the short-time stability of each frequency band's phase, and the relative hysteresis with the temperature signal's phase is as follows: For any two electrical frequency bands within the same time slice, the angle difference between them is constrained to the range of -180° to 180°, and then the cosine is taken to map the phase relationship quantity. The closer the value is to 1, the stronger the synchronization. For each electrical frequency band, the mean square amplitude of the first-order difference is calculated within a short-time sliding window based on the phase data. This is then normalized to obtain the short-time stability value. A higher value indicates less phase jitter and better stability in the frequency band. The difference between the temperature phase and the phase of each electrical frequency band is compared in the same time slice, and the average value is taken within the sliding window and normalized to obtain the relative lag. The larger the value, the longer the temperature change leads the electrical change. The aging parameters include the cumulative years of operation of the surge arrester, the number of lightning strikes, the long-term growth rate of leakage current, the change in the nonlinear coefficient of the valve plate, the percentage decrease in insulation resistance, and the increase in surface conductivity due to environmental deposition. A phase coupling graph is constructed using a set of phase coupling features. The nodes of the phase coupling graph correspond to each frequency band and temperature. The weights of the edges are jointly determined by the phase relationship quantity and the short-time stability quantity. An aging parameter is introduced to uniformly adjust the weights of all edges. Based on the phase coupling graph, a deterministic embedding is performed while maintaining the consistency of the edge weight order and the main direction to generate a manifold coordinate vector. The manifold coordinate vector and the corresponding phase coupling graph weights are incorporated to form a phase coupling risk manifold, which is jointly defined by frequency domain phase relationship and dual-domain consistency constraints of temperature and aging context. A risk potential energy index is calculated on the phase-coupled risk manifold. This risk potential energy index is obtained by a weighted aggregation of cross-band phase inconsistency, temperature phase hysteresis, aging level, and phase change rate penalty, wherein: Cross-band phase inconsistency is obtained by statistically analyzing the dispersion of phase differences between electrical frequency bands within the same time slice and then normalizing the dispersion proportionally. Temperature phase lag is obtained by measuring the average lag time of the temperature signal phase relative to the fundamental electrical phase and converting the lag time to a 0-1 scale according to the operational experience baseline. The aging level and phase change rate penalty are obtained by mapping the combined years of operation, lightning strike count, long-term leakage current growth rate and recent phase change rate into a single aging penalty value, and then weighting and normalizing it.

[0024] In this embodiment, obtaining the multi-step risk potential prediction sequence includes: Set a fixed historical window length and prediction step size, and combine the embedding vectors of the risk potential index and the phase-coupled risk manifold into time-series input samples in chronological order, keeping the input dimension consistent. An improved FEDformer model is constructed, which sequentially includes an input embedding layer, a dual-frequency domain cross-attention module, a phase consistency gating module, a residual frequency band aggregation module, and a prediction output head, wherein: The dual-frequency domain cross-attention module separates the input embedding vector into low-frequency and high-frequency channels, calculates the attention weights of each channel, performs a cross-product operation, and reassembles the two cross-features to output a multi-frequency coupled feature sequence. The phase consistency gating module receives a multi-band coupled feature sequence and a phase difference sequence at the same time step. It performs a short-time stability assessment of the phase difference, multiplies the assessment value as a gating factor into the corresponding feature channel, suppresses channels with instantaneous phase jitter exceeding a preset threshold, strengthens channels with phase consistency within a stable threshold range, and outputs a feature sequence. Specifically, the short-time stability assessment of the phase difference involves: A fixed-length sliding time window is set for each phase difference channel, and the phase difference values ​​continuously sampled within the window are uniformly expanded to a continuous range of -180° to 180°. Calculate the absolute average value of the phase difference change within the window, and simultaneously calculate the range of the phase difference; if both are lower than the stability threshold determined based on historical operating data, the channel is determined to be stable within the current window; otherwise, it is considered unstable. By mapping stability to a gating factor of 0 or 1, the features of unstable channels are attenuated, while those of stable channels are preserved or appropriately enhanced, resulting in a feature sequence. The residual frequency band aggregation module superimposes the current feature sequence with historical frequency band features from each layer using residual connections, fuses them according to preset weights to maintain consistent output dimensions, performs weighted fusion on the superposition result according to preset frequency band weights, and outputs the aggregated feature sequence to the prediction output head, where: The preset weights are 0.60 for the current feature and 0.40 for the historical frequency band features; The preset frequency band weights are: low-frequency main band 0.35, second-low frequency harmonic band 0.25, mid-high frequency harmonic band 0.20, and high-frequency pulse band 0.20. In the improved FEDformer model, the dual-frequency domain cross-attention module and the phase consistency gating module are connected in series, and the residual frequency band aggregation module is located at the last stage and is connected in parallel with the prediction output head. Set the training batch size, learning rate, window length and regularization coefficient, and use a joint loss function composed of risk potential error term and short-term trend smoothing term to supervise the training of the improved FEDformer model. After each training cycle, execute an early stopping strategy to avoid overfitting. During the inference phase, the improved FEDformer model is driven by online input time-series samples of equal dimension. This is transformed by the prediction output head to obtain a multi-step risk potential prediction sequence, specifically as follows: The aggregated feature sequence is fed into a linear mapping unit that is executed step by step over time. The linear mapping unit uses a weight matrix that matches the number of input feature channels to compress high-dimensional features into a single risk potential channel. To eliminate the risk of overfitting, a small dropout layer is cascaded after the linear mapping unit to randomly mask the responses of some neurons in the mapping result, and batch normalization is used to unify the output distribution at different time steps. The continuous output is segmented and shaped according to the predicted step size, and then combined in chronological order to form a sequence of predicted risk potential values ​​for each future control cycle.

[0025] The existing FEDformer architecture employs a single-branch structure consisting of input embedding, frequency domain selection, time-frequency decomposition, time-series feature decoding, and a linear output layer. It primarily relies on frequency domain selection operators to extract low-frequency trends and perform time-domain mapping at the decoding end. It lacks explicit modeling of the mutual influence between different frequency bands and does not consider the short-term stability of electrical-thermal phase coupling. This invention introduces three structural innovations: First, a dual-frequency domain cross-attention module is added, performing cross-product between low-frequency and high-frequency channels to explicitly capture the coupling characteristics of multiple frequency bands. Second, a phase consistency gating module is inserted, utilizing windowed phase difference stability to gating and suppress and enhance feature channels, improving robustness to instantaneous phase jitter. Third, a residual frequency band aggregation module is added, superimposing current-historical feature residuals and adaptively fusing them according to frequency band weights to unify multi-layer outputs into a risk potential-sensitive feature space. This forms a three-level enhancement path of cross-attention, gating-stabilized phase, and residual aggregation, improving FEDformer's coupling perception capability, short-term stability, and output accuracy in multi-step prediction scenarios of surge arrester risk potential.

[0026] In this embodiment, the constraints for forming safe and feasible actions at the current moment include: Obtain the current risk potential energy index and multi-step risk potential energy prediction sequence, form a prediction sequence covering multiple future control cycles in chronological order, and record the corresponding time index; For each future control cycle, a return coefficient and step weight associated with the current risk potential energy are set. The upper limit of the allowed risk potential energy for the cycle is calculated, and a return judgment list arranged in chronological order is generated. A frequency domain allocation table is established simultaneously, and the incremental contribution of each frequency band in the multi-step risk potential energy prediction sequence is recorded as an independent entry. Specifically, the upper limit of the allowed risk potential energy for the cycle is calculated as follows: Read the risk potential energy index at the current moment, combine it with equipment type, years of operation and seasonal environmental factors, look up the pre-established safe operation benchmark table, and give the corresponding benchmark upper limit starting value for each future cycle; The reversal coefficient is set according to the speed of risk evolution. The faster the risk potential energy rises, the smaller the reversal coefficient is. The reversal coefficient is multiplied by the starting value of the benchmark upper limit to obtain the upper limit value after the initial reduction, so that the acceptable upper limit in the rapid rise phase is automatically tightened. For multiple consecutive forecast periods, the characteristic of step weight increasing with time is introduced. The time weight is then applied to the lowered upper limit value to make the upper limit of the longer period more lenient, and finally a list of upper limits arranged in chronological order is formed. Establish a frequency domain allocation table, specifically as follows: Based on the results of multi-band frequency domain analysis, the risk contribution of each frequency band in the multi-step risk potential energy prediction sequence is divided into low-frequency main band, second low-frequency harmonic band, mid-high frequency harmonic band and high-frequency pulse band, and a one-to-one correspondence between prediction time step and frequency band category is established as the basic structure of frequency domain allocation table. For each future control cycle, the change in the predicted risk potential energy of the cycle relative to the current risk potential energy is calculated, and the change is broken down by frequency band source. The risk increment corresponding to each frequency band in the cycle is recorded separately, forming a record entry with time step-frequency band-increment contribution as the core fields. The frequency band increment contribution entries for each future control cycle are summarized in chronological order to generate a frequency domain allocation table, which is then refreshed synchronously with each prediction update. Based on the comparison of predicted values ​​with corresponding upper limits item by item in the turnaround judgment list, a multi-step safety judgment result is generated. When an over-limit period occurs, a step-by-step risk budget recovery mechanism is triggered, and the over-limit amount is recovered from the frequency domain allocation table according to a preset priority, forming a turnaround correction instruction. Specifically, the triggering of the step-by-step risk budget recovery mechanism and the recovery of over-limit amount from the frequency domain allocation table according to a preset priority are as follows: When the predicted risk potential value of any future control cycle is higher than the allowable upper limit of the cycle, the cycle is marked as an over-limit cycle, the corresponding over-limit range is recorded, the risk budget recovery process is triggered, and the over-limit range is used as the target risk amount to be recovered. Based on the magnitude of the risk increment contribution of each frequency band in the frequency domain allocation table within the period, the priority order for recovery is determined. Priority is given to recovering the frequency bands that contribute more to the risk potential energy and have greater volatility, followed by processing the frequency bands with lower contributions or more gradual changes. Following the established frequency band order, the risk increment of the corresponding frequency band is gradually reduced until the cumulative recovery amount reaches the limit or the recoverable amount of the frequency band is exhausted; after the recovery is completed, the frequency bands, time steps and recovery ratios involved are summarized to generate a backtracking correction instruction; The turnaround decision list, future multi-step safety decision results, and turnaround correction instructions are combined to generate a time-turnaround forward-looking safety barrier, which is given in the form of executable constraints, including: the allowable set and priority of the main track risk level, the value range and restriction direction of the auxiliary track continuous parameters, and the joint feasible combination relationship between the main track and the auxiliary track. The time-reversal forward-looking safety barrier is mapped inversely to the current control action space, forming the safe and feasible action constraints at the current moment.

[0027] In this embodiment, the actions for generating candidate risk levels and candidate continuous parameters include: At the current moment, obtain the state input consisting of the risk potential energy index and the phase-coupled risk manifold embedding vector, receive the safe and feasible action constraints, and set the discrete value set of the main track risk level action and the value range of the auxiliary track continuous parameter action. The topological potential energy guiding layer generates a guiding direction along the direction of risk potential decrease on the phase-coupled risk manifold. It samples along the guiding direction at a preset step size and sampling number to generate a set of candidate guiding actions. After curvature adaptive step size control and direction calibration, it retains candidates that are consistent with the guiding direction and meet the upper limit of the step size, where: The preset step size is set to 5% to 10% of the allowable adjustment range of the current control parameter as the base step size; The number of samples is fixed at 8 to 12. The backtracking constraint layer guides the matching and filtering of candidate action sets with safe and feasible action constraints, eliminating candidates that are not within the allowed set and value range. The remaining candidates are processed sequentially according to backtracking priority, adjusting the discrete level of the main track and the continuous parameters of the auxiliary track in turn. Constraint satisfaction verification is completed in the order of main track first, then auxiliary track, until the constraint conditions are met or the candidate is confirmed to be unusable. Specifically, the sequential processing of the remaining candidates according to backtracking priority is as follows: For each remaining candidate action, only check whether the main track risk level meets the current safe action constraints. When the main track risk level exceeds the allowable set, adjust the risk level in descending order until the main track risk level enters the allowable set or reaches the minimum usable level. Under the premise that the risk level of the main rail meets the constraints, the continuous parameters of the auxiliary rail corresponding to the candidate action are verified; when the continuous parameters exceed the safe range, the adjustment range is gradually reduced along the parameter constraint boundary. If, after adjusting the order of the main track and auxiliary track, the candidate action simultaneously satisfies the risk level constraint of the main track and the continuous parameter constraint of the auxiliary track, then the candidate action is marked as available and retained; if, during the adjustment process, any track reaches the boundary and still cannot meet the constraint conditions, then the candidate action is determined to be unavailable and removed from the candidate set. The dual-track strategy fusion layer performs dual-track consistency checks on the retained candidates. It matches the primary track candidate risk level actions with the secondary track candidate continuous parameter actions according to a mapping table, eliminating interlocking and inconsistent combinations. The combinations that pass the check are sorted according to a preset score, outputting an ordered set of candidate risk level actions and candidate continuous parameter action combinations. The mapping table is as follows: When the risk level of the main rail is level one or level two, the continuous parameters of the auxiliary rail are limited to the normal operating range, and only small-amplitude, low-frequency parameter adjustments are allowed to keep the surge arrester in a stable monitoring and normal operating state. When the risk level of the main rail is level three, the continuous parameters of the auxiliary rail are allowed to enter the enhanced adjustment range, and the parameters such as monitoring frequency and alarm sensitivity are adjusted to a moderate degree. When the risk level of the main rail is level four, the continuous parameters of the auxiliary rail are limited to the high-risk control range, allowing the maximum range of parameter adjustments to be performed, which can be used to trigger high-priority responses such as key monitoring, rapid alarm or maintenance preparation. The preset scoring and ranking is based on the expected decrease in risk potential energy as the primary ranking indicator. When the expected decrease in risk potential energy is the same, the combination with the smaller adjustment range of continuous parameters is given priority. When the two indicators are still indistinguishable, the ranking is based on historical execution stability. The action combination that has shown less fluctuation and higher execution success rate in historical operation is given priority, forming a sequence of candidate risk level actions and candidate continuous parameter action combinations arranged from best to worst.

[0028] In this embodiment, the step of generating a dynamic risk level adjustment command and sending it to the surge arrester control device includes: Receive candidate risk level actions and candidate continuous parameter actions, and construct a set of candidate control action pairs according to a one-to-one correspondence. Based on the safety and feasible action constraints, the set of candidate control action pairs is checked one by one for safety, and control action pairs that do not meet the safety and feasible action constraints are eliminated to obtain a set of safe candidate control action pairs. In the set of candidate control action pairs for safety, each control action pair is comprehensively evaluated based on the trend of risk potential energy change, the adjustment range of control action, and historical execution stability indicators. The control action pair with the best evaluation result is selected as the target control action pair. The target control action is converted into a risk level dynamic adjustment command and sent to the surge arrester control device for execution. At the same time, the operation feedback data after execution is collected. The operation feedback data includes the updated risk potential energy index and the corresponding operation status information. Based on the operational feedback data, the embedding parameters of the phase-coupled risk manifold, the prediction parameters of the improved FEDformer model, and the policy parameters of the reinforcement learning adjustment strategy are updated to form the state input for the next control cycle.

[0029] Example 1:

[0030] To verify the feasibility of this invention in practice, it was applied to a 110kV substation. In this area, the lowest winter temperature can reach -40℃, and the average annual low-temperature operation time exceeds 5 months. The surge arresters are subjected to complex conditions of extreme cold, frequent lightning strikes, and superimposed operating voltages. Operational practice shows that in this environment, the degradation process of the zinc oxide varistors in the surge arresters exhibits characteristics of slow changes in leakage current but continuous accumulation of internal risks. Traditional alarm methods based on fixed thresholds are unable to reflect the risk evolution trend in a timely manner, resulting in problems such as delayed risk level adjustments and a lack of continuity in adjustment actions.

[0031] In this scenario, the reinforcement learning-based dynamic adjustment and control method for surge arrester risk levels described in this invention is deployed. During operation, the system continuously collects multi-source operating status data such as surge arrester leakage current, lightning strike count, valve plate temperature, and ambient temperature, and performs time synchronization and preprocessing to form a standard dataset. Subsequently, multi-band frequency domain analysis is performed on the leakage current signal to extract phase features of different frequency bands. Combined with temperature and operating age information, a phase coupling feature set is constructed and embedded into the risk manifold space to form a phase-coupled risk manifold that can reflect the multi-factor coupling effect under extremely cold environments, and the corresponding risk potential energy index is calculated.

[0032] In actual operation, the risk potential energy index is continuously updated with a 10-minute time resolution and input into an improved FEDformer model to predict the trend of risk potential energy changes over the next 30 minutes. The prediction results are used to construct a time-return look-ahead safety barrier, mapping the potential risk change constraints over multiple future control cycles back to the current control action space, forming safe and feasible action constraints at the current moment. Under these constraints, the system executes a three-layer reinforcement learning adjustment strategy to complete risk potential energy guidance, look-ahead safety constraint screening, and dual-track action fusion, generating candidate risk level actions and candidate continuous parameter actions, and outputting dynamic risk level adjustment instructions.

[0033] During a 60-day trial run, the system conducted online monitoring and dynamic adjustment of six 110kV surge arresters within the substation. The results showed that, compared to previous methods, this invention can identify a continuously rising trend in risk potential even before the leakage current shows obvious abnormalities, and adjust the risk level in advance. Because the adjustment process is constrained by a time-retrograde look-ahead safety barrier, the risk level change process is smooth, without frequent jumps or false triggers, and the overall adjustment behavior is more consistent with the operating characteristics of equipment in extremely cold environments.

[0034] Table 1. Statistical Table of Risk Identification and Adjustment Effect of 110kV Surge Arresters in Extremely Cold Environments

[0035] As shown in Table 1, among the multiple 110kV surge arresters operating in extremely cold environments, their service life ranges from 7 to 11 years, generally placing them in the mid-to-late stage of service and making them highly representative of the projects. The average leakage current of each surge arrester is concentrated in the range of 1.76mA to 2.02mA, with relatively small fluctuations, falling within the allowable range for normal operation. This indicates that under extremely cold conditions, relying solely on the leakage current amplitude is insufficient to directly reflect the degree of internal degradation of the equipment. This phenomenon also confirms the limitations of the traditional single-index threshold method in this type of scenario.

[0036] Further observation of the average risk potential energy reveals a correlation between its trend and the years of operation and leakage current. Surge arresters with longer service lives or relatively higher average leakage currents generally exhibit higher average risk potential energy levels. For example, the average risk potential energy of LA-05, which has been in operation for 11 years, reaches 0.52, while the average risk potential energy of LA-02, which has been in operation for 7 years, is only 0.41. This indicates that the risk potential energy index can comprehensively reflect the combined effects of multi-source operating status information and long-term aging factors, making it more effective in characterizing the potential risk evolution characteristics of surge arresters in extremely cold environments.

[0037] Regarding the adjustment effect, using the traditional method, the six surge arresters triggered alarms 20 times within the statistical period. The alarms were mainly concentrated on equipment with longer service lives, and some alarms occurred before the risk had fully evolved, exhibiting a certain degree of lag and dispersion. In contrast, the method of this invention performed risk level adjustments 13 times within the same period, significantly reducing the number of adjustments and showing a more reasonable distribution, focusing more on surge arresters with higher risk potential energy levels. The results demonstrate that by introducing phase-coupled risk modeling, look-ahead prediction, and reinforcement learning adjustment mechanisms, this invention can reduce unnecessary intervention while ensuring safety, making risk level adjustments more closely reflect actual operating conditions, demonstrating good engineering feasibility and application value.

[0038] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for dynamic adjustment and control of risk level of lightning arresters based on reinforcement learning, characterized in that, The method comprises the following steps: Collecting multi-source operating state data of the lightning arrester during operation, preprocessing the multi-source operating state data, and generating a standard data set; Performing multi-frequency band frequency domain analysis on the standard data set to extract phase characteristics, constructing a phase coupling feature set and embedding it in a risk manifold space, forming a phase coupling risk manifold and calculating a risk potential index; Taking the risk potential index and the phase coupling risk manifold as inputs, constructing an improved FEDformer model to predict the evolution trend of the risk potential of the lightning arrester in multiple future control periods, and obtaining a multi-step risk potential prediction sequence; According to the multi-step risk potential prediction sequence, a time folding forward-looking safety barrier is constructed, the risk potential change in the future multiple control periods is reversely mapped to the current control action space, and a safe and actionable action constraint condition at the current time is formed; Under the safe and actionable action constraint condition, a three-layer reinforcement learning regulation strategy including a topological potential energy guiding layer, a folding constraint layer, and a double-track strategy fusion layer combining a main track and an auxiliary track is constructed to generate candidate risk level actions and candidate continuous parameter actions; The candidate risk level actions and the candidate continuous parameter actions are fused and safety checked to generate a risk level dynamic regulation instruction and issue it to a lightning arrester control device, and the phase coupling risk manifold, the improved FEDformer model and the reinforcement learning regulation strategy are updated according to the execution feedback.

2. The method of claim 1, wherein the method is characterized by, The multi-source operating state data includes lightning arrester leakage current signals, multi-frequency band frequency domain components of the leakage current, lightning impulse signals, lightning count information, temperature data of the lightning arrester body or the inside, temperature rise rate data, operating parameters reflecting the operating time and aging degree of the lightning arrester, and environmental parameters representing the external operating environment.

3. The method of claim 1, wherein the method further comprises: The preprocessing of the multi-source operating state data includes time alignment processing, outlier detection and elimination processing, missing data compensation processing, and amplitude normalization or standardization processing of the collected multi-source operating state data.

4. The method of claim 1, wherein the method further comprises: The formation of the phase coupling risk manifold and the calculation of the risk potential index include: In the standard data set, a fixed multi-frequency band set is set, the phases of each frequency band of the leakage current and the lightning impulse signal are extracted and time-aligned and phase-expanded to obtain continuous and consistent phase sequences; Based on the phase sequences, the phase relationship quantity between each frequency band, the short-time stability quantity of each frequency band phase, and the relative lag quantity of the temperature signal phase are calculated according to the frequency band pairing relationship, and the phase coupling feature set is formed in a fixed field order combined with the aging parameter; A phase coupling graph is constructed based on the phase coupling feature set, the nodes of the phase coupling graph correspond to each frequency band and temperature, the weights of the edges are determined by the phase relationship quantity and the short-time stability quantity, the aging parameter is introduced to uniformly adjust the weights of all edges, and the manifold coordinate vector is generated by performing deterministic embedding under the premise of maintaining the consistency of the edge weight order and the main direction according to the phase coupling graph; The manifold coordinate vector and the corresponding phase coupling graph weight are included to form a phase coupling risk manifold, which is jointly defined by the dual-domain consistency constraints of the frequency domain phase relationship and the temperature and aging context. The risk potential index is calculated on the phase coupling risk manifold, and the risk potential index is obtained by weighted aggregation of cross-band phase inconsistency, temperature phase lag, aging level and phase change rate penalty.

5. The method of claim 1, wherein the method further comprises: The multi-step risk potential prediction sequence is obtained by: Set the fixed history window length and the prediction step, and form the time sequence input sample by the risk potential index and the embedding vector of the phase coupling risk manifold in time sequence, and keep the input dimension consistent; An improved FEDformer model is constructed, which sequentially includes an input embedding layer, a dual-frequency cross-attention module, a phase consistency gating module, a residual frequency band aggregation module and a prediction output head, wherein: The dual-frequency cross-attention module separates the input embedding vector into low-frequency channels and high-frequency channels, respectively calculates the attention weights, performs cross product operation, reassembles the two features after cross, and outputs a multi-band coupled feature sequence; The phase consistency gating module receives the multi-band coupled feature sequence and the phase difference sequence at the same time step, evaluates the short-time stability of the phase difference, and multiplies the evaluation value into the corresponding feature channel as a gating factor to suppress the channel whose instantaneous phase jitter exceeds the preset threshold and to strengthen the channel whose phase consistency is within the stable threshold range, and outputs the feature sequence; The residual frequency band aggregation module superimposes the current feature sequence and the historical frequency band features of each layer through the residual connection method, fuses them according to the preset weight to keep the output dimension consistent, and performs weighted fusion on the superimposed results according to the preset frequency band weight, and outputs the aggregated feature sequence to the prediction output head; In the improved FEDformer model, the dual-frequency cross-attention module and the phase consistency gating module are connected in series, and the residual frequency band aggregation module is located at the last stage and is connected in parallel with the prediction output head; Set the training batch size, learning rate, window length and regularization coefficient, and use the joint loss function composed of the risk potential error term and the short-term trend smoothing term to supervise the training of the improved FEDformer model, and execute the early stopping strategy at the end of each training period to avoid overfitting; In the inference stage, the improved FEDformer model is driven by online input time sequence samples with the same dimension, and the multi-step risk potential prediction sequence is obtained through the prediction output head.

6. The method of claim 1, wherein the method further comprises: The safety actionable constraint condition at the current time is formed, including: Obtain the risk potential index and the multi-step risk potential prediction sequence at the current time, form the prediction sequence covering multiple future control periods in time sequence and record the corresponding time index; Set the return coefficient and the step weight associated with the current risk potential for each future control period, calculate the upper limit of the risk potential allowed in the period, generate a return judgment list arranged in time sequence, and simultaneously establish a frequency domain allocation table, record the incremental contribution of each frequency band in the multi-step risk potential prediction sequence as an independent item; Based on the return judgment list, compare the predicted value with the corresponding upper limit item by item, generate a future multi-step safety judgment result, and when an over-limit period occurs, trigger a step-by-step risk budget recovery mechanism to recover the over-limit amount from the frequency domain allocation table according to the preset priority, and form a return correction instruction; The turn-back judgment list, the future multi-step safety judgment result and the turn-back correction instruction are combined to generate a time turn-back forward safety barrier in the form of executable constraints, including: an allowed set and priority of a main track risk level, a value range and a limit direction of a secondary track continuous parameter, and a joint feasible combination relationship of the main track and the secondary track; The time turn-back forward safety barrier is reversely mapped to a current control action space to form a safe actionable action constraint condition at the current time.

7. The method of claim 1, wherein the method further comprises: The candidate risk level action and the candidate continuous parameter action are generated, including: At the current time, a state input composed of a risk potential index and a phase-coupled risk manifold embedding vector is obtained, a safe actionable action constraint condition is received, and a discrete value set of the main track risk level action and a value range of the secondary track continuous parameter action are set; A topological potential guide layer generates a guide direction on the phase-coupled risk manifold along a risk potential descent direction, samples the guide direction according to a preset step size and a sampling number to generate a guide candidate action set, and after curvature self-adaptive step size control and direction calibration, candidate items consistent with the guide direction and meeting a step size upper limit are retained; A turn-back constraint layer matches and filters the guide candidate action set with the safe actionable action constraint condition, removes candidate items not in the allowed set and the value range, and in a sequential processing manner according to a turn-back priority, adjusts between the main track discrete level and the secondary track continuous parameter in turn, and completes constraint satisfaction verification according to a processing order of the main track first and the secondary track second, until the constraint condition is met or it is confirmed that the candidate item is unavailable; A dual-track strategy fusion layer performs dual-track consistency checking on the retained candidate items, matches the main track candidate risk level action and the secondary track candidate continuous parameter action according to a mapping table, removes combinations that do not match each other, ranks combinations that pass the checking according to a preset scoring, and outputs an ordered candidate risk level action and candidate continuous parameter action combination set.

8. The method of claim 1, wherein the method further comprises: The risk level dynamic adjustment instruction is generated and delivered to the lightning arrester control device, including: candidate control action pairs are constructed according to a one-to-one correspondence relationship; According to the safe actionable action constraint condition, the candidate control action pair set is safety checked one by one, and control action pairs that do not meet the safe actionable action constraint condition are removed to obtain a safe candidate control action pair set; In the safe candidate control action pair set, based on the risk potential change trend, the control action adjustment amplitude and the historical execution stability index, each control action pair is comprehensively evaluated, and the control action pair with the optimal evaluation result is selected as the target control action pair; The target control action pair is converted into a risk level dynamic adjustment instruction and delivered to the lightning arrester control device for execution, and running feedback data after execution is collected, including updated risk potential indexes and corresponding running state information; Based on the running feedback data, the embedding parameters of the phase-coupled risk manifold, the prediction parameters of the improved FEDformer model and the policy parameters of the reinforcement learning adjustment strategy are updated to form a state input for the next control period.