A method for cable performance optimization based on reinforcement learning

By constructing a partial discharge precursor trajectory library and using reinforcement learning methods, the complexity of the cable's evolution process before partial discharge is solved, enabling adaptive optimization of cable performance and risk identification, thereby improving the stability and efficiency of cable operation.

CN122472271APending Publication Date: 2026-07-28GUANGZHOU YEBEN ELECTRICAL EQUIPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU YEBEN ELECTRICAL EQUIPMENT CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively characterize the complex evolution process of cables before partial discharge occurs, resulting in delayed risk identification. They also lack comprehensive utilization of multi-source time-series evolution information and adaptive optimization mechanisms, making it difficult to achieve continuous optimization of cable performance.

Method used

By constructing a sample set of partial discharge precursor trajectories, a partial discharge precursor trajectory library is generated. Reinforcement learning methods are used for trajectory matching and reward shaping to construct a reinforcement learning state space and action space, outputting cable performance optimization strategies to achieve continuous characterization and adaptive control of cable operating status.

Benefits of technology

It significantly improves the ability to detect potential faults in advance, enables refined risk assessment and adaptive control of cable operating status, and improves system operating efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472271A_ABST
    Figure CN122472271A_ABST
Patent Text Reader

Abstract

The application discloses a kind of cable performance optimization methods based on reinforcement learning, comprising the following steps: obtaining partial discharge precursor track sample set;Extracting track evolution characteristics and carrying out track clustering and template extraction;Current cable state track is matched with precursor track template, and the proximity of partial discharge precursor is calculated and target risk track template is determined;Reinforcement learning state space is constructed;Reinforcement learning action space is constructed;Evolution deviation degree is calculated and reinforcement learning reward function is shaped in reverse reward, and partial discharge precursor reverse reward function is generated;Reinforcement learning decision model is constructed, and cable performance optimization strategy is output, and performance optimization control is executed to target cable, and cable performance optimization result is output.The application utilizes partial discharge precursor track and reinforcement learning method, realizes cable performance optimization control, has the advantages of high risk prediction accuracy and strong regulation and control adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cable operation optimization, and in particular to a cable performance optimization method based on reinforcement learning. Background Technology

[0002] As a critical transmission device in power systems, the operating status of cables directly affects power supply reliability and system security. Existing technologies typically acquire multi-dimensional data such as cable operating load, temperature, humidity, and partial discharge through online monitoring, and then assess cable operating risks using threshold judgment or statistical analysis methods. While some solutions incorporate machine learning methods to identify partial discharge characteristics and provide early warnings of potential cable faults, the overall approach remains primarily based on static rules or single-model analysis, lacking comprehensive utilization of multi-source time-series evolution information.

[0003] Existing technologies struggle to effectively characterize the complex evolution of cables before partial discharge occurs in practical applications, making it difficult to dynamically predict and finely control risk states. Traditional methods often rely on fixed thresholds or single-step judgments, lacking the ability to jointly model state change trends and risk evolution paths, resulting in delayed risk identification. Furthermore, existing control strategies are mostly experience-driven or rule-triggered, lacking adaptive optimization mechanisms that match operating states and risk levels, making it difficult to continuously optimize cable performance while ensuring safety boundaries.

[0004] Therefore, how to provide a cable performance optimization method based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose a cable performance optimization method based on reinforcement learning. This invention utilizes partial discharge precursor trajectories and reinforcement learning methods to achieve optimized control of cable performance, possessing the advantages of high accuracy in risk prediction and strong adaptive control.

[0006] A cable performance optimization method based on reinforcement learning according to an embodiment of the present invention includes the following steps: Historical cable state data of the target cable is acquired, continuous state evolution segments before the occurrence of partial discharge events are extracted, and the continuous state evolution segments are preprocessed to generate a sample set of partial discharge precursor trajectories. Trajectory evolution features are extracted from the partial discharge precursor trajectory sample set, and trajectory clustering and template extraction are performed based on the trajectory evolution features to generate a partial discharge precursor trajectory library; Construct the current cable status trajectory, match the current cable status trajectory with the precursor trajectory templates in the partial discharge precursor trajectory library, calculate the proximity of partial discharge precursors and determine the target risk trajectory template; The reinforcement learning basic state variables are constructed based on the current cable state trajectory, the target risk trajectory template, and the proximity of partial discharge precursors. The reinforcement learning basic state variables are then constrained and mapped based on the target cable operation constraint parameters and the target cable safety boundary parameters to generate the reinforcement learning state space. Based on the cable operation control capability of the target cable, various candidate control actions are determined and a reinforcement learning action space is generated. The degree of evolution deviation of the current cable state trajectory relative to the target risk trajectory template is calculated based on the proximity of partial discharge precursors, and the reinforcement learning reward function is shaped by the proximity of partial discharge precursors and the degree of evolution deviation to generate the partial discharge precursor inverse reward function. A reinforcement learning decision model is constructed based on the reinforcement learning state space, reinforcement learning action space, and inverse reward function for partial discharge precursors. The model outputs a cable performance optimization strategy and performs performance optimization control on the target cable, outputting the cable performance optimization result.

[0007] Optionally, the historical cable status data includes historical operating load data, historical temperature data, historical humidity data, historical insulation status data, historical partial discharge pulse data, historical phase spectrum data, and historical spectrum data. The preprocessing includes trajectory alignment, feature normalization, abnormal segment removal, and time window segmentation.

[0008] Optionally, the generation of the partial discharge precursor trajectory library specifically includes: Based on the partial discharge precursor trajectory sample set, load change features, temperature change features, humidity change features, insulation state change features, partial discharge pulse change features, phase distribution change features, and spectral energy change features are extracted respectively. The various change features are then correlated and fused according to time window identifiers, target cable section identifiers, and partial discharge event identifiers to generate trajectory evolution features. Based on the trajectory evolution characteristics, the partial discharge precursor trajectory sample set is grouped according to the partial discharge type, target cable section, operating conditions and environmental conditions to obtain multiple partial discharge precursor trajectory groups. The trajectory evolution characteristics within each partial discharge precursor trajectory group are then clustered to obtain the corresponding precursor trajectory cluster. Extract precursor trajectory templates that represent the commonalities of trajectory changes within the same precursor trajectory cluster from each precursor trajectory cluster, and establish an index relationship for the precursor trajectory templates according to the partial discharge type, target cable section, operating condition, and environmental condition corresponding to the precursor trajectory cluster to generate a partial discharge precursor trajectory library.

[0009] Optionally, determining the target risk trajectory template specifically includes: The system collects real-time data on the current operating status of the target cable, associates the data with the data collection timestamp, the target cable section identifier, and the monitoring channel identifier, and then extracts the associated data using a sliding time window to generate the current cable status trajectory. The current cable state trajectory is aligned with each precursor trajectory template in the partial discharge precursor trajectory library within a time window, and the trajectory difference is calculated according to the feature dimension order corresponding to the trajectory evolution characteristics to generate the trajectory difference value of the current cable state trajectory relative to each precursor trajectory template. The trajectory difference value is reverse normalized, and the normalized reverse difference characterization value is determined as the proximity of the current cable state trajectory to the partial discharge precursors of each precursor trajectory template. The partial discharge precursor trajectory templates in the partial discharge precursor trajectory library are sorted according to the proximity of partial discharge precursors. The precursor trajectory template with the highest proximity of partial discharge precursors is determined as the target risk trajectory template corresponding to the current cable status trajectory. The partial discharge type, target cable section, operating conditions and environmental conditions corresponding to the target risk trajectory template are output.

[0010] Optionally, the generation of the reinforcement learning state space specifically includes: Based on the current cable status trajectory, acquisition time window identifier, target cable section identifier, and monitoring channel identifier, extract and associate the current trajectory status components of the target cable; Based on the time window order of the target risk trajectory template and the current trajectory state components, extract and arrange the template state components corresponding to the target risk trajectory template; Based on the current trajectory state components, template state components, and proximity to partial discharge precursors, a reinforcement learning-based state variable is constructed to characterize the degree to which the current trajectory deviates from the risk template. Based on the permissible operating range and risk monitoring boundary of the target cable, determine the target cable operating constraint parameters and target cable safety boundary parameters used to constrain the values ​​of the basic state variables in reinforcement learning; Based on the target cable's operating constraint parameters and target cable safety boundary parameters, the reinforcement learning basic state variables are mapped to generate operating constraint state components that characterize the target cable's controllable state and safety boundary state components that characterize the target cable's partial discharge risk approaching state. Based on the operational constraint state components and the safety boundary state components, the controllability and risk approximation within the corresponding time window are determined. Based on the controllability and risk approximation, the state adjustment priority is generated. The controllability, risk approximation, and state adjustment priority are combined with the basic state variables of reinforcement learning to generate constraint-enhanced state variables. State encoding is performed on the constraint-enhancing state variables based on the target cable section identifier, time window identifier, and monitoring channel identifier to generate a reinforcement learning state space.

[0011] Optionally, the generation of the reinforcement learning action space specifically includes: Obtain the allowable load regulation range, cooling equipment start / stop conditions, operating mode switching conditions, adjustable range of inspection tasks, adjustable range of maintenance timing, and configurable range of alarm levels of the power supply system where the target cable is located, and associate the above control conditions according to the target cable section identification, operating conditions, and environmental conditions to determine the cable operation control capability corresponding to the target cable. Based on the load regulation range in the cable operation regulation capability, and combined with the current load status, determine the load regulation actions that the target cable can perform within the corresponding time window; Based on the start-up and shutdown conditions of the cooling equipment, the current temperature status, and the current humidity status, determine the cooling control actions used to adjust the thermal and humidity status of the target cable. Within the range of operating modes that meet the conditions for switching operating modes, the current operating mode status variable is matched with the switchable operating modes to determine the operating mode switching action; Based on the target cable section identification, proximity of partial discharge precursors, degree of risk approach, priority of status control, and adjustable range of inspection tasks, determine the inspection priority adjustment actions for the corresponding target cable section. Based on the partial discharge type, target cable section, operating conditions and environmental conditions corresponding to the target risk trajectory template, determine the maintenance timing adjustment action within the adjustable range of maintenance timing; Based on the proximity of partial discharge precursors, the degree of risk approach, and the priority of status control, determine the alarm level adjustment action within the configurable range of alarm levels; Perform executability checks and action coding on load adjustment actions, cooling control actions, operating mode switching actions, inspection priority adjustment actions, maintenance timing adjustment actions, and alarm level adjustment actions to generate a reinforcement learning action space.

[0012] Optionally, the generation of the partial discharge precursor inverse reward function specifically includes: Based on the time window sequence of the current cable status trajectory and the target risk trajectory template, the current cable status trajectory and the target risk trajectory template are aligned to generate the aligned current trajectory sequence and the aligned risk template sequence. The same-dimensional difference is calculated by performing the same-dimensional difference calculation on the current trajectory sequence and the risk template sequence after alignment, and the trajectory difference component of the current cable state trajectory relative to the target risk trajectory template is obtained. Based on the trajectory difference components and the deviation weights corresponding to each trajectory dimension, the degree of evolutionary deviation of the current cable state trajectory relative to the target risk trajectory template is calculated. Based on the proximity and evolutionary deviation of partial discharge precursors, determine the reverse deviation benefit of the current cable state trajectory from the target risk trajectory template; The precursor approximation penalty is determined based on the proximity of partial discharge precursors and the degree of evolutionary deviation, and the precursor approximation penalty is used as a risk penalty term in the reinforcement learning reward function; The reinforcement learning reward function is shaped by reverse reward based on the reverse deviation benefit and risk penalty term, generating a partial discharge precursor reverse reward function.

[0013] Optionally, the output of the cable performance optimization results specifically includes: By associating the reinforcement learning state space, the reinforcement learning action space, and the inverse reward function for partial discharge precursors, a set of decision samples for various control actions corresponding to the target cable under different states is established. Based on the set of decision samples, a reinforcement learning decision model is constructed, consisting of a state input layer, an action candidate layer, a reward evaluation layer, and a policy output layer. Based on the state input layer, action candidate layer, reward evaluation layer and policy output layer, a policy combination search is performed on the reinforcement learning state space and reinforcement learning action space to generate candidate cable performance optimization policies, and the policy reward evaluation value corresponding to the candidate cable performance optimization policy is determined according to the partial discharge precursor inverse reward function. Candidate cable performance optimization strategies that do not meet the target cable operation constraint parameters or trigger the target cable safety boundary parameters are eliminated, and the candidate cable performance optimization strategy with the highest strategy benefit evaluation value among the remaining candidate cable performance optimization strategies is determined as the cable performance optimization strategy. Generate a sequence of control execution instructions for the target cable according to the cable performance optimization strategy, perform performance optimization control on the target cable according to the sequence of control execution instructions, and collect the operation feedback data of the target cable after performance optimization control; The effectiveness of the cable performance optimization strategy is evaluated based on the target cable operation feedback data, and the cable performance optimization results are output.

[0014] The beneficial effects of this invention are: This invention constructs a partial discharge precursor trajectory sample set and forms a partial discharge precursor trajectory library, structurally expressing the multidimensional evolution of cable operating states before partial discharge occurs. This allows cable operating risks to no longer rely on threshold judgments at a single moment, but can be matched and identified based on historical evolution patterns, thus significantly improving the ability to detect potential faults in advance. Simultaneously, by introducing a matching mechanism between target risk trajectory templates and current cable state trajectories, and combining this with the proximity of partial discharge precursors to quantify the degree of risk, continuous characterization and refined risk assessment of cable operating states are achieved.

[0015] Building upon this foundation, this invention further constructs a reinforcement learning state space, a reinforcement learning action space, and a reverse reward function for partial discharge precursors, transforming the cable operation control process into a risk-avoidance-oriented decision optimization problem. By jointly modeling the degree of evolutionary deviation and the proximity of partial discharge precursors, a reverse reward shaping mechanism is formed, enabling the optimization process to proactively guide the cable operation state away from high-risk trajectories, thereby achieving adaptive control in a dynamically changing environment. This mechanism overcomes the limitations of traditional rule-driven or single-model prediction, allowing the cable control strategy to be updated in real time with changes in the operation state, improving the overall flexibility and intelligence of the control.

[0016] Meanwhile, this invention employs a dual screening process of operational constraint parameters and safety boundary parameters for candidate cable performance optimization strategies, combined with strategy benefit evaluation values ​​for optimal selection. This ensures that the generated control strategy achieves optimal performance output while meeting safety constraints. Coupled with a control execution command sequence and operational feedback closed-loop mechanism, a complete closed-loop optimization process is achieved from risk identification and strategy generation to execution evaluation. This not only ensures the safe operation of cables but also improves system operating efficiency and stability, demonstrating significant engineering application value. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a cable performance optimization method based on reinforcement learning proposed in this invention; Figure 2 This is a flowchart illustrating the determination of the target risk trajectory template for a cable performance optimization method based on reinforcement learning proposed in this invention. Figure 3 This is a flowchart illustrating the output of a cable performance optimization strategy proposed in this invention, which is based on reinforcement learning. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figures 1-3 A cable performance optimization method based on reinforcement learning includes the following steps: Historical cable state data of the target cable is acquired, continuous state evolution segments before the occurrence of partial discharge events are extracted, and the continuous state evolution segments are preprocessed to generate a sample set of partial discharge precursor trajectories. Trajectory evolution features are extracted from the partial discharge precursor trajectory sample set, and trajectory clustering and template extraction are performed based on the trajectory evolution features to generate a partial discharge precursor trajectory library; Construct the current cable status trajectory, match the current cable status trajectory with the precursor trajectory templates in the partial discharge precursor trajectory library, calculate the proximity of partial discharge precursors and determine the target risk trajectory template; The reinforcement learning basic state variables are constructed based on the current cable state trajectory, the target risk trajectory template, and the proximity of partial discharge precursors. The reinforcement learning basic state variables are then constrained and mapped based on the target cable operation constraint parameters and the target cable safety boundary parameters to generate the reinforcement learning state space. Based on the cable operation control capability of the target cable, various candidate control actions are determined and a reinforcement learning action space is generated. The degree of evolution deviation of the current cable state trajectory relative to the target risk trajectory template is calculated based on the proximity of partial discharge precursors, and the reinforcement learning reward function is shaped by the proximity of partial discharge precursors and the degree of evolution deviation to generate the partial discharge precursor inverse reward function. A reinforcement learning decision model is constructed based on the reinforcement learning state space, reinforcement learning action space, and inverse reward function for partial discharge precursors. The model outputs a cable performance optimization strategy and performs performance optimization control on the target cable, outputting the cable performance optimization result.

[0020] In this embodiment, the historical cable status data includes historical operating load data, historical temperature data, historical humidity data, historical insulation status data, historical partial discharge pulse data, historical phase spectrum data, and historical spectrum data. The extraction of continuous state evolution segments includes time positioning based on the time reference point of the partial discharge event, extracting the historical data sequence before the event according to the preset backtracking time length, verifying the time continuity of the extracted data sequence, and performing segment marking processing on the data sequence that meets the time continuity condition. The preprocessing includes trajectory alignment, feature normalization, abnormal segment removal, and time window segmentation processing.

[0021] In this embodiment, the generation of the partial discharge precursor trajectory database specifically includes: Based on the partial discharge precursor trajectory sample set, load change features, temperature change features, humidity change features, insulation state change features, partial discharge pulse change features, phase distribution change features, and spectral energy change features are extracted respectively. The various change features are then correlated and fused according to time window identifiers, target cable section identifiers, and partial discharge event identifiers to generate trajectory evolution features. Load variation characteristics include load variation amplitude, load variation rate, load fluctuation frequency, and load duration; temperature variation characteristics include temperature rise rate, temperature fall rate, temperature fluctuation amplitude, and duration of high temperature; humidity variation characteristics include humidity variation amplitude, humidity variation rate, humidity fluctuation frequency, and duration of high humidity; insulation condition variation characteristics include insulation resistance variation, dielectric loss variation, insulation condition decline rate, and duration of low insulation condition; partial discharge pulse variation characteristics include partial discharge pulse amplitude variation, partial discharge pulse repetition rate variation, partial discharge pulse energy variation, and partial discharge pulse duration; phase distribution variation characteristics include phase concentration variation, phase distribution offset, phase interval proportion variation, and number of phase repetitions; spectral energy variation characteristics include spectral main energy variation, spectral energy concentration variation, spectral peak frequency offset, and frequency band energy proportion variation. Based on the trajectory evolution characteristics, the partial discharge precursor trajectory sample set is grouped according to the partial discharge type, target cable section, operating conditions and environmental conditions to obtain multiple partial discharge precursor trajectory groups. The trajectory evolution characteristics within each partial discharge precursor trajectory group are then clustered to obtain the corresponding precursor trajectory cluster. Partial discharge type is used to characterize the defect forms of partial discharge in the cable insulation structure, including internal air gap discharge, surface discharge, surface discharge, floating electrode discharge, and tip discharge; target cable section is used to identify the spatial location of the cable along its length or structural nodes, including cable body section, joint section, terminal section, and cable sections corresponding to different laying intervals; operating condition is used to characterize the electrical and load status of the cable during operation, including operating voltage level, load current magnitude, load change state, and operating time period; environmental condition is used to characterize the external environmental conditions of the cable, including ambient temperature, ambient humidity, ambient electromagnetic interference intensity, and laying environment type; The trajectory clustering process includes constructing feature vectors for the trajectory evolution features within each partial discharge precursor trajectory group according to the time window order, generating a precursor trajectory feature vector set, performing scale unification processing on the precursor trajectory feature vector set to generate a standardized precursor trajectory feature vector set, constructing a trajectory similarity matrix based on the trajectory similarity between any two standardized precursor trajectory feature vectors in the standardized precursor trajectory feature vector set, performing clustering based on the trajectory similarity matrix to generate initial precursor trajectory clusters, updating cluster centers and reassigning cluster members for the standardized precursor trajectory feature vectors within the initial precursor trajectory clusters, until the cluster member assignment results of two adjacent times meet the preset convergence condition, and determining the initial precursor trajectory cluster that meets the preset convergence condition as the precursor trajectory cluster; Precursor trajectory templates representing the commonalities of trajectory changes within the same precursor trajectory cluster are extracted from each precursor trajectory cluster. An index relationship for the precursor trajectory templates is established according to the partial discharge type, target cable section, operating condition, and environmental condition corresponding to the precursor trajectory cluster, generating a partial discharge precursor trajectory library. The extraction of precursor trajectory templates includes statistically analyzing the common features of the standardized precursor trajectory feature vectors within each precursor trajectory cluster according to the time window order to obtain the common change sequence within the cluster. Based on the common change sequence within the cluster, the key feature nodes, trajectory duration, and trajectory evolution direction are determined. The key feature nodes, trajectory duration, and trajectory evolution direction are then associated with the corresponding partial discharge type, target cable section, operating condition, and environmental condition to generate precursor trajectory templates.

[0022] In this embodiment, the determination of the target risk trajectory template specifically includes: The system collects real-time data on the current operating status of the target cable, associates the data with the timestamp, the target cable section identifier, and the monitoring channel identifier, and then extracts the associated data using a sliding time window to generate the current cable status trajectory. The current operating status data includes current operating load data, current temperature data, current humidity data, current insulation status data, current partial discharge pulse data, current phase spectrum data, and current spectrum data. The current cable state trajectory is aligned with each precursor trajectory template in the partial discharge precursor trajectory library within a time window, and the trajectory difference is calculated according to the feature dimension order corresponding to the trajectory evolution characteristics to generate the trajectory difference value of the current cable state trajectory relative to each precursor trajectory template. The trajectory difference calculation includes: calculating the feature difference values ​​between the current cable status trajectory and each precursor trajectory template in terms of load change characteristics, temperature change characteristics, humidity change characteristics, insulation status change characteristics, partial discharge pulse change characteristics, phase distribution change characteristics, and spectral energy change characteristics, and then weighting and summing each feature difference value according to the preset feature weight to generate the trajectory difference value; The trajectory difference value is reverse normalized, and the normalized reverse difference characterization value is determined as the proximity of the current cable state trajectory to the partial discharge precursors of each precursor trajectory template. The reverse normalization process includes numerically iterating through the trajectory difference values ​​of the current cable state trajectory relative to each precursor trajectory template to determine the maximum and minimum difference values ​​among all trajectory difference values; calculating the difference between the maximum and minimum difference values ​​to obtain the difference range value; subtracting the minimum difference value from each trajectory difference value to obtain the corresponding difference offset value, and dividing the difference offset value by the difference range value to obtain the normalized difference value; subtracting the normalized difference value from one to obtain the reverse normalized difference value; limiting the reverse normalized difference value to the range of zero to one, and determining the limited reverse normalized difference value as the proximity of the partial discharge precursor. The partial discharge precursor trajectory templates in the partial discharge precursor trajectory library are sorted according to the proximity of partial discharge precursors. The precursor trajectory template with the highest proximity of partial discharge precursors is determined as the target risk trajectory template corresponding to the current cable status trajectory. The partial discharge type, target cable section, operating conditions and environmental conditions corresponding to the target risk trajectory template are output.

[0023] In this embodiment, the generation of the reinforcement learning state space specifically includes: Based on the current cable status trajectory, acquisition time window identifier, target cable section identifier, and monitoring channel identifier, the current trajectory status components of the target cable are extracted and associated. Specifically, the current load status quantity, current temperature status quantity, current humidity status quantity, current insulation status quantity, current partial discharge pulse status quantity, current phase distribution status quantity, current spectrum energy status quantity, and current operating mode status quantity are extracted from the current cable status trajectory. The current status quantities are then associated according to the acquisition time window identifier, target cable section identifier, and monitoring channel identifier to generate the current trajectory status components. Based on the time window order of the target risk trajectory template and the current trajectory state components, the template state components corresponding to the target risk trajectory template are extracted and arranged. Specifically, the risk template load change component, risk template temperature change component, risk template humidity change component, risk template insulation state change component, risk template partial discharge pulse change component, risk template phase distribution change component, and risk template spectral energy change component are extracted from the target risk trajectory template and arranged according to the time window order corresponding to the current trajectory state components to generate template state components. Based on the current trajectory state components, template state components, and proximity to partial discharge precursors, a reinforcement learning-based state variable is constructed to characterize the degree to which the current trajectory deviates from the risk template. The construction of the basic state variables for reinforcement learning includes: calculating the corresponding differences between the current trajectory state components and the template state components according to the same time window and the same feature dimension to obtain load deviation components, temperature deviation components, humidity deviation components, insulation state deviation components, partial discharge pulse deviation components, phase distribution deviation components, and spectral energy deviation components; setting deviation weights according to the feature dimension corresponding to each deviation component, multiplying each deviation component by its corresponding deviation weight, and summing the results to obtain the overall trajectory deviation; using the proximity of partial discharge precursors as a risk amplification coefficient to weight and correct the overall trajectory deviation to obtain the risk-corrected deviation; and concatenating each deviation component, the overall trajectory deviation, the proximity of partial discharge precursors, and the risk-corrected deviation in the order of the time window to generate the basic state variables for reinforcement learning. Based on the target cable's permissible operating range and risk monitoring boundaries, the target cable's operating constraint parameters and target cable safety boundary parameters are determined to constrain the values ​​of the basic state variables used in reinforcement learning. Specifically, the load adjustment range, permissible temperature range, permissible humidity range, permissible operating mode range, permissible cooling control range, adjustable range of inspection tasks, adjustable range of maintenance timing, and configurable range of alarm levels are obtained as target cable operating constraint parameters. The lower limit of insulation state safety, the upper limit of partial discharge pulse safety, the abnormal boundary of phase distribution, the abnormal boundary of spectral energy, and the proximity risk threshold of partial discharge precursors are obtained as target cable safety boundary parameters. Based on the target cable's operating constraint parameters and target cable safety boundary parameters, the reinforcement learning basic state variables are mapped to generate operating constraint state components that characterize the target cable's controllable state and safety boundary state components that characterize the target cable's partial discharge risk approaching state. The operation constraint mapping process includes: performing operation constraint mapping on the current load state, current temperature state, current humidity state, and current operation mode state variables in the reinforcement learning basic state variables according to the target cable operation constraint parameters; mapping state variables within the corresponding constraint range to controllable states; mapping state variables outside the corresponding constraint range to restricted states; and generating operation constraint state components. The safety boundary mapping process includes: mapping the current insulation state, current partial discharge pulse state, current phase distribution state, current spectral energy state, and partial discharge precursor proximity in the reinforcement learning basic state variables according to the target cable safety boundary parameters; mapping state quantities that have not reached the safety boundary to safe states; mapping state quantities that have reached the corresponding safety boundary preset buffer range or whose corresponding partial discharge precursor proximity has reached the risk threshold to risk approach states; and mapping state quantities that have crossed the safety boundary to risk over-limit states, thereby generating safety boundary state components. Based on the operational constraint state components and the safety boundary state components, the controllability and risk approximation within the corresponding time window are determined. Based on the controllability and risk approximation, the state adjustment priority is generated. The controllability, risk approximation, and state adjustment priority are combined with the basic state variables of reinforcement learning to generate constraint-enhanced state variables. The generation of constraint-enhanced state variables includes: calculating the distance values ​​between the current load state variable, current temperature state variable, current humidity state variable, and current operating mode state variable and the corresponding upper and lower boundaries of the allowable range; dividing each distance value by the corresponding allowable range span to obtain a normalized control margin value; weighting and summing each normalized control margin value according to a preset weight to obtain a comprehensive control margin; limiting the comprehensive control margin to a value range of zero to one; and determining the limited value as the controllability level. The generation of constraint-enhanced state variables also includes: calculating the distance values ​​between the current insulation state variable, current partial discharge pulse state variable, current phase distribution state variable, current spectral energy state variable, and partial discharge precursor proximity value and the corresponding safety boundary; reverse normalizing each distance value according to the safety boundary direction to obtain a risk proximity characterization value; and then applying a weighted summation method to each risk proximity characterization value. A weighted sum is calculated based on preset risk weights to generate a comprehensive risk approximation quantity, which is then limited to a value between zero and one. This limited value is defined as the risk approximation degree. The risk approximation degree is used as the dominant risk component, and the controllability degree is used as the controllability component. A weighted combination calculation is performed on the risk approximation degree and the controllability degree to obtain a priority score. The priority score is then graded according to a preset grading rule, and the grading result is determined as the state control priority. The controllability degree, risk approximation degree, and state control priority are used as controllability features, risk approximation features, and priority features, respectively. These are then appended to the corresponding deviation components, trajectory comprehensive deviation, partial discharge precursor proximity, and risk correction deviation in the reinforcement learning basic state variables according to the corresponding time window identifiers, generating constrained enhancement state variables. State encoding is performed on the constraint enhancement state variables based on the target cable section identifier, time window identifier, and monitoring channel identifier to generate a reinforcement learning state space. The state encoding includes: determining the spatial location code corresponding to the constraint enhancement state variable according to the target cable section identifier, determining the temporal location code corresponding to the constraint enhancement state variable according to the time window identifier, and determining the data source code corresponding to the constraint enhancement state variable according to the monitoring channel identifier; the spatial location code, temporal location code, and data source code are respectively appended to the current trajectory state component, each deviation component, control capability feature, risk approximation feature, and priority feature corresponding to the constraint enhancement state variable to generate a state encoding vector; the state encoding vector is arranged dimensionally and its value range is limited according to the input dimension requirements of the reinforcement learning decision model to generate the reinforcement learning state space.

[0024] In this embodiment, the generation of the reinforcement learning action space specifically includes: Obtain the allowable load regulation range, cooling equipment start / stop conditions, operating mode switching conditions, adjustable range of inspection tasks, adjustable range of maintenance timing, and configurable range of alarm levels of the power supply system where the target cable is located, and associate the above control conditions according to the target cable section identification, operating conditions, and environmental conditions to determine the cable operation control capability corresponding to the target cable. Based on the load regulation range in the cable operation regulation capability, and combined with the current load status, determine the load regulation actions that the target cable can perform within the corresponding time window. The determination of the load regulation actions includes: determining the load increase boundary, load decrease boundary, and load hold boundary according to the load regulation range; comparing the current load status with the load regulation range to generate load increase actions, load decrease actions, and load hold actions; and classifying the load increase actions and load decrease actions according to the load regulation magnitude to obtain the load regulation actions. Based on the start-up and shutdown conditions of the cooling equipment, the current temperature status, and the current humidity status, determine the cooling control actions used to adjust the thermal and humidity status of the target cable. The determination of the cooling control actions includes: determining the cooling start-up action, cooling enhancement action, cooling maintenance action, and cooling reduction action based on the allowable temperature range, allowable humidity range, and start-up and shutdown conditions of the cooling equipment; and classifying the cooling start-up action, cooling enhancement action, and cooling reduction action according to the output intensity of the cooling equipment to obtain the cooling control actions. Within the range of operating modes that meet the conditions for switching operating modes, the current operating mode status quantity is matched with the switchable operating modes to determine the operating mode switching action. The determination of the operating mode switching action includes: obtaining the normal operating mode, load reduction operating mode, standby power supply mode, maintenance preparation mode and risk isolation mode that the target cable is allowed to switch to; matching the current operating mode status quantity with each allowed switching operating mode; eliminating operating modes that do not meet the conditions for switching operating modes; and retaining the operating modes that meet the conditions for switching operating modes as the operating mode switching action. Based on the target cable section identification, proximity of partial discharge precursors, risk approach level, status control priority, and the adjustable range of the inspection task, the inspection priority adjustment actions for the corresponding target cable section are determined. The determination of the inspection priority adjustment actions includes: determining the inspection priority enhancement actions, inspection priority maintenance actions, and inspection priority reduction actions for each section of the target cable according to the target cable section identification, proximity of partial discharge precursors, risk approach level, and status control priority. The inspection priority enhancement actions and inspection priority reduction actions are then classified into levels according to the adjustable range of the inspection task to obtain the inspection priority adjustment actions. Based on the partial discharge type, target cable section, operating conditions, and environmental conditions corresponding to the target risk trajectory template, maintenance timing adjustment actions are determined within the adjustable range of maintenance timing. The determination of maintenance timing adjustment actions includes: determining advance maintenance actions, maintenance hold actions, and maintenance postponement actions according to the partial discharge type, target cable section, operating conditions, and environmental conditions corresponding to the target risk trajectory template; and classifying the advance maintenance actions and maintenance postponement actions according to the time range allowed for adjustment in the maintenance plan to obtain the maintenance timing adjustment actions. Based on the proximity of partial discharge precursors, the degree of risk approach, and the priority of state control, the alarm level adjustment action is determined within the configurable range of alarm levels. The determination of the alarm level adjustment action includes: determining the alarm level upgrade action, alarm level maintenance action, and alarm level down action according to the proximity of partial discharge precursors, the degree of risk approach, and the priority of state control, and limiting the alarm level upgrade action and alarm level down action according to the preset alarm level classification rules to obtain the alarm level adjustment action. Perform executability verification and action coding on load adjustment actions, cooling control actions, operation mode switching actions, inspection priority adjustment actions, maintenance timing adjustment actions, and alarm level adjustment actions to generate a reinforcement learning action space. Executability verification and action coding include: obtaining the corresponding action target, action segment, action execution time window, and action amplitude parameters for each load adjustment action, cooling control action, operation mode switching action, inspection priority adjustment action, maintenance timing adjustment action, and alarm level adjustment action; comparing each candidate action with the target cable's operation constraint parameters, including load adjustment range, allowable temperature range, allowable humidity range, allowable operation mode range, allowable cooling control range, adjustable inspection task range, adjustable maintenance timing range, and configurable alarm level range, eliminating candidate actions that exceed the corresponding operation constraint range; and matching the remaining candidate actions with the target cable's safety boundary parameters, including the insulation state safety lower limit, partial discharge pulse safety upper limit, phase distribution abnormal boundary, spectral energy abnormal boundary, and partial discharge precursor proximity risk threshold, eliminating any remaining candidate actions. Candidate actions that, upon execution, will cause the corresponding state variable to exceed the safety boundary are identified. Candidate actions that pass the executability verification are categorized by action type into load adjustment actions, cooling control actions, operating mode switching actions, inspection priority adjustment actions, maintenance timing adjustment actions, and alarm level adjustment actions. Within each action type, hierarchical coding is performed according to action direction and action amplitude to generate action type code, action direction code, and action amplitude code. The action type code, action direction code, and action amplitude code are associated and concatenated with the target cable section identifier and time window identifier to generate action code vectors. The entire set of action code vectors is defined as the reinforcement learning action space, where the action amplitude parameters include load adjustment amplitude, cooling intensity adjustment amplitude, operating mode switching span, inspection priority change level, maintenance time adjustment length, and alarm level change level.

[0025] In this embodiment, the generation of the inverse reward function for partial discharge precursors specifically includes: Based on the time window sequence of the current cable status trajectory and the target risk trajectory template, the current cable status trajectory and the target risk trajectory template are aligned to generate the aligned current trajectory sequence and the aligned risk template sequence. The same-dimensional difference is calculated by performing the same-dimensional difference calculation on the current trajectory sequence and the risk template sequence after alignment, and the trajectory difference component of the current cable state trajectory relative to the target risk trajectory template is obtained. The trajectory difference components are obtained by calculating the following: the load difference between the current load state quantity and the load change component of the risk template, the temperature difference between the current temperature state quantity and the temperature change component of the risk template, the humidity difference between the current humidity state quantity and the humidity change component of the risk template, the insulation state difference between the current insulation state quantity and the insulation state change component of the risk template, the partial discharge pulse difference between the current partial discharge pulse state quantity and the partial discharge pulse change component of the risk template, the phase distribution difference between the current phase distribution state quantity and the phase distribution change component of the risk template, and the spectral energy difference between the current spectral energy state quantity and the spectral energy change component of the risk template. Based on the trajectory difference components and the deviation weights corresponding to each trajectory dimension, the degree of evolutionary deviation of the current cable state trajectory relative to the target risk trajectory template is calculated. The calculation of the degree of evolution deviation includes: setting deviation weights according to the trajectory dimensions corresponding to load differences, temperature differences, humidity differences, insulation state differences, partial discharge pulse differences, phase distribution differences, and spectral energy differences; multiplying each trajectory difference component with its corresponding deviation weight and summing the results to obtain the comprehensive trajectory difference; and normalizing the comprehensive trajectory difference according to the trajectory duration corresponding to the target risk trajectory template to generate the degree of evolution deviation. Based on the proximity and evolutionary deviation of partial discharge precursors, determine the reverse deviation benefit of the current cable state trajectory from the target risk trajectory template; The determination of the reverse deviation return includes: taking the degree of evolutionary deviation as the risk deviation component and the proximity of partial emission precursors as the risk suppression weight, and weighting the risk deviation component to obtain the precursor deviation return; when the proximity of partial emission precursors reaches a preset high proximity threshold and the degree of evolutionary deviation is lower than a preset low deviation threshold, the precursor deviation return is adjusted downward; when the proximity of partial emission precursors is lower than the preset low proximity threshold and the degree of evolutionary deviation reaches a preset high deviation threshold, the precursor deviation return is adjusted upward, and the adjusted precursor deviation return is determined as the reverse deviation return. The precursor approximation penalty is determined based on the proximity of partial discharge precursors and the degree of evolutionary deviation, and the precursor approximation penalty is used as a risk penalty term in the reinforcement learning reward function; The determination of the precursor approximation penalty includes: using the proximity of partial emission precursors as a risk approximation component, and the inverse representation value of the degree of evolutionary deviation as a template approximation component; weighting the risk approximation component and the template approximation component to obtain the precursor approximation penalty; when the proximity of partial emission precursors reaches a preset high proximity threshold and the degree of evolutionary deviation is lower than a preset low deviation threshold, the precursor approximation penalty is increased; when the proximity of partial emission precursors is lower than a preset low proximity threshold and the degree of evolutionary deviation reaches a preset high deviation threshold, the precursor approximation penalty is decreased. The corrected precursor approximation penalty is determined as the risk penalty term of the reinforcement learning reward function. The reinforcement learning reward function is shaped by the reverse deviation benefit and risk penalty term to generate the partial discharge precursor reverse reward function; The reverse reward shaping process includes: obtaining the basic operational benefits, regulatory action costs, and safety constraint penalties from the reinforcement learning reward function; weighting and summing the reverse deviation benefits and basic operational benefits according to preset fusion weights to generate deviation guidance benefits; weighting and superimposing the risk penalty and safety constraint penalties according to preset penalty weights to generate comprehensive risk penalty; subtracting the comprehensive risk penalty from the deviation guidance benefits and deducting it in conjunction with the regulatory action costs to generate an initial shaping reward value; performing segmented normalization on the initial shaping reward value based on the target cable segment identifier and time window identifier, and limiting the normalized shaping reward value to a preset reward value range; establishing a reward mapping relationship between the limited shaping reward value and the proximity of partial discharge precursors, the degree of evolutionary deviation, the state regulation priority, and the action encoding vector in the reinforcement learning action space to generate a reverse reward function for partial discharge precursors.

[0026] In this embodiment, the output of the cable performance optimization results specifically includes: By associating the reinforcement learning state space, the reinforcement learning action space, and the inverse reward function for partial discharge precursors, a set of decision samples for various control actions corresponding to the target cable under different states is established. The establishment of the decision sample set includes: combining the state encoding vector in the reinforcement learning state space with the action encoding vector in the reinforcement learning action space according to the target cable section identifier and time window identifier to form state-action samples; calculating the reward output value corresponding to each state-action sample based on the partial discharge precursor inverse reward function, and associating the state encoding vector, action encoding vector, reward output value, partial discharge precursor proximity, risk approximation degree and state control priority to generate the decision sample set; Based on the set of decision samples, a reinforcement learning decision model is constructed, consisting of a state input layer, an action candidate layer, a reward evaluation layer, and a policy output layer. The construction of the reinforcement learning decision-making model includes: constructing a state input layer based on the state encoding vectors in the decision sample set, and using partial discharge precursor proximity, risk approach, state control priority, and target cable section identification as state input fields; constructing an action candidate layer based on the action encoding vectors in the decision sample set, and using load adjustment action, cooling control action, operating mode switching action, inspection priority adjustment action, maintenance timing adjustment action, and alarm level adjustment action as action candidate fields; constructing a reward evaluation layer based on the reward output value in the decision sample set, and using the inverse reward function of partial discharge precursor as the basis for reward calculation; and constructing a strategy output layer based on the state input fields, action candidate fields, and reward calculation basis, and using the candidate action sequence and strategy benefit evaluation value as strategy output fields. Based on the state input layer, action candidate layer, reward evaluation layer and policy output layer, a policy combination search is performed on the reinforcement learning state space and reinforcement learning action space to generate candidate cable performance optimization policies, and the policy reward evaluation value corresponding to the candidate cable performance optimization policy is determined according to the partial discharge precursor inverse reward function. The strategy combination search includes: inputting the state encoding vector output from the state input layer in sequence according to the time window identifier and the target cable section identifier; within each time window, selecting a set of candidate actions that meet the current state constraints from the action candidate layer based on the proximity of partial discharge precursors, the degree of risk approach, and the state control priority in the state encoding vector; combining and splicing the candidate action sets corresponding to each time window in chronological order to form multiple candidate action sequences; and determining the candidate action sequences that meet the preset strategy selection conditions as candidate cable performance optimization strategies by the strategy output layer. The determination of the strategy benefit evaluation value includes: inputting the candidate action sequence corresponding to the candidate cable performance optimization strategy into the reward evaluation layer in the order of time windows; calculating the reward output value corresponding to the candidate action in each time window according to the inverse reward function of partial discharge precursor; accumulating each reward output value in the order of time windows to obtain the cumulative reward value of the candidate cable performance optimization strategy; aggregating the cumulative reward value into segments according to the target cable segment identifier to obtain the segment strategy benefit corresponding to each target cable segment; and weighting and summing the segment strategy benefits according to the risk proximity degree and state control priority corresponding to the target cable segment to obtain the strategy benefit evaluation value corresponding to the candidate cable performance optimization strategy. Candidate cable performance optimization strategies that do not meet the target cable operation constraint parameters or trigger the target cable safety boundary parameters are eliminated, and the candidate cable performance optimization strategy with the highest strategy benefit evaluation value among the remaining candidate cable performance optimization strategies is determined as the cable performance optimization strategy. The control execution command sequence for the target cable is generated according to the cable performance optimization strategy. The performance optimization control of the target cable is performed according to the control execution command sequence, and the operation feedback data of the target cable after the performance optimization control is collected. The control execution command sequence includes load adjustment command, cooling control command, operation mode switching command, inspection priority adjustment command, maintenance timing adjustment command and alarm level adjustment command. The generation of the control execution instruction sequence includes: according to the action type, action direction, action amplitude, target cable section and applicable time window in the cable performance optimization strategy, the load adjustment action, cooling control action, operation mode switching action, inspection priority adjustment action, maintenance timing adjustment action and alarm level adjustment action are processed into instructions to generate the control execution instruction sequence; Performance optimization control includes: performing load adjustment, cooling control, operation mode switching, inspection priority adjustment, maintenance timing adjustment, and alarm level adjustment on the target cable according to the control execution command sequence, and collecting the target cable's operating load data, temperature data, humidity data, insulation status data, partial discharge pulse data, phase spectrum data, and spectrum data after execution to generate target cable operation feedback data; The effectiveness of the cable performance optimization strategy is evaluated based on the target cable operation feedback data, and the cable performance optimization results are output. The performance evaluation includes: recalculating the proximity of partial discharge precursors, the degree of risk approach, and the priority of state control based on the target cable operation feedback data, and comparing them with the proximity of partial discharge precursors, the degree of risk approach, and the priority of state control before the performance optimization control was implemented. The results of precursor risk changes, state control changes, and strategy execution records are generated. The results of precursor risk changes, state control changes, strategy execution records, and corresponding target cable section identifiers are integrated to output the cable performance optimization results.

[0027] Example 1: To verify the feasibility of this invention in practice, it was applied to the operation monitoring scenario of a medium-voltage cable line in a main power transmission channel of a coastal city. This line is in a high-humidity environment for a long time, and the cable joint section is significantly damp. Traditional monitoring methods mainly rely on threshold alarms and periodic inspections, which are difficult to identify the hidden evolution process before partial discharge occurs in a timely manner, resulting in problems of delayed risk identification and untimely control response.

[0028] In this scenario, an online monitoring system continuously acquires data on the target cable's operating load, temperature, humidity, insulation status, partial discharge pulses, phase spectra, and spectrum. The system automatically backtracks historical data prior to partial discharge occurrence, extracts continuous state evolution segments, and constructs a sample set of partial discharge precursor trajectories. Based on this, feature extraction and cluster analysis are performed on the trajectories to form a partial discharge precursor trajectory library. When new state data is generated during cable operation, the system generates the current cable state trajectory in real time and performs time window alignment and difference calculation with each precursor trajectory template in the trajectory library. Weighted fusion yields the trajectory difference value, which is further converted into partial discharge precursor proximity, thereby identifying the target risk trajectory template that best matches the current state.

[0029] Subsequently, the current cable state trajectory, the target risk trajectory template, and the proximity of partial discharge precursors are jointly mapped into a reinforcement learning state space, and an action space is constructed by combining cable operation constraint parameters and safety boundary parameters. The system uses an inverse reward function to guide the optimization direction, causing the cable operation state to gradually deviate from the high-risk trajectory. The reinforcement learning decision model performs strategy combination search based on state input and candidate actions, generating multiple sets of candidate cable performance optimization strategies, and selects the control scheme that meets the operation constraints and has the optimal risk through a strategy benefit evaluation mechanism. Finally, the system transforms the optimization strategy into a sequence of control execution instructions, dynamically adjusting load distribution, cooling control, and operating mode, while simultaneously achieving collaborative optimization by adjusting inspection priority and maintenance timing.

[0030] In actual operation, by comparing the operation records within the same time window before and after implementation, it can be observed that the trend of partial discharge precursor proximity changes is more gradual, the risk approach process is effectively delayed, and the state control response is more timely. The system can continuously output stable cable performance optimization strategies under different sections and operating conditions, and maintain the ability to detect partial discharge risks in advance throughout the continuous operation cycle. Through comprehensive analysis of operation logs, control records, and feedback data, the dynamic control capability and risk prediction capability of the present invention for cable operation status in complex environments have been verified, which can effectively improve the safety and stability of cable operation.

[0031] Table 1. Performance Comparison of the Invention and Traditional Cable Performance Optimization Methods

[0032] As can be clearly seen from Table 1, the method of the present invention is superior to the traditional method in many indicators.

[0033] Regarding the lead time for partial discharge precursor identification, the traditional method is 5.8 hours, while the method of this invention is 7.1 hours, an increase of 1.3 hours. This result demonstrates that by constructing a partial discharge precursor trajectory database and performing trajectory matching, this invention can identify the state evolution trend before partial discharge occurs, exhibiting stronger foresight compared to traditional threshold-based judgment methods.

[0034] Regarding the reduction in proximity of partial discharge precursors, the traditional method achieves a reduction of 12.6%, while the method of this invention achieves 18.9%, representing an improvement of 6.3%. In terms of the reduction in risk approach, the traditional method achieves a reduction of 10.4%, while the method of this invention achieves 16.2%, representing an improvement of 5.8%. This improvement stems from the inverse reward shaping mechanism introduced in this invention, which causes the reinforcement learning decision-making model to tend to select regulatory strategies that reduce the proximity of partial discharge precursors and increase the degree of evolutionary deviation during the optimization process, thereby effectively suppressing the development of risk towards higher value ranges.

[0035] Regarding the false alarm rate, the traditional method has a rate of 8.7%, while the method of this invention has a rate of 6.1%, a reduction of 2.6%. This is because the present invention uses multi-dimensional trajectory feature fusion for judgment, rather than a single partial discharge signal to trigger the alarm. This can filter out abnormal judgments caused by environmental interference and short-term fluctuations, thereby reducing false alarms.

[0036] Regarding policy response time, the traditional method takes 18.5 minutes, while the method of this invention takes 13.2 minutes, a reduction of 5.3 minutes. This improvement is due to the predefined action space of reinforcement learning and the policy combination search mechanism, which enables the system to quickly generate and filter candidate cable performance optimization policies after acquiring state information.

[0037] Regarding the inspection task hit rate, the traditional method achieved 76.4%, while the method of this invention achieved 83.9%, an improvement of 7.5%. The cable temperature fluctuation range decreased from 6.8℃ to 5.4℃; and the operational stability score improved from 82.3 to 88.6. These results demonstrate that this invention can optimize the allocation of inspection and control resources based on the degree of risk proximity and the priority of state control, and reduce operational fluctuations through multi-dimensional control actions, thereby improving the overall stability of cable operation.

[0038] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A cable performance optimization method based on reinforcement learning, characterized in that, Includes the following steps: Historical cable state data of the target cable is acquired, continuous state evolution segments before the occurrence of partial discharge events are extracted, and the continuous state evolution segments are preprocessed to generate a sample set of partial discharge precursor trajectories. Trajectory evolution features are extracted from the partial discharge precursor trajectory sample set, and trajectory clustering and template extraction are performed based on the trajectory evolution features to generate a partial discharge precursor trajectory library; Construct the current cable status trajectory, match the current cable status trajectory with the precursor trajectory templates in the partial discharge precursor trajectory library, calculate the proximity of partial discharge precursors and determine the target risk trajectory template; The reinforcement learning basic state variables are constructed based on the current cable state trajectory, the target risk trajectory template, and the proximity of partial discharge precursors. The reinforcement learning basic state variables are then constrained and mapped based on the target cable operation constraint parameters and the target cable safety boundary parameters to generate the reinforcement learning state space. Based on the cable operation control capability of the target cable, various candidate control actions are determined and a reinforcement learning action space is generated. The degree of evolution deviation of the current cable state trajectory relative to the target risk trajectory template is calculated based on the proximity of partial discharge precursors, and the reinforcement learning reward function is shaped by the proximity of partial discharge precursors and the degree of evolution deviation to generate the partial discharge precursor inverse reward function. A reinforcement learning decision model is constructed based on the reinforcement learning state space, reinforcement learning action space, and inverse reward function for partial discharge precursors. The model outputs a cable performance optimization strategy and performs performance optimization control on the target cable, outputting the cable performance optimization result.

2. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The historical cable status data includes historical operating load data, historical temperature data, historical humidity data, historical insulation status data, historical partial discharge pulse data, historical phase spectrum data, and historical spectrum data. The preprocessing includes trajectory alignment, feature normalization, abnormal segment removal, and time window segmentation.

3. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The generation of the partial discharge precursor trajectory library specifically includes: Based on the partial discharge precursor trajectory sample set, load change features, temperature change features, humidity change features, insulation state change features, partial discharge pulse change features, phase distribution change features, and spectral energy change features are extracted respectively. The various change features are then correlated and fused according to time window identifiers, target cable section identifiers, and partial discharge event identifiers to generate trajectory evolution features. Based on the trajectory evolution characteristics, the partial discharge precursor trajectory sample set is grouped according to the partial discharge type, target cable section, operating conditions and environmental conditions to obtain multiple partial discharge precursor trajectory groups. The trajectory evolution characteristics within each partial discharge precursor trajectory group are then clustered to obtain the corresponding precursor trajectory cluster. Extract precursor trajectory templates that represent the commonalities of trajectory changes within the same precursor trajectory cluster from each precursor trajectory cluster, and establish an index relationship for the precursor trajectory templates according to the partial discharge type, target cable section, operating condition, and environmental condition corresponding to the precursor trajectory cluster to generate a partial discharge precursor trajectory library.

4. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The determination of the target risk trajectory template specifically includes: The system collects real-time data on the current operating status of the target cable, associates the data with the data collection timestamp, the target cable section identifier, and the monitoring channel identifier, and then extracts the associated data using a sliding time window to generate the current cable status trajectory. The current cable state trajectory is aligned with each precursor trajectory template in the partial discharge precursor trajectory library within a time window, and the trajectory difference is calculated according to the feature dimension order corresponding to the trajectory evolution characteristics to generate the trajectory difference value of the current cable state trajectory relative to each precursor trajectory template. The trajectory difference value is reverse normalized, and the normalized reverse difference characterization value is determined as the proximity of the current cable state trajectory to the partial discharge precursors of each precursor trajectory template. The partial discharge precursor trajectory templates in the partial discharge precursor trajectory library are sorted according to the proximity of partial discharge precursors. The precursor trajectory template with the highest proximity of partial discharge precursors is determined as the target risk trajectory template corresponding to the current cable status trajectory. The partial discharge type, target cable section, operating conditions and environmental conditions corresponding to the target risk trajectory template are output.

5. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The generation of the reinforcement learning state space specifically includes: Based on the current cable status trajectory, acquisition time window identifier, target cable section identifier, and monitoring channel identifier, extract and associate the current trajectory status components of the target cable; Based on the time window order of the target risk trajectory template and the current trajectory state components, extract and arrange the template state components corresponding to the target risk trajectory template; Based on the current trajectory state components, template state components, and proximity to partial discharge precursors, a reinforcement learning-based state variable is constructed to characterize the degree to which the current trajectory deviates from the risk template. Based on the permissible operating range and risk monitoring boundary of the target cable, determine the target cable operating constraint parameters and target cable safety boundary parameters used to constrain the values ​​of the basic state variables in reinforcement learning; Based on the target cable's operating constraint parameters and target cable safety boundary parameters, the reinforcement learning basic state variables are mapped to generate operating constraint state components that characterize the target cable's controllable state and safety boundary state components that characterize the target cable's partial discharge risk approaching state. Based on the operational constraint state components and the safety boundary state components, the controllability and risk approximation within the corresponding time window are determined. Based on the controllability and risk approximation, the state adjustment priority is generated. The controllability, risk approximation, and state adjustment priority are combined with the basic state variables of reinforcement learning to generate constraint-enhanced state variables. State encoding is performed on the constraint-enhancing state variables based on the target cable section identifier, time window identifier, and monitoring channel identifier to generate a reinforcement learning state space.

6. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The generation of the reinforcement learning action space specifically includes: Obtain the allowable load regulation range, cooling equipment start / stop conditions, operating mode switching conditions, adjustable range of inspection tasks, adjustable range of maintenance timing, and configurable range of alarm levels of the power supply system where the target cable is located, and associate the above control conditions according to the target cable section identification, operating conditions, and environmental conditions to determine the cable operation control capability corresponding to the target cable. Based on the load regulation range in the cable operation regulation capability, and combined with the current load status, determine the load regulation actions that the target cable can perform within the corresponding time window; Based on the start-up and shutdown conditions of the cooling equipment, the current temperature status, and the current humidity status, determine the cooling control actions used to adjust the thermal and humidity status of the target cable. Within the range of operating modes that meet the conditions for switching operating modes, the current operating mode status variable is matched with the switchable operating modes to determine the operating mode switching action; Based on the target cable section identification, proximity of partial discharge precursors, degree of risk approach, priority of status control, and adjustable range of inspection tasks, determine the inspection priority adjustment actions for the corresponding target cable section. Based on the partial discharge type, target cable section, operating conditions and environmental conditions corresponding to the target risk trajectory template, determine the maintenance timing adjustment action within the adjustable range of maintenance timing; Based on the proximity of partial discharge precursors, the degree of risk approach, and the priority of status control, determine the alarm level adjustment action within the configurable range of alarm levels; Perform executability checks and action coding on load adjustment actions, cooling control actions, operating mode switching actions, inspection priority adjustment actions, maintenance timing adjustment actions, and alarm level adjustment actions to generate a reinforcement learning action space.

7. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The generation of the partial discharge precursor inverse reward function specifically includes: Based on the time window sequence of the current cable status trajectory and the target risk trajectory template, the current cable status trajectory and the target risk trajectory template are aligned to generate the aligned current trajectory sequence and the aligned risk template sequence. The same-dimensional difference is calculated by performing the same-dimensional difference calculation on the current trajectory sequence and the risk template sequence after alignment, and the trajectory difference component of the current cable state trajectory relative to the target risk trajectory template is obtained. Based on the trajectory difference components and the deviation weights corresponding to each trajectory dimension, the degree of evolutionary deviation of the current cable state trajectory relative to the target risk trajectory template is calculated. Based on the proximity and evolutionary deviation of partial discharge precursors, determine the reverse deviation benefit of the current cable state trajectory from the target risk trajectory template; The precursor approximation penalty is determined based on the proximity of partial discharge precursors and the degree of evolutionary deviation, and the precursor approximation penalty is used as a risk penalty term in the reinforcement learning reward function; The reinforcement learning reward function is shaped by reverse reward based on the reverse deviation benefit and risk penalty term, generating a partial discharge precursor reverse reward function.

8. The cable performance optimization method based on reinforcement learning according to claim 1, characterized in that, The output of the cable performance optimization results specifically includes: By associating the reinforcement learning state space, the reinforcement learning action space, and the inverse reward function for partial discharge precursors, a set of decision samples for various control actions corresponding to the target cable under different states is established. Based on the set of decision samples, a reinforcement learning decision model is constructed, consisting of a state input layer, an action candidate layer, a reward evaluation layer, and a policy output layer. Based on the state input layer, action candidate layer, reward evaluation layer and policy output layer, a policy combination search is performed on the reinforcement learning state space and reinforcement learning action space to generate candidate cable performance optimization policies, and the policy reward evaluation value corresponding to the candidate cable performance optimization policy is determined according to the partial discharge precursor inverse reward function. Candidate cable performance optimization strategies that do not meet the target cable operation constraint parameters or trigger the target cable safety boundary parameters are eliminated, and the candidate cable performance optimization strategy with the highest strategy benefit evaluation value among the remaining candidate cable performance optimization strategies is determined as the cable performance optimization strategy. Generate a sequence of control execution instructions for the target cable according to the cable performance optimization strategy, perform performance optimization control on the target cable according to the sequence of control execution instructions, and collect the operation feedback data of the target cable after performance optimization control; The effectiveness of the cable performance optimization strategy is evaluated based on the target cable operation feedback data, and the cable performance optimization results are output.