Power transmission line equipment parallel control method and system based on deep reinforcement learning
By improving the Hippo algorithm and TD3 algorithm to construct an optimized ELM neural network and TCAMD algorithm, and combining them with parallel control theory, the control accuracy and stability problems of complex power systems were solved, and efficient dynamic optimization and adaptive control of transmission line equipment were realized.
Patent Information
- Application Number
- CN202510771997.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies suffer from insufficient accuracy and poor stability in the control of complex dynamic power transmission and distribution systems. Traditional control methods are difficult to adapt to frequent changes, and optimization algorithms have limitations. The random generation of input weights and hidden layer thresholds in the ELM model leads to model instability, while the Hippo algorithm suffers from insufficient optimization accuracy due to its small population sample size and numerous control parameters.
The hidden and input layer weights of the ELM neural network are optimized by improving the Hippo algorithm based on the Jaya algorithm. The TCAMD algorithm, which is a triple Critic network, is constructed by combining the TD3 algorithm. The control strategy of the transmission line equipment is trained through the Markov decision process and parallel control training is carried out in a virtual environment.
It significantly improves the control accuracy and disturbance rejection stability of transmission line equipment, enhances the generalization ability and prediction accuracy of the model, and realizes dynamic optimization and adaptive decision-making for complex systems.
Smart Images

Figure CN120879526A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to control methods for power transmission equipment, belonging to the field of digital control of power transmission lines, and particularly to a parallel control method and system for power transmission line equipment based on deep reinforcement learning. Background Technology
[0002] With the increasing complexity of power transmission and distribution systems, higher demands are placed on the real-time performance and accuracy of control. However, traditional control methods based on static models are ill-suited to the frequent changes in complex dynamic systems, resulting in poor control performance. Although the widely adopted time-series analysis and optimization control algorithms have improved control accuracy to some extent, they also have limitations when dealing with complex data from large-scale power grids. They reveal numerous shortcomings when facing the control and optimization of complex power transmission and distribution systems. For example, parallel control methods based on digital twins and parallel intelligence theory heavily rely on accurate mathematical models, but accurate modeling is often difficult to achieve in complex dynamic environments, limiting their application scope and effectiveness. Load forecasting models based on the ELM algorithm suffer from the problem of randomly generated input weights and hidden layer thresholds, leading to insufficient model stability. Although optimization algorithms can address issues such as unstable weight outputs and susceptibility to local minima to some extent, their optimization capabilities remain limited. Furthermore, the traditional Hippo algorithm suffers from problems such as a small population sample size, numerous control parameters, insufficient optimization accuracy, and poor stability when optimizing BP neural networks. Although introducing a dynamic attenuation factor strategy can enhance its adaptability, its overall performance still needs improvement. Summary of the Invention
[0003] The purpose of this invention is to overcome the above-mentioned defects and problems in the prior art and provide a parallel control method and system for transmission line equipment based on deep reinforcement learning with better control accuracy and disturbance rejection stability.
[0004] To achieve the above objectives, the technical solution of this invention is: a parallel control method for transmission line equipment based on deep reinforcement learning, comprising:
[0005] The Hippo algorithm is improved based on the Jaya algorithm, and the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network are optimized based on the improved Hippo algorithm to construct an optimized ELM neural network.
[0006] Based on the dual-Critic network of the TD3 algorithm, a single-Critic network is introduced to construct the TCAMD algorithm, and a Markov decision process for the transmission line system based on the TCAMD algorithm is constructed to obtain an improved TD3 decision transmission line system.
[0007] The agent is trained based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment.
[0008] Based on the control strategy of transmission lines, parallel control theory is introduced to train and optimize the control strategy in a virtual environment, thereby obtaining a parallel control method for transmission line equipment.
[0009] The construction and optimization of the ELM neural network specifically includes:
[0010] Data from power transmission line equipment was used as training samples. , construct with A standard ELM neural network structure with 1 hidden layer neurons; wherein: The input vector for the training samples of the ELM neural network; The output vector of the training samples for the ELM neural network;
[0011] The expression for the standard ELM neural network structure is as follows:
[0012] ;
[0013] in: For the output function, To output weights, For input weights, For the first The unit threshold of each hidden layer For the first Activation functions of neurons in a hidden layer. Input the data matrix. The input vector;
[0014] Based on the standard ELM neural network structure, the improved Hippo algorithm is used to adjust the input weights between the hidden layer and the input layer. and unit threshold Optimization was performed by constructing an optimized ELM neural network; improvements to the Hippo algorithm included:
[0015] Based on Latin hypercube sampling, the population of the hippopotamus algorithm is uniformly sampled to improve the diversity of the hippopotamus population samples.
[0016] Based on the Jaya algorithm, control parameters are introduced to control the step size of the hippopotamus population moving towards the optimal and worst solutions, respectively. The expressions are as follows:
[0017] ;
[0018] in: For the first Individuals updated in the next iteration The variable state value, For the first In the next iteration, individuals The state value, For the current dimension, , Let [the variable] be a random variable with control parameters in the range [0,1]. For the first The state value of the optimal individual in the next iteration. For the first The state value of the worst individual in the next iteration;
[0019] The minimum value of the development stage of the Hippo algorithm is found based on the smooth development mutation method; the smooth development mutation method includes three stages: random sampling, random crossover, and sequence mutation.
[0020] The expression for the random sampling is as follows:
[0021] ;
[0022] in: The ratio related to the dimension, To The result of the function is rounded up. For the iteration progress, This represents the current iteration number. The maximum number of iterations, For data dimensions;
[0023] The expression for the random crossover is as follows:
[0024] ;
[0025] in: For nodes exist The state value at time t, For nodes exist The state value at time t, , They are nodes and nodes The state value;
[0026] The expression for the sequence variation is as follows:
[0027] ;
[0028] in: This is the state value of the previous node.
[0029] The improved TD3 decision transmission line system specifically includes:
[0030] Based on the dual-Critic network structure of the TD3 algorithm, a single-Critic network structure is introduced to obtain the TCD algorithm with a triple-Critic network, the expression of which is as follows:
[0031] ;
[0032] in: For instant rewards, As a discount factor, These are the weighting factors for the triple Critic network. This is the output value of the triple Critic network. This represents the minimum value of the dual Critic network output.
[0033] Based on the TCD algorithm, the maximum value is selected from the dual-Critic network and the triple-Critic network to replace the introduced single-Critic network structure, thus obtaining the TCMD algorithm, whose expression is as follows:
[0034] ;
[0035] in: To obtain the maximum value between the triple Critic network and the double Critic network;
[0036] Based on the multi-time-step averaging method, the target value of the TCMD algorithm is reduced. The bias in the value estimation is used to obtain the TCAMD algorithm, whose expression is as follows:
[0037] ;
[0038] in: The network parameters of the target Critic network are... To calculate the average number of time steps.
[0039] The construction of the Markov decision process for the transmission line system based on the TCAMD algorithm specifically includes:
[0040] The set of components of the Markov decision process is as follows: ,in For state space, For action space; For the reward function;
[0041] The state space comprises the transmission line states, predicted values of the transmission line states, and operational constraints of the transmission line equipment; the transmission line states include current, voltage, equipment, and load states; the expression for the state space is as follows:
[0042] ;
[0043] in: , , The transmission lines are respectively in Voltage, current, and frequency at any given moment; for The line status is constantly predicted by optimizing the ELM neural network. for The load power of the transmission line equipment at all times;
[0044] The action space refers to the adjustment of transmission line equipment parameters and the operation of switching. It needs to be defined within the acceptable range of the physical system and conform to actual operation. Its expression is as follows:
[0045] ;
[0046] in: This represents the change in load power. This is the transformer tap position. This is the reactive power compensation amount for the capacitor. For the switching operation of circuit breakers and disconnectors;
[0047] The reward function uses the deviations in fault range, fault time, line loss, and voltage frequency as penalty terms, and load balancing and power supply reliability as reward terms. Its expression is as follows:
[0048] ;
[0049] ;
[0050] in: , These are the weights of the reward and penalty items for the indicators, respectively. , The first The value of the first reward item and the first Each penalty item value.
[0051] The agent is trained based on an optimized ELM neural network and an improved TD3 decision-making transmission line system to obtain the optimal control decision for the transmission line, specifically including:
[0052] S1. Obtain the raw input electrical data of the transmission line, including current, voltage, frequency, and power; extract features from the raw input data based on the optimized ELM neural network, and use the extracted features as the state input for improving the TD3 decision.
[0053] The expression for extracting features from the original input data is as follows:
[0054] ;
[0055] in: For output features, For activation function, The original input data, For input weights, For the first The unit threshold of each hidden layer;
[0056] S2. Set up the Actor network and the triple Critic network, and initialize the target network parameters of the Actor network and the triple Critic network; wherein: the Actor network outputs features For state Input and output actions ; The Critic network is state-based With action The concatenated vector is the input, and the output is... value;
[0057] S3, Collection Status And based on Actor networks and Strategy selection action ;
[0058] S4. Execute the selected action. And observe the next state. and instant rewards ;
[0059] S5, will Store experiences in the experience pool for updates, and enable priority experience replay when the capacity exceeds the threshold;
[0060] S6. Randomly sample from the experience pool to calculate the TD error, and delay updating the weights of the Actor network and the Critic network.
[0061] S7. Generate actions through the Actor target network. And calculate the comprehensive objective of the triple Critic target network. The value, the steps of which include:
[0062] S71. Take the minimum value between the Critic1 target network and the Critic2 target network in the triple Critic network as the target network. And make a copy;
[0063] S72. Take the target network of the Critic3 target network in the triple Critic network. Values and target networks The maximum value in is And the average was processed using a multi-time-step averaging method;
[0064] S73, Using Weighting Factors right Perform weighting and use weighting factors For copying Perform weighting; calculate the weighted result. Compared with weighted replication The sum of these yields the overall objective. value;
[0065] S8. After each update of the triple Critic network, delay the update of the Actor network, and then perform a soft update of the Actor target network and the triple Critic target network.
[0066] S9. Repeat steps S3-S8 until the predetermined training period is reached or the convergence condition is met, then output the real-time electrical parameter characteristics.
[0067] S10. Input the output real-time electrical parameter characteristics into the improved TD3 decision transmission line system for decision analysis to obtain the control strategy of the transmission line equipment.
[0068] A parallel control system for power transmission line equipment based on deep reinforcement learning, the system comprising:
[0069] ELM neural network building unit is used to improve the Hippo algorithm based on the Jaya algorithm, and to optimize the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network based on the improved Hippo algorithm, thereby constructing an optimized ELM neural network.
[0070] The TD3 decision system construction unit is used to introduce a single Critic network on the basis of the dual Critic network of the TD3 algorithm, construct the TCAMD algorithm, and construct the Markov decision process of the transmission line system based on the TCAMD algorithm to obtain an improved TD3 decision transmission line system.
[0071] The control strategy training unit is used to train the agent based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment.
[0072] The control method acquisition unit is used to train and optimize the control strategy based on the control strategy of the transmission line by introducing parallel control theory in a virtual environment, and obtain the parallel control method of the transmission line equipment.
[0073] Furthermore, for the specific implementation schemes of the ELM neural network construction unit, TD3 decision system construction unit, control strategy training unit, and control method acquisition unit, please refer to the relevant descriptions in the embodiments.
[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0075] This invention discloses a parallel control method and system for transmission line equipment based on deep reinforcement learning. The method first improves the Hippo algorithm based on the Jaya algorithm to construct an optimized ELM neural network. Next, it constructs an improved TD3 decision-making transmission line system based on the TD3 algorithm using a triple Critic network. Then, it trains the agent to obtain the control strategy for the transmission line equipment and introduces parallel control theory to train and optimize the control strategy in a virtual environment, thus obtaining a parallel control method for the transmission line equipment. In application, the optimized algorithm combines the global search capability of the Jaya algorithm with the local search advantage of the Hippo algorithm, effectively avoiding getting trapped in local optima. Furthermore, the use of a triple Critic network solves the Q-value underestimation problem of the TD3 algorithm, significantly improving the generalization ability, prediction accuracy, and disturbance rejection stability of the ELM network. Simultaneously, by incorporating parallel control theory, it achieves dynamic optimization, predictive management, and adaptive decision-making for complex systems in a virtual system, effectively improving the physical control accuracy of the transmission line equipment. Attached Figure Description
[0076] Figure 1 This is a flowchart of the method of the present invention.
[0077] Figure 2 This is an optimization flowchart of the optimized ELM neural network structure in Embodiment 1 of the present invention.
[0078] Figure 3 This is the algorithm network diagram of TCAMD in Embodiment 1 of the present invention.
[0079] Figure 4 This is a schematic diagram of the structure of the parallel control model for transmission lines in Embodiment 1 of the present invention.
[0080] Figure 5 This is a comparison chart of voltage control results in a simulated case scenario in Embodiment 1 of the present invention.
[0081] Figure 6 This is a comparison chart of frequency control results in a simulated case scenario of Embodiment 1 of the present invention.
[0082] Figure 7 This is a comparison chart of indicator parameter values for a simulated case scenario in Embodiment 1 of the present invention.
[0083] Figure 8 This is a comparison chart of the convergence of different algorithms in Embodiment 1 of the present invention.
[0084] Figure 9 This is a system structure diagram of the present invention.
[0085] Figure 10 This is a structural diagram of the device of the present invention.
[0086] In the diagram: ELM neural network building unit 1, TD3 decision system building unit 2, control strategy training unit 3, control method acquisition unit 4, processor 5, memory 6, computer program code 61. Detailed Implementation
[0087] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0088] Example 1:
[0089] See Figure 1 A parallel control method for transmission line equipment based on deep reinforcement learning, comprising:
[0090] The Hippo algorithm is improved based on the Jaya algorithm, and the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network are optimized based on the improved Hippo algorithm to construct an optimized ELM neural network.
[0091] Furthermore, an optimized ELM neural network is constructed, specifically including:
[0092] Data from power transmission line equipment was used as training samples. , construct with A standard ELM neural network structure with 1 hidden layer neurons; wherein: The input vector for the training samples of the ELM neural network; The output vector of the training samples for the ELM neural network;
[0093] The expression for the standard ELM neural network structure is as follows:
[0094] ;
[0095] in: For the output function, To output weights, For input weights, For the first The unit threshold of each hidden layer For the first Activation functions of neurons in a hidden layer. Input the data matrix. The input vector;
[0096] Based on the standard ELM neural network structure, the improved Hippo algorithm is used to adjust the input weights between the hidden layer and the input layer. and unit threshold Optimize and construct an optimized ELM neural network;
[0097] Traditional ELM neural networks use randomly initialized weights for both the input and hidden layers. The input weights and unit thresholds between the hidden and input layers significantly impact the prediction accuracy of the ELM. While the Hippo algorithm can optimize weight parameters and unit thresholds through global search, reducing the impact of randomness, the traditional Hippo algorithm suffers from problems such as a small sample size, numerous control parameters, low convergence efficiency, and low optimization accuracy. To address these issues, this solution improves the Hippo algorithm and uses this improved algorithm to optimize the input weights and unit thresholds between the hidden and input layers of the ELM. The improvements to the Hippo algorithm include:
[0098] Based on Latin hypercube sampling, the population of the hippopotamus algorithm is uniformly sampled to improve the diversity of the hippopotamus population samples.
[0099] Latin hypercube sampling is a method for uniform sampling in a multidimensional space to generate a set of samples that are as uniform as possible and with as few repetitions as possible in each parameter space. The process is as follows:
[0100] Divide [0,1] into several equal parts, and randomly select a sampling point in each interval;
[0101] The position of each sampling point is randomly switched to ensure that the sampling points on each parameter axis are evenly distributed and non-repeating. Based on the hippo algorithm, the overall position is initialized using a Latin hypercube uniformly distributed from the interval [0,1] Sample, and the position of the target hippo individual is calculated. The expression is as follows:
[0102] ;
[0103] in: Location of the target hippopotamus. The upper bound of the interval is... This is the lower bound of the interval; The values generated by the Latin hypercube sampling method are typically located in the interval [0,1]. Indicates the sample size. Related to sampling dimensions;
[0104] To address the issues of traditional Hippo algorithm having numerous control parameters and weak global adaptability, this solution improves upon the traditional Hippo algorithm by incorporating the Jaya algorithm. The Jaya algorithm is a novel metaheuristic algorithm that utilizes the idea of continuous improvement, constantly approaching the optimal solution while moving away from the worst, thereby continuously improving the quality of the solution. The Jaya algorithm is characterized by fewer control parameters and strong global search capability.
[0105] Based on the Jaya algorithm, control parameters are introduced to control the step size of the hippopotamus population moving towards the optimal and worst solutions, respectively. The expressions are as follows:
[0106] ;
[0107] in: For the first Individuals updated in the next iteration The variable state value, For the first In the next iteration, individuals The state value, For the current dimension, , Let [the variable] be a random variable with control parameters in the range [0,1]. For the first The state value of the optimal individual in the next iteration. For the first The state value of the worst individual in the next iteration;
[0108] To address the issues of low convergence efficiency and low optimization accuracy in the traditional Hippo algorithm, this scheme improves the traditional Hippo algorithm by using a smooth evolution mutation method. Smooth evolution mutation includes unordered dimension sampling, random crossover, and sequence mutation. Furthermore, in the fourth stage of the Hippo algorithm, namely the development stage, this scheme uses smooth evolution mutation to find the minimum value, thereby improving the algorithm's convergence efficiency and optimization accuracy.
[0109] The minimum value of the development stage of the Hippo algorithm is found based on the smooth development mutation method; the smooth development mutation method includes three stages: random sampling, random crossover, and sequence mutation.
[0110] Random sampling can prevent the reduction of population sparsity and promote vectorization of programming to reduce runtime; the expression for random sampling is as follows:
[0111] ;
[0112] in: The ratio related to the dimension, To The result of the function is rounded up. For the iteration progress, This represents the current iteration number. The maximum number of iterations, For data dimensions;
[0113] Random crossover can improve exploration capabilities by randomly selecting individuals; the expression for random crossover is as follows:
[0114] ;
[0115] in: For nodes exist The state value at any given time; For nodes exist The state value at any given time; , They are nodes and nodes The state value is used to provide The update is provided for reference;
[0116] Random crossover relies on two random individuals, and when the radius is too small, it will prematurely fall into a local optimum. Therefore, this scheme proposes sequence mutation to solve this problem; thus, sequence mutation and random crossover complement each other, improving the exploration capability; the expression for the sequence mutation is as follows:
[0117] ;
[0118] in: The state value is the state value of the previous node; the formula expresses the state value of the current node. State value The average value of the state of the previous node is taken as the average value. Represents a node exist The state value at any given time.
[0119] See Figure 2 , Figure 2This is an optimization flowchart of the ELM neural network structure constructed using the improved Hippo algorithm in this scheme. The process is as follows: First, determine the network structure of the Extreme Learning Machine (ELM), such as the number of neurons in the input layer, hidden layer, and output layer; and initialize the parameters of the ELM, such as input weights and biases; then initialize the population of the optimization algorithm (such as the set of particles in a particle swarm optimization), calculate the fitness value of each individual (parameter combination), measure the performance of the ELM model under the current parameters (such as prediction error), and then adjust the parameter values of individuals in the population according to the optimization algorithm rules. Calculate the new fitness value of the individuals using the updated parameters, evaluate the model performance, and then check whether the fitness no longer significantly improves (such as reaching the number of iterations or accuracy threshold). If not, continue iterative optimization. After the termination condition is met, output the optimal parameter configuration, and finally use the optimized ELM algorithm for prediction, ending the process.
[0120] After improving the Hippo algorithm using the above method, the specific steps for using the improved Hippo algorithm to optimize the input weights and unit threshold between the hidden layer and the input layer of the ELM neural network are as follows:
[0121] Obtain the initial transmission line equipment dataset, which includes input features and corresponding target values. Normalize the input data and divide it into training set, validation set and test set.
[0122] Initialize the input weights between the hidden layers and the input layer of the ELM neural network. and unit threshold As the initial solution for improving the hippo algorithm; set the hippo population size, maximum number of iterations, search space range (upper and lower limits of weights and thresholds), and define the fitness function (mean squared error of the ELM neural network on the training set); randomly generate the initial population, where each individual represents a set of ELM input weights and hidden layer thresholds;
[0123] Calculate fitness: For each individual, train the ELM using the current weights and thresholds, and calculate the validation set error as the fitness value.
[0124] Based on the social hierarchy of hippos (leaders, followers, loners), the individual positions are updated through an improved search strategy, that is, the input weights and unit thresholds are updated.
[0125] For each individual, decode it into ELM parameters and calculate the validation set error;
[0126] Retain the best individuals, update the population until convergence, and obtain the optimal input weights that minimize the ELM error. and unit threshold Use optimal input weights and unit threshold Construct an optimized ELM neural network.
[0127] The TD3 algorithm, by combining deep learning and policy gradient methods, can learn optimal control decisions in a continuous action space. It is a reinforcement learning method that combines Actor networks and Critic networks, and it performs well in a continuous action space. However, TD3 still has some problems. To address these issues, this proposal improves the TD3 algorithm to achieve the goal of learning optimal control decisions.
[0128] Based on the dual-Critic network of the TD3 algorithm, a single-Critic network is introduced to construct the TCAMD algorithm, and a Markov decision process for the transmission line system based on the TCAMD algorithm is constructed to obtain an improved TD3 decision transmission line system.
[0129] Furthermore, improvements to the TD3 decision transmission line system are obtained, specifically including:
[0130] For the traditional TD3 algorithm To address the issue of underestimation, this solution introduces a single Critic network structure on top of the original dual Critic network structure, transforming the TD3 algorithm into a Triple Critic network depth deterministic strategy gradient (TCD) algorithm, the expression of which is as follows:
[0131] ;
[0132] in: For instant rewards; This is a discount factor used to measure the importance of future rewards; , which are the weighting factors of the triple Critic network, ranging from 0 to 1; This is the output value of the triple Critic network. This represents the minimum value of the dual Critic network output.
[0133] However, the underestimation problem in the TD3 algorithm has not been completely solved, and there is still room for improvement. Based on this observation, this scheme draws inspiration from the method of selecting the minimum value in the dual Critic network of the TD3 algorithm. On the basis of the TCD algorithm, a method of selecting the maximum value in the dual Critic network and the new Critic network is adopted to replace the newly introduced single Critic, thus forming the Triple Critic network goal maximization depth deterministic strategy gradient algorithm (TCMD), the expression of which is as follows:
[0134] ;
[0135] in: To obtain the maximum value between the triple Critic network and the double Critic network;
[0136] However, weighting and maximizing multiple Critic networks may lead to... Overestimating the value can lead to strategy instability and reduce [the effectiveness of the strategy]. To address these issues and improve the accuracy of value estimation, this scheme employs a multi-time-step averaging method and proposes a Triple Critics Average Maximization Deep Deterministic Policy Gradient (TCAMD) algorithm, whose expression is as follows:
[0137] ;
[0138] in: The network parameters of the target Critic network; To determine the optimal number of time steps for averaging, a value of 5 is preferred. Employing multi-time-step averaging helps smooth the algorithm's performance. The variance and volatility of the value and strategy are analyzed to enhance stability.
[0139] See Figure 3 , Figure 3 This is the network diagram of the TCAMD algorithm proposed in this scheme. The diagram illustrates a reinforcement learning system integrating an Actor-Critic architecture, which corely comprises three main parts: environment interaction, experience storage, and network training. Specifically, the agent interacts with the environment, generating state transition data (…). , , , The current state ,action ,award Next state The experience pool stores batches of experience generated from environmental interactions, which are then sampled during training to enable experience replay and improve data utilization efficiency. The Actor refers to the policy network, which generates actions based on the current state and outputs them through the Critic network. Optimize strategies.
[0140] Critic network refers to The value estimation network, in this scheme, is a triple Critic network (Critic1, Critic2, Critic3) used to evaluate action value. The TD error is calculated to update the network parameters. Each network corresponds to a target network, the Critic target network, which is used to process the next state. Actor target network combined with policy noise generation Input the Critic2 target network to obtain the target Then, the target is generated by taking the minimum and maximum values. (such as target) Maximum value (etc.), and finally a stable comprehensive objective is obtained by weighting. It is used to calculate the loss during training to avoid training fluctuations.
[0141] The agent is trained based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment.
[0142] Because the operating environment of transmission line equipment is complex and affected by random factors such as weather, load fluctuations, equipment aging, and human interference, Markov decision processes can be used to quantify these uncertainties through state transition probabilities and reward functions. This helps to formulate adaptive decisions and achieve dynamic adjustment of the optimal control strategy in the actual control center.
[0143] Furthermore, a Markov decision process is constructed, specifically including:
[0144] The set of components of the Markov decision process is as follows: ,in For state space, For action space; For the reward function;
[0145] The state space comprises the transmission line states, predicted values of the transmission line states, and operational constraints of the transmission line equipment; the transmission line states include current, voltage, equipment, and load states; the expression for the state space is as follows:
[0146] ;
[0147] in: , , The transmission lines are respectively in Voltage, current, and frequency at any given moment; for The line status is constantly predicted by optimizing the ELM neural network. for The load power of the transmission line equipment at all times;
[0148] The action space refers to the adjustment of transmission line equipment parameters and the operation of switching. It needs to be defined within the acceptable range of the physical system and conform to actual operation. Its expression is as follows:
[0149] ;
[0150] in: This represents the change in load power. This is the transformer tap position. This is the reactive power compensation amount for the capacitor. For the switching operation of circuit breakers and disconnectors;
[0151] The reward function uses the deviations in fault range, fault time, line loss, and voltage frequency as penalty terms, and load balancing and power supply reliability as reward terms. When the reward function aligns with the optimization objectives of the control center, it can improve power quality, enhance operational stability, and increase power supply reliability. Its expression is as follows:
[0152] ;
[0153] ;
[0154] in: , These are the weights of the reward and penalty items for the indicators, respectively. , The first The value of the first reward item and the first Each penalty item value.
[0155] Based on the control strategy of transmission lines, parallel control theory is introduced to train and optimize the control strategy in a virtual environment, thereby obtaining a parallel control method for transmission line equipment.
[0156] Further, the training optimization steps are as follows:
[0157] S1. Obtain the raw input electrical data of the transmission line, including current, voltage, frequency, and power; extract features from the raw input data based on the optimized ELM neural network, and use the extracted features as the state input for improving the TD3 decision.
[0158] The expression for extracting features from the original input data is as follows:
[0159] ;
[0160] in: For output features, For activation function, The original input data, For input weights, For the first The unit threshold of each hidden layer;
[0161] S2. Set up the Actor network and the triple Critic network, and initialize the target network parameters of the Actor network and the triple Critic network; wherein: the Actor network outputs features For state Input and output actions ; The Critic network is state-based With action The concatenated vector is the input, and the output is... value;
[0162] S3, Collection Status And based on Actor networks and Strategy selection action ;
[0163] S4. Execute the selected action. And observe the next state. and instant rewards ;
[0164] S5, will Store experiences in the experience pool for updates, and enable priority experience replay when the capacity exceeds the threshold;
[0165] S6. Randomly sample from the experience pool to calculate the TD error, and delay updating the weights of the Actor network and the Critic network.
[0166] S7. Generate actions through the Actor target network. And calculate the comprehensive objective of the triple Critic target network. The value, the steps of which include:
[0167] S71. Take the minimum value between the Critic1 target network and the Critic2 target network in the triple Critic network as the target network. And make a copy;
[0168] S72. Take the target network of the Critic3 target network in the triple Critic network. Values and target networks The maximum value in is And the average was processed using a multi-time-step averaging method;
[0169] In this scheme, the multi-time-step averaging method for averaging means that at each time step... Value estimation, for example The estimation may be affected by random factors such as environmental noise, strategy fluctuations, or parameter updates, resulting in estimation errors. Assuming these errors are independently or approximately independent at different time steps, then according to the law of large numbers, when averaging the estimates over K time steps: That is, the variance after averaging decreases as K increases, making the estimated value more stable and closer to the true value.
[0170] S73, Using Weighting Factors right Perform weighting and use weighting factors For copying Perform weighting; calculate the weighted result. Compared with weighted replication The sum of these yields the overall objective. value;
[0171] S8. After each update of the triple Critic network, update the Actor network after a delay of D steps, and then perform a soft update on the Actor target network and the triple Critic target network.
[0172] S9. Repeat steps S3-S8 until the predetermined training period is reached or the convergence condition is met, then output the real-time electrical parameter characteristics.
[0173] S10. Input the output real-time electrical parameter characteristics into the improved TD3 decision transmission line system for decision analysis to obtain the control strategy of the transmission line equipment.
[0174] After acquiring the control strategy, it is trained and optimized in a virtual environment using parallel control theory. Parallel control theory is an intelligent control method for complex systems. By constructing a virtual system parallel to the actual system, a closed-loop control optimization logic of "learning-optimization-verification" is formed. By synchronizing real-time data from the transmission line equipment of the real system, such as electrical parameters like current and voltage, the virtual environment parameters are updated to maintain model consistency. Multiple control strategies are pre-simulated in the virtual environment. For example, candidate actions generated by the TD3 algorithm proposed in this invention are predicted to have an impact on system stability. The optimal strategy is then selected based on the simulation results in the virtual environment and deployed to the real system. Finally, the response data from the real system continuously feeds back into the virtual environment to optimize model accuracy and strategy adaptability.
[0175] In power transmission scenarios, TD3 needs to attempt new actions and execute known optimal actions during continuous control tasks. Blindly executing these actions can lead to serious consequences. Furthermore, TD3 relies on a large amount of environmental interaction data, but collecting the required data through a real system would be prohibitively expensive. To address this issue, this solution introduces a parallel control method to simulate power transmission scenarios in an artificial system. This allows TD3 to explore risk-free in a virtual space, collecting diverse training data to train and optimize control strategies. The control strategies are represented as multi-dimensional continuous action vectors, with each dimension corresponding to a regulation command for a power device. These commands correspond to the aforementioned action space, such as adjusting transformer taps, controlling reactive power compensation equipment, and generating circuit breaker opening and closing suggestions.
[0176] The control decision-making process for transmission line equipment is as follows: First, real-time data related to system status, such as equipment operating parameters like current, voltage, frequency, and power, is collected. Then, the status data is transmitted to the TD3 decision-making transmission line system model to update the model's internal state, accurately reflecting the real-time state and environmental changes of the actual system. Next, control commands are generated by an intelligent agent, and by introducing a parallel control method, the control commands generated by the TD3 decision-making transmission line system model are accurately and effectively reflected in the actual physical system. Then, the response of the actual physical system is monitored in real-time, i.e., the state changes of the actual transmission line after implementation. Finally, an online state detection and evaluation module is set up to check the rationality of resource scheduling and load allocation after the physical system receives control commands from the intelligent agent. If the control efficiency is lower than expected, a new round of digital simulation prediction and control updates is initiated to complete the dynamic adjustment of the physical system. The response data is then fed back to the TD3 decision-making transmission line system model to adjust model parameters, improve its fit to the actual system, and ensure continuous model optimization to improve adaptability.
[0177] See Figure 4 , Figure 4This is the parallel control model for transmission lines in this embodiment. The model mainly consists of three key parts: the implementation layer, the architecture layer, and the infrastructure layer. In the transmission line system, a large amount of data originates from physical physical equipment. Data is collected through various acquisition devices in the infrastructure layer. The data is first aggregated and parsed on the rack, and then the edge computing devices transmit the data to the cloud platform via fiber optic and wireless networks. The cloud platform receives the data and transmits it to the perception layer. Then, data integration and simulation calculations are performed through technologies such as modeling management and simulation services. Finally, human-computer interaction is achieved through virtualization and visualization. Users can issue control commands through a visual interactive interface, which are ultimately sent to the physical equipment of the transmission line in the perception layer, realizing intelligent control of the equipment.
[0178] Figure 5 and Figure 6 This is a comparison chart of voltage and frequency control results in the simulated case scenario of this embodiment. To verify the effectiveness of the parallel control method proposed in this scheme, several comparison schemes were designed, including: Scheme 1 is the parallel control method proposed in this scheme; Scheme 2 is parallel control based solely on improved TD3 decision-making; Scheme 3 is parallel control based solely on ELM prediction based solely on the improved Hippo algorithm; and Scheme 4 is the traditional control method. In this embodiment, a case of abnormal load is set up. In this case scenario, there are two important industrial customers whose user load has increased by 20%. Adjacent splittable lines are set up, and the power supply pressure of the evaluation node is analyzed. The control center formulates a priority power supply strategy for important users and, based on voltage and frequency stability analysis, provides a controllable splitting scheme to improve power supply reliability by balancing the load.
[0179] The case study simulates and verifies a priority load recovery strategy based on state assessment and prediction. This strategy addresses voltage and frequency anomalies caused by abnormal load increases through a priority power supply strategy and an effective load splitting scheme. When the local line experiences insufficient power supply due to load growth, a portion of the load is controlled and split to adjacent lines to support it, ensuring stable local voltage and frequency and meeting the power needs of critical users. As shown in the figure, as the load abnormally increases, voltage and frequency decrease simultaneously. To ensure power supply reliability, this scheme restores voltage and frequency to normal levels in the shortest possible time by reducing user energy consumption and supplementing adjacent lines.
[0180] Figure 7This is a comparison chart of the indicator parameter values for the simulated case scenario in this embodiment. The chart clearly shows that in dynamic quality control tasks such as frequency and voltage regulation, the closed-loop parallel control driven by the digital twin achieved the lowest voltage / frequency settling time of 0.021 / 0.512 and the lowest maximum voltage / frequency dynamic deviation of 0.022 / 0.19. Compared to other solutions, the maximum deviation of voltage and frequency dynamic indicators is reduced by more than 40% and 17% respectively, and the stability overshoot time is shortened by nearly 3.5% and 34% respectively. This significantly reduces the dependence of power grid operation on stability and the effort required for regulation. Furthermore, the strategy proposed in this solution achieves optimal results after the digital twin control commands reach the physical end compared to traditional static physical control. The voltage qualification rate for important users is increased by 11%, the power supply satisfaction rate for important users is increased by 9.8%, and the load balancing is improved by 5.8%, which will significantly reduce the possibility of line overload.
[0181] Figure 8 This is a comparison chart of the convergence of our proposed solution with other algorithms. To verify the effectiveness of the improved algorithm proposed in this solution, the following comparison methods were set up in simulated case scenarios: Method 1 is the improved algorithm proposed in this solution; Method 2 is the algorithm based on HO-ELM-improved TD3; Method 3 is the algorithm based solely on ELM-TD3 prediction; and Method 4 is the algorithm based on ELM-improved TD3.
[0182] As can be seen from the figure, the proposed improved algorithm outperforms other methods. Specifically, the proposed improved algorithm converges after approximately 200 iterations, while comparative methods 2 / 3 / 4 converge after approximately 380 / 400 / 780 iterations, respectively. This indicates that the proposed method improves the speed of exploring the optimal strategy by 47% / 50% / 74%, respectively. Furthermore, compared to other methods, the proposed method achieves the highest reward value of -1.03. These results significantly demonstrate the effectiveness of the proposed method in improving TD3 through a triple Critic network and optimizing the ELM neural network with an improved Hippo algorithm.
[0183] Example 2:
[0184] See Figure 9 A parallel control system for transmission line equipment based on deep reinforcement learning, the system comprising:
[0185] ELM Neural Network Building Unit 1 is used to improve the Hippo algorithm based on the Jaya algorithm, and optimize the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network based on the improved Hippo algorithm, thereby constructing an optimized ELM neural network.
[0186] Furthermore, the ELM neural network building unit 1 constructs and optimizes the ELM neural network according to the following steps:
[0187] Data from power transmission line equipment was used as training samples. , construct with A standard ELM neural network structure with 1 hidden layer neurons; wherein: The input vector for the training samples of the ELM neural network; The output vector of the training samples for the ELM neural network;
[0188] The expression for the standard ELM neural network structure is as follows:
[0189] ;
[0190] in: For the output function, To output weights, For input weights, For the first The unit threshold of each hidden layer For the first Activation functions of neurons in a hidden layer. Input the data matrix. The input vector;
[0191] Based on the standard ELM neural network structure, the improved Hippo algorithm is used to adjust the input weights between the hidden layer and the input layer. and unit threshold Optimization was performed by constructing an optimized ELM neural network; improvements to the Hippo algorithm included:
[0192] Based on Latin hypercube sampling, uniform sampling is performed on the hippopotamus algorithm to improve the diversity of hippopotamus population samples;
[0193] Based on the Jaya algorithm, control parameters are introduced to control the step size of the hippopotamus population moving towards the optimal and worst solutions, respectively. The expressions are as follows:
[0194] ;
[0195] in: For the first Individuals updated in the next iteration The variable state value, For the first In the next iteration, individuals The state value, For the current dimension, , Let [the variable] be a random variable with control parameters in the range [0,1]. For the first The state value of the optimal individual in the next iteration. For the first The state value of the worst individual in the next iteration;
[0196] The minimum value of the development stage of the Hippo algorithm is found based on the smooth development mutation method; the smooth development mutation method includes three stages: random sampling, random crossover, and sequence mutation.
[0197] The expression for the random sampling is as follows:
[0198] ;
[0199] in: The ratio related to the dimension, To The result of the function is rounded up. For the iteration progress, This represents the current iteration number. The maximum number of iterations, For data dimensions;
[0200] The expression for the random crossover is as follows:
[0201] ;
[0202] in: For nodes exist The state value at time t, For nodes exist The state value at time t, , They are nodes and nodes The state value;
[0203] The expression for the sequence variation is as follows:
[0204] ;
[0205] in: This is the state value of the previous node.
[0206] TD3 decision system construction unit 2 is used to introduce a single Critic network on the basis of the dual Critic network of the TD3 algorithm, construct the TCAMD algorithm, and construct the Markov decision process of the transmission line system based on the TCAMD algorithm to obtain an improved TD3 decision transmission line system.
[0207] Furthermore, the TD3 decision system construction unit 2 obtains the TCAMD algorithm according to the following steps:
[0208] Based on the dual-Critic network structure of the TD3 algorithm, a single-Critic network structure is introduced to obtain the TCD algorithm with a triple-Critic network, the expression of which is as follows:
[0209] ;
[0210] in: For instant rewards, As a discount factor, These are the weighting factors for the triple Critic network. This is the output value of the triple Critic network. This represents the minimum value of the dual Critic network output.
[0211] Based on the TCD algorithm, the maximum value is selected from the dual-Critic network and the triple-Critic network to replace the introduced single-Critic network structure, thus obtaining the TCMD algorithm, whose expression is as follows:
[0212] ;
[0213] in: To obtain the maximum value between the triple Critic network and the double Critic network;
[0214] Based on the multi-time-step averaging method, the TCMD algorithm is reduced. The bias in the value estimation is used to obtain the TCAMD algorithm, whose expression is as follows:
[0215] ;
[0216] in: The network parameters of the target Critic network are... To calculate the average number of time steps.
[0217] Furthermore, the TD3 decision system construction unit 2 constructs a Markov decision process for the transmission line system based on the TCAMD algorithm according to the following steps:
[0218] The set of components of the Markov decision process is as follows: ,in For state space, For action space; For the reward function;
[0219] The state space comprises the transmission line states, predicted values of the transmission line states, and operational constraints of the transmission line equipment; the transmission line states include current, voltage, equipment, and load states; the expression for the state space is as follows:
[0220] ;
[0221] in: , , The transmission lines are respectively in Voltage, current, and frequency at any given moment; for The line status is constantly predicted by optimizing the ELM neural network. for The load power of the transmission line equipment at all times;
[0222] The action space refers to the adjustment of transmission line equipment parameters and the operation of switching. It needs to be defined within the acceptable range of the physical system and conform to actual operation. Its expression is as follows:
[0223] ;
[0224] in: This represents the change in load power. This is the transformer tap position. This is the reactive power compensation amount for the capacitor. For the switching operation of circuit breakers and disconnectors;
[0225] The reward function uses the deviations in fault range, fault time, line loss, and voltage frequency as penalty terms, and load balancing and power supply reliability as reward terms. Its expression is as follows:
[0226] ; ;
[0227] in: , These are the weights of the reward and penalty items for the indicators, respectively. , The first The value of the first reward item and the first Each penalty item value.
[0228] Control strategy training unit 3 is used to train the agent based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment.
[0229] Furthermore, for the specific implementation scheme of the control strategy training unit 3, please refer to Embodiment 1.
[0230] Control method acquisition unit 4 is used to train and optimize the control strategy based on the control strategy of the transmission line by introducing parallel control theory in a virtual environment, and obtain the parallel control method of the transmission line equipment.
[0231] Example 3:
[0232] See Figure 10 A parallel control device for power transmission line equipment based on deep reinforcement learning, the device including a processor 5 and a memory 6;
[0233] The memory 6 is used to store computer program code 61 and transmit the computer program code 61 to the processor 5;
[0234] The processor 5 is used to execute the parallel control method for transmission line equipment based on deep reinforcement learning as described in Embodiment 1 according to the instructions in the computer program code 61.
[0235] This embodiment also includes a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed on a computer, the parallel control method for transmission line equipment based on deep reinforcement learning described in Embodiment 1 is implemented.
[0236] Generally, the computer instructions for implementing the method of the present invention can be carried on any combination of one or more computer-readable storage media. Non-transitory computer-readable storage media can include any computer-readable medium except for the signal itself, which is temporarily propagating.
[0237] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EKROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0238] The aforementioned equipment and non-transitory computer-readable storage media can be found in the detailed description of a parallel control method for transmission line equipment based on deep reinforcement learning and its beneficial effects, which will not be repeated here.
[0239] Although embodiments of the present invention have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A parallel control method for transmission line equipment based on deep reinforcement learning, characterized in that, include: The Hippo algorithm is improved based on the Jaya algorithm, and the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network are optimized based on the improved Hippo algorithm to construct an optimized ELM neural network. Based on the dual-Critic network of the TD3 algorithm, a single-Critic network is introduced to construct the TCAMD algorithm, and a Markov decision process for the transmission line system based on the TCAMD algorithm is constructed to obtain an improved TD3 decision transmission line system. The agent is trained based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment. Based on the control strategy of transmission lines, parallel control theory is introduced to train and optimize the control strategy in a virtual environment, thereby obtaining a parallel control method for transmission line equipment.
2. The parallel control method for transmission line equipment based on deep reinforcement learning according to claim 1, characterized in that: The construction and optimization of the ELM neural network specifically includes: Data from power transmission line equipment was used as training samples. , construct with A standard ELM neural network structure with 1 hidden layer neurons; wherein: The input vector for the training samples of the ELM neural network; The output vector of the training samples for the ELM neural network; The expression for the standard ELM neural network structure is as follows: ; in: For the output function, To output weights, For input weights, For the first The unit threshold of each hidden layer For the first Activation functions of neurons in a hidden layer. Input the data matrix. The input vector; Based on the standard ELM neural network structure, the improved Hippo algorithm is used to adjust the input weights between the hidden layer and the input layer. and unit threshold Optimization was performed by constructing an optimized ELM neural network; improvements to the Hippo algorithm included: Based on Latin hypercube sampling, the population of the hippopotamus algorithm is uniformly sampled to improve the diversity of the hippopotamus population samples. Based on the Jaya algorithm, control parameters are introduced to control the step size of the hippopotamus population moving towards the optimal and worst solutions, respectively. The expressions are as follows: ; in: For the first Individuals updated in the next iteration The variable state value, For the first In the next iteration, individuals The state value, For the current dimension, , Let [the variable be] a random variable with control parameters in the range [0,1]. For the first The state value of the optimal individual in the next iteration. For the first The state value of the worst individual in the next iteration; The minimum value of the development stage of the Hippo algorithm is found based on the smooth development mutation method; the smooth development mutation method includes three stages: random sampling, random crossover, and sequence mutation. The expression for the random sampling is as follows: ; in: The ratio related to the dimension, To The result of the function is rounded up. For the iteration progress, This represents the current iteration number. The maximum number of iterations, For data dimensions; The expression for the random crossover is as follows: ; in: For nodes exist The state value at time t, For nodes exist The state value at time t, , They are nodes and nodes The state value; The expression for the sequence variation is as follows: ; in: This is the state value of the previous node.
3. The parallel control method for transmission line equipment based on deep reinforcement learning according to claim 1, characterized in that: The improved TD3 decision transmission line system specifically includes: Based on the dual-Critic network structure of the TD3 algorithm, a single-Critic network structure is introduced to obtain the TCD algorithm with a triple-Critic network, the expression of which is as follows: ; in: For instant rewards, As a discount factor, These are the weighting factors for the triple Critic network. This is the output value of the triple Critic network. This represents the minimum output value of the dual Critic network. Based on the TCD algorithm, the maximum value is selected from the dual-Critic network and the triple-Critic network to replace the introduced single-Critic network structure, thus obtaining the TCMD algorithm, whose expression is as follows: ; in: To obtain the maximum value between the triple Critic network and the dual Critic network; Based on the multi-time-step averaging method, the target value of the TCMD algorithm is reduced. The bias in the value estimation is used to obtain the TCAMD algorithm, whose expression is as follows: ; in: The network parameters of the target Critic network are... To calculate the average number of time steps.
4. The parallel control method for transmission line equipment based on deep reinforcement learning according to claim 1, characterized in that: The construction of the Markov decision process for the transmission line system based on the TCAMD algorithm specifically includes: The set of components of the Markov decision process is as follows: ,in For state space, For action space; For the reward function; The state space comprises the transmission line states, predicted values of the transmission line states, and operational constraints of the transmission line equipment; the transmission line states include current, voltage, equipment, and load states; the expression for the state space is as follows: ; in: , , The transmission lines are respectively in Voltage, current, and frequency at any given moment; for The line status is constantly predicted by optimizing the ELM neural network. for The load power of the power transmission line equipment at all times; The action space refers to the adjustment of transmission line equipment parameters and the operation of switching. It needs to be defined within the acceptable range of the physical system and conform to actual operation. Its expression is as follows: ; in: This represents the change in load power. This is the transformer tap position. This is the reactive power compensation amount for the capacitor. For the switching operation of circuit breakers and disconnectors; The reward function uses the deviations in fault range, fault time, line loss, and voltage frequency as penalty terms, and load balancing and power supply reliability as reward terms. Its expression is as follows: ; ; in: , These are the weights of the reward and penalty items for the indicators, respectively. , The first The value of the first reward item and the first Each penalty item value.
5. The parallel control method for transmission line equipment based on deep reinforcement learning according to claim 1, characterized in that: The agent is trained based on an optimized ELM neural network and an improved TD3 decision-making transmission line system to obtain the optimal control decision for the transmission line, specifically including: S1. Obtain the raw input electrical data of the transmission line, including current, voltage, frequency, and power; extract features from the raw input data based on the optimized ELM neural network, and use the extracted features as the state input for improving the TD3 decision. The expression for extracting features from the original input data is as follows: ; in: For output features, For activation function, The original input data, For input weights, For the first The unit threshold of each hidden layer; S2. Set up the Actor network and the triple Critic network, and initialize the target network parameters of the Actor network and the triple Critic network; wherein: the Actor network outputs features For state Input and output actions ; The Critic network is state-based With action The concatenated vector is the input, and the output is... value; S3, Collection Status And based on Actor networks and Strategy selection action ; S4. Execute the selected action. And observe the next state. and instant rewards ; S5, will Store experiences in the experience pool for updates, and enable priority experience replay when the capacity exceeds the threshold; S6. Randomly sample from the experience pool to calculate the TD error, and delay updating the weights of the Actor network and the Critic network. S7. Generate actions through the Actor target network. And calculate the comprehensive objective of the triple Critic target network. The value is determined by the following steps: S71. Take the minimum value between the Critic1 target network and the Critic2 target network in the triple Critic network as the target network. And make a copy; S72. Take the target network of the Critic3 target network in the triple Critic network. Values and target networks The maximum value in is And the average was processed using a multi-time-step averaging method; S73, Using Weighting Factors right Perform weighting and use weighting factors For copying Perform weighting; calculate the weighted result. Compared with weighted replication The sum of these yields the overall objective. value; S8. After each update of the triple Critic network, delay the update of the Actor network, and then perform a soft update of the Actor target network and the triple Critic target network. S9. Repeat steps S3-S8 until the predetermined training period is reached or the convergence condition is met, then output the real-time electrical parameter characteristics. S10. Input the output real-time electrical parameter characteristics into the improved TD3 decision transmission line system for decision analysis to obtain the control strategy of the transmission line equipment.
6. A parallel control system for transmission line equipment based on deep reinforcement learning, characterized in that, The system includes: ELM neural network building unit (1) is used to improve the Hippo algorithm based on the Jaya algorithm, and optimize the input weights and unit thresholds between the hidden layer and the input layer of the ELM neural network based on the improved Hippo algorithm, and build an optimized ELM neural network. The TD3 decision system construction unit (2) is used to introduce a single Critic network based on the dual Critic network of the TD3 algorithm, construct the TCAMD algorithm, and construct the Markov decision process of the transmission line system based on the TCAMD algorithm to obtain the improved TD3 decision transmission line system. The control strategy training unit (3) is used to train the agent based on the optimized ELM neural network and the improved TD3 decision transmission line system to obtain the control strategy of the transmission line equipment. The control method acquisition unit (4) is used to train and optimize the control strategy in a virtual environment based on the control strategy of the transmission line, and obtain the parallel control method of the transmission line equipment.
7. The parallel control system for transmission line equipment based on deep reinforcement learning according to claim 6, characterized in that: The ELM neural network building unit (1) constructs and optimizes the ELM neural network according to the following steps: Data from power transmission line equipment was used as training samples. , construct with A standard ELM neural network structure with 1 hidden layer neurons; wherein: The input vector for the training samples of the ELM neural network; The output vector of the training samples for the ELM neural network; The expression for the standard ELM neural network structure is as follows: ; in: For the output function, To output weights, For input weights, For the first The unit threshold of each hidden layer For the first Activation functions of neurons in a hidden layer. Input the data matrix. The input vector; Based on the standard ELM neural network structure, the improved Hippo algorithm is used to adjust the input weights between the hidden layer and the input layer. and unit threshold Optimization was performed by constructing an optimized ELM neural network; improvements to the Hippo algorithm included: Based on Latin hypercube sampling, uniform sampling is performed on the hippopotamus algorithm to improve the diversity of hippopotamus population samples; Based on the Jaya algorithm, control parameters are introduced to control the step size of the hippopotamus population moving towards the optimal and worst solutions, respectively. The expressions are as follows: ; in: For the first Individuals updated in the next iteration The variable state value, For the first In the next iteration, individuals The state value, For the current dimension, , Let [the variable be] a random variable with control parameters in the range [0,1]. For the first The state value of the optimal individual in the next iteration. For the first The state value of the worst individual in the next iteration; The minimum value of the development stage of the Hippo algorithm is found based on the smooth development mutation method; the smooth development mutation method includes three stages: random sampling, random crossover, and sequence mutation. The expression for the random sampling is as follows: ; in: The ratio related to the dimension, To The result of the function is rounded up. For the iteration progress, This represents the current iteration number. The maximum number of iterations, For data dimensions; The expression for the random crossover is as follows: ; in: For nodes exist The state value at time t, For nodes exist The state value at time t, , They are nodes and nodes The state value; The expression for the sequence variation is as follows: ; in: This is the state value of the previous node.
8. The parallel control system for transmission line equipment based on deep reinforcement learning according to claim 6, characterized in that: The TD3 decision system construction unit (2) obtains the TCAMD algorithm according to the following steps: Based on the dual-Critic network structure of the TD3 algorithm, a single-Critic network structure is introduced to obtain the TCD algorithm with a triple-Critic network, the expression of which is as follows: ; in: For instant rewards, As a discount factor, These are the weighting factors for the triple Critic network. This is the output value of the triple Critic network. This represents the minimum output value of the dual Critic network. Based on the TCD algorithm, the maximum value is selected from the dual-Critic network and the triple-Critic network to replace the introduced single-Critic network structure, thus obtaining the TCMD algorithm, whose expression is as follows: ; in: To obtain the maximum value between the triple Critic network and the dual Critic network; Based on the multi-time-step averaging method, the TCMD algorithm is reduced. The bias in the value estimation is used to obtain the TCAMD algorithm, whose expression is as follows: ; in: The network parameters of the target Critic network are... To calculate the average number of time steps.
9. The parallel control system for transmission line equipment based on deep reinforcement learning according to claim 6, characterized in that: The TD3 decision system construction unit (2) constructs a Markov decision process for the transmission line system based on the TCAMD algorithm according to the following steps: The set of components of the Markov decision process is as follows: ,in For state space, For action space; For the reward function; The state space comprises the transmission line states, predicted values of the transmission line states, and operational constraints of the transmission line equipment; the transmission line states include current, voltage, equipment, and load states; the expression for the state space is as follows: ; in: , , The transmission lines are respectively in Voltage, current, and frequency at any given moment; for The line status is constantly predicted by optimizing the ELM neural network. for The load power of the power transmission line equipment at all times; The action space refers to the adjustment of transmission line equipment parameters and the operation of switching. It needs to be defined within the acceptable range of the physical system and conform to actual operation. Its expression is as follows: ; in: This represents the change in load power. This is the transformer tap position. This is the reactive power compensation amount for the capacitor. For the switching operation of circuit breakers and disconnectors; The reward function uses the deviations in fault range, fault time, line loss, and voltage frequency as penalty terms, and load balancing and power supply reliability as reward terms. Its expression is as follows: ; ; in: , These are the weights of the reward and penalty items for the indicators, respectively. , The first The value of the first reward item and the first Each penalty item value.
10. The parallel control system for transmission line equipment based on deep reinforcement learning according to claim 6, characterized in that: The control strategy training unit (3) is used to obtain the control strategy of the transmission line equipment according to the following steps: S1. Obtain the raw input electrical data of the transmission line, including current, voltage, frequency, and power; extract features from the raw input data based on the optimized ELM neural network, and use the extracted features as the state input for improving the TD3 decision. The expression for extracting features from the original input data is as follows: ; in: For output features, For activation function, The original input data, For input weights, For the first The unit threshold of each hidden layer; S2. Set up the Actor network and the triple Critic network, and initialize the target network parameters of the Actor network and the triple Critic network; wherein: the Actor network outputs features For state Input and output actions ; The Critic network is state-based With action The concatenated vector is the input, and the output is... value; S3, Collection Status And based on Actor networks and Strategy selection action ; S4. Execute the selected action. And observe the next state. and instant rewards ; S5, will Store experiences in the experience pool for updates, and enable priority experience replay when the capacity exceeds the threshold; S6. Randomly sample from the experience pool to calculate the TD error, and delay updating the weights of the Actor network and the Critic network. S7. Generate actions through the Actor target network. And calculate the comprehensive objective of the triple Critic target network. The value is determined by the following steps: S71. Take the minimum value between the Critic1 target network and the Critic2 target network in the triple Critic network as the target network. And make a copy; S72. Take the target network of the Critic3 target network in the triple Critic network. Values and target networks The maximum value in is And the average was processed using a multi-time-step averaging method; S73, Using Weighting Factors right Perform weighting and use weighting factors For copying Perform weighting; calculate the weighted result. Compared with weighted replication The sum of these yields the overall objective. value; S8. After each update of the triple Critic network, delay the update of the Actor network, and then perform a soft update of the Actor target network and the triple Critic target network. S9. Repeat steps S3-S8 until the predetermined training period is reached or the convergence condition is met, then output the real-time electrical parameter characteristics. S10. Input the output real-time electrical parameter characteristics into the improved TD3 decision transmission line system for decision analysis to obtain the control strategy of the transmission line equipment.
Citation Information
Cited By
Method and system for predicting granularity of lithium battery material
CN121725957A