Multi-port inverter control method based on PPO reinforcement learning algorithm
Through the Actor-Critic network framework of PPO reinforcement learning algorithm, the nonlinear control problem of multi-port inverters is solved, efficient multi-objective control under model-free conditions, adaptive optimization of power distribution and voltage control, and the dynamic response and steady-state performance of the system are improved.
Patent Information
- Application Number
- CN202510491557.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
When facing nonlinear control targets, the existing multi-port inverter control methods have large power ripple, slow dynamic response, and require precise modeling, resulting in high distortion of the total harmonic distortion of the grid-side voltage and unable to achieve efficient multi-objective control.
Using the Actor-Critic network framework based on PPO reinforcement learning algorithm, the strategy network is trained through the state observer and reward function to realize multi-objective control under model-free conditions. Combining the power distribution factor and modulation wave as control variables, the optimal action is output to adjust the switching tube duty cycle of the multi-port inverter.
Without modeling, high-quality control is achieved with small port power ripple, fast dynamic response, and small total harmonic distortion of AC voltage, taking into account dynamic performance and steady-state accuracy.
Smart Images

Figure CN120357757A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-port inverters, and more specifically, relates to a control method for a multi-port inverter based on a PPO reinforcement learning algorithm. Background Art
[0002] The development and utilization of new energy is a feasible solution to global climate problems and fossil energy shortages. However, there are still some limitations in current new energy power generation systems. At present, the new energy power system has the problem of relatively poor power supply stability, mainly manifested in that when encountering extreme weather, the operation of the new energy power supply system will be affected, resulting in problems such as a decrease in the power generation of the new energy power system and a lower power transmission efficiency. The hybrid energy storage system is a feasible solution to this problem. This system uses a battery energy system as an auxiliary unit of the new energy system to smooth power fluctuations and achieve ideal energy management.
[0003] Traditional hybrid energy storage system structures, such as DC parallel architectures, series architectures, and dual-port DC-DC architectures, involve two conversions, have low efficiency, and inevitably require high-power passive filters with large volume and weight, increasing the implementation cost. Compared with traditional topologies, the single-stage multi-port structure eliminates the intermediate DC-DC converter and has advantages such as high conversion efficiency, high power density, and low hardware cost.
[0004] The design of the single-stage multi-port inverter control method faces two major challenges. One is that its control objectives include both DC port power distribution and AC side voltage control, and the two are coupled. The other is that the power distribution link has strong non-linearity. Existing methods partially use multi-loop lead controllers to gradually achieve each objective. However, linear controllers cannot handle non-linear control objectives, ultimately resulting in large port power ripples and slow dynamic responses. In addition, there are also methods using MPC for multi-objective control, but its accurate modeling based on prior knowledge is complex and the total harmonic distortion of the grid-side voltage is high. Therefore, there is an urgent need for a control method that can achieve high-quality multi-objective control without modeling. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a control method for a multi-port inverter based on a PPO reinforcement learning algorithm, which solves the technical problems of large power ripples, slow dynamic responses, high total harmonic distortion of the grid-side voltage, and the need for accurate modeling existing in the prior art by introducing an adaptive control mechanism of joint training of a policy network and a value network.
[0006] To achieve the above invention objective, a control method for a multi-port inverter based on a PPO reinforcement learning algorithm according to the present invention is characterized by including the following steps:
[0007] (1). Build a multi-port inverter and a grid-side load, and design the action space of the controller according to the control objectives of the multi-port inverter: A = [k, e refd , e refq , where k represents the power distribution factor, and e refd , e refq represent the dq-axis components obtained after the CLARK transformation of the modulation wave e ref ;
[0008] (2). Calculate the duty cycle of the switching tube based on the equivalent two-level model;
[0009] (3). Design the state observer of the controller:
[0010]
[0011] Among them, P h , are the output power and the reference output power of the high-voltage port of the multi-port inverter respectively, V h , V l are the voltages of the high-voltage port and the low-voltage port of the multi-port inverter respectively, i g is the grid-side current of the multi-port inverter, u gq , u gd are the dq-axis components obtained after the CLARK transformation of the grid-side voltage u g respectively, are the dq-axis components obtained after the CLARK transformation of the grid-side reference voltage respectively;
[0012] (4). Establish an Actor-Critic network;
[0013] The Actor-Critic network includes an Actor network and a Critic network;
[0014] The Actor network has a mean path branch and a standard deviation path branch. Among them, the mean path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a tanh layer, a scaling layer, a mean layer, and a merged output layer connected in series; the standard deviation path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a softplus layer, a standard deviation layer, and a merged output layer connected in series;
[0015] The Critic network consists of a state input layer, a fully connected layer, a normalization layer, a relu layer, a fully connected layer, a normalization layer, a relu layer, and an output layer connected in series;
[0016] (5). Design the reward function for guided training:
[0017]
[0018] Among them, ω1, ω2, and ω3 are weight parameters, and γ is a reward coefficient; min and max are respectively 1% and 5% of the reference power P * ; P h 、 are respectively the output power of the high-voltage port and the reference output power of the multi-port inverter;
[0019] (6) Train the Actor-Critic network based on the PPO reinforcement learning algorithm;
[0020] (7) Select the optimal Actor-Critic network
[0021] Suppose a total of H Actor-Critic networks are trained in step (6);
[0022] Given and the grid-side load value load, then collect the output of the multi-port inverter to form the test state S * , and input the to-be-tested state S * into H Actor networks at the same time. Obtain the action A through the Actor network, then substitute the action A * into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, control the output of the multi-port inverter through the duty cycle of the switching tube, and then select the Actor network corresponding to the minimum output power ripple and the minimum total harmonic distortion of the voltage of the multi-port inverter;
[0023] (8) Given and the grid-side load value load, then collect the output of the multi-port inverter to form the to-be-tested state S, input the to-be-tested state S into the trained Actor network, obtain the action A through the Actor network, then substitute the action A into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, thereby realizing the power control of the multi-port inverter.
[0024] A control method for a multi-port inverter based on the PPO reinforcement learning algorithm is provided. This method introduces a reinforcement learning intelligent control strategy to replace the traditional control algorithm and realizes efficient power distribution and grid-side voltage control of the multi-port inverter.
[0025] The invention object of the present invention is realized as follows:
[0026] The multi-port inverter control method based on the PPO reinforcement learning algorithm of the present invention first constructs an Actor-Critic network framework with the operating state of the multi-port inverter as the environment, the port power and the grid-side output voltage as the control objectives, and the power distribution factor and the modulation wave as the direct control variables. The PPO algorithm is used to jointly train the Actor and the Critic. The state quantity of the state observer is used as the observation input, and the optimal action for adjusting the modulation strategy of the multi-port inverter is output. Through interactive training with the environment, the adaptive optimization of the power distribution and voltage control strategies can be realized under the condition of no model.
[0027] Meanwhile, the multi-port inverter control method based on the PPO reinforcement learning algorithm of the present invention also has the following
[0028] beneficial effects:
[0029] (1) By adopting a model-free non-linear control algorithm, the technical problem of the system non-linear control of the port power regulation environment is solved, and the obtained port power ripple is small.
[0030] (2) The reinforcement learning controller based on PPO is applied to simultaneously control the DC-side power and the AC-side voltage. Without the need for modeling or adjusting the cooperation of multiple controllers, the accurate and rapid control of the port power and the AC voltage can be realized simultaneously, and the total harmonic distortion of the AC voltage is small and the dynamic response is fast.
[0031] (3) The trained policy network can be directly deployed in the control system to output the corresponding control quantity in real time, realizing the unity of multi-port power coordinated control and AC-side voltage stable control, and taking into account the dynamic performance and steady-state accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of a multi-port inverter control method based on the PPO reinforcement learning algorithm according to an embodiment of the present invention;
[0033] Figure 2 is a schematic diagram of the Actor-Critic network setting in an embodiment of the present invention;
[0034] Figure 3 is the working condition of the single-stage multi-port inverter under the port power given value in an embodiment of the present invention;
[0035] Figure 4 is the working condition of the single-stage multi-port inverter under the change of the port power given value in an embodiment of the present invention;
[0036] Figure 5 is the working condition of the single-stage multi-port inverter under the change of the AC-side total power given value in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following describes the specific implementation manners of the present invention in conjunction with the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed descriptions of known functions and designs may obscure the main content of the present invention, these descriptions will be omitted here.
[0038] Embodiment
[0039] In this embodiment, as Figure 1 shown, the multi-port inverter control method based on the PPO reinforcement learning algorithm of the present invention includes the following steps:
[0040] S1. Build a multi-port inverter and a grid-side load in Matlab / Simulink, and design the action space of the controller according to the control objectives of the multi-port inverter: A = [k, e refd , e refq , where k represents the power distribution factor, and e refd , e refq represent the dq-axis components obtained after the CLARK transformation of the modulation wave e ref ;
[0041] S2. Calculate the duty cycle of the switching tubes based on the equivalent two-level model;
[0042] S2.1. Build an equivalent two-level model of the multi-port inverter in the modulation link in Matlab / Simulink. The equivalent two-level model contains 6 non-zero vectors and 2 zero vectors, so that the space vectors are converted from 27 to 8;
[0043] S2.2. Perform the CLARK transformation on the modulation wave e ref and determine the sector, and then calculate the dwell times T1 and T2 corresponding to the two non-zero vectors in each sector, as well as the dwell time T0 of the zero vector;
[0044] S2.3. Calculate the duty cycle d x of each phase switching tube in the equivalent two-level model, where x = a, b, c represents three phases;
[0045]
[0046] Among them, T s is the switching period, and k is the zero vector adjustment parameter;
[0047] S2.4. Calculate the duty cycles d x1 , d x2 of the switching tubes S x1 , d x2 in the multi-port inverter;
[0048]
[0049] S3. Design the state observer of the controller:
[0050]
[0051] Among them, P and P * are the output power and reference output power of the upper DC port of the multi-port inverter respectively, and V h , V l are the upper DC port voltage and lower DC port voltage of the multi-port inverter respectively, and i g is the grid-side current of the multi-port inverter, and u gq , u gd are the dq-axis components of the grid-side voltage u g after CLARK transformation respectively, are the dq-axis components of the grid-side reference voltage after CLARK transformation respectively;
[0052] In this embodiment, the CLARK transformation formula is:
[0053]
[0054] S4. Establish an Actor-Critic network;
[0055] As Figure 2 shown, the Actor-Critic network includes an Actor network and a Critic network;
[0056] As Figure 2 (a) shown, the Actor network has a mean path branch and a standard deviation path branch. Among them, the mean path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a tanh layer, a scaling layer, a mean layer, and a merged output layer connected in series; the standard deviation path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a softplus layer, a standard deviation layer, and a merged output layer connected in series;
[0057] As Figure 2 (b) shown, the Critic network consists of a state input layer, a fully connected layer, a normalization layer, a relu layer, a fully connected layer, a normalization layer, a relu layer, and an output layer connected in series;
[0058] S5. Design the reward function for guided training:
[0059]
[0060] Among them, ω1, ω2, and ω3 are weight parameters, and γ is a reward coefficient; min and max are 1% and 5% of the reference power P * respectively; P h and are the output power of the high-voltage port and the reference output power of the multi-port inverter respectively;
[0061] S6. Train the Actor-Critic network based on the PPO reinforcement learning algorithm;
[0062] S6.1. Set the maximum number of PPO reinforcement learning epochs T, and each epoch contains N sampling operations; set the variable i = 0, 2,..., N - 1; initialize the Actor-Critic network; initialize the weight parameters ω1, ω2, and ω3;
[0063] S6.2. Randomly and initially give P * and the grid-side load value load, and initialize the initial state S0 of the state observer;
[0064] S6.3. Substitute the initial state S0 into Actor: π(A|S) and Critic: V(S) to obtain the action A0, the output of the Critic, and calculate the reward value R1;
[0065] Substitute the action A0 into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, thereby controlling the output of the multi-port inverter;
[0066] Randomly give P * and the grid-side load value load again, and combine the output of the multi-port inverter to obtain the next state S1,
[0067] S6.4. Execute step S6.3 according to the state S1, and so on until N sampling operations are completed;
[0068] S6.5. Calculate the average reward value of N sampling operations
[0069] S6.6. If the average reward value or the PPO reinforcement learning reaches the maximum number of epochs T, if either of them is satisfied, the training stops and jumps to step S6.10; otherwise, enter step S6.7;
[0070] S6.7. Calculate the advantage function D i and the return value G i ;
[0071]
[0072] Among them, V(s i) Represents the output of the Critic network at state S i ; δ k is the temporal difference error; b is a constant, if S i+N is the final state, then b = 0, otherwise b = 1; λ is the smoothing factor, is the discount factor;
[0073] Calculate the return G i :
[0074] G i = D i + V(S i )
[0075] S6.8, Update the Actor-Critic network;
[0076] S6.8.1, Randomly select M values from N groups of D i to calculate the loss value L actor of the Actor network;
[0077]
[0078] c i (θ) = max(min(r i (θ), 1 + ε), 1 - ε)
[0079] where π(A i |S i ; θ) represents the probability of the Actor network executing action A i at state S i with the current network parameter θ; π(A i |S i ; θ old ) represents the probability of the Actor network executing action A i at state S old with the previous round of network parameter θ i ; ε is the set clipping factor;
[0080] S6.8.2, Randomly select M values from N groups of G i to calculate the loss value L critic of the Critic network;
[0081]
[0082] S6.8.3, Use the gradient descent method to update the network parameters of the Actor network through the loss value L actor , and update the weight parameters of the Critic network using the loss value L critic ;
[0083] S6.9. Increment the number of epochs of PPO reinforcement learning by 1, and then return to step S6.2 for the next round of training;
[0084] S6.10. Update the weight parameters ω1, ω2, and ω3 in the reward function for guided training, then return to step S6.2, retrain an Actor-Critic network under the updated weight parameters, and so on until the number of networks to be trained is reached.
[0085] In this embodiment, the weight parameters ω1, ω2, and ω3 are respectively set to: 1 / 1 / 1, 2 / 1 / 1, 3 / 1 / 1. Three sets of data are used for test experiments in the training process, and a total of 3 convergent Actor-Critic networks are obtained;
[0086] S7. Select the optimal Actor-Critic network
[0087] Given and the grid-side load value load, then collect the outputs of the multi-port inverter to form the test state S * , and input the state S to be measured * into 3 Actor networks simultaneously. Obtain the action A through the Actor network, and then substitute the action A * into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube. Control the output of the multi-port inverter through the duty cycle of the switching tube, and then select the Actor network corresponding to the minimum output power ripple and the minimum total harmonic distortion of the voltage of the multi-port inverter;
[0088] S8. Given and the grid-side load value load, then collect the outputs of the multi-port inverter to form the state S to be measured. Input the state S to be measured into the trained Actor network, obtain the action A through the Actor network, and then substitute the action A into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, thereby realizing the power control of the multi-port inverter.
[0089] This embodiment is illustrated by an example, as shown in Figures 3 - 5 . Let the high-voltage port voltage be v h = 400V, and the low-voltage port voltage be v l = 200V. Figure 3 shows the steady-state performance of the motor under different port voltages, where the total power on the AC side remains P ac = 1000W. Figure 3 (a) The given output power of the high-voltage port is Figure 3 (b) The given output power of the high-voltage port is Figure 3(c) The given output power of the medium-high voltage port is P * = 1200 W. By observing the controller's output power distribution factor k, the actual output power P h of the high voltage port, the output power P l of the low voltage port, the phase-a voltage u g and phase-a current i g on the AC side, and the THD of the AC side voltage, it can be seen that under the power distribution scheme proposed in this embodiment, flexible power distribution can be achieved under different reference power requirements, and a high-quality AC side voltage output can be obtained.
[0090] As Figure 4 shown, with the total power on the AC side remaining unchanged, by changing the given value of the output power of the high voltage port, jumping from 800 W to 1200 W and then to 800 W, observe the operating state of the single-stage multi-port inverter and the output power of the upper DC power supply: Figure 4 (a) represents the situation of jumping from 800 W to 1200 W, Figure 4 (b) represents the situation of jumping from 1200 W to 800 W. It can be seen that the power of the upper DC power supply port can accurately and quickly track the change of the given value, indicating that the scheme proposed in the present invention has successfully achieved flexible port power distribution.
[0091] As Figure 5 shown, with the given value of the output power of the upper DC power supply remaining unchanged, by changing the total power on the AC side, P ac jumping from 1000 W to 1250 W and then to 1000 W, observe the operating state of the single-stage multi-port inverter and the output power of the upper DC power supply: Figure 5 (a) represents P ac jumping from 1000 W to 1250 W, Figure 5 (b) represents P ac jumping from 1250 W to 1000 W. It can be seen that the power of the upper DC power supply port can accurately maintain tracking of the given value, and the lower port makes compensation, indicating that the scheme proposed in the present invention has successfully achieved flexible port power distribution.
[0092] By observing the steady-state control performance and dynamic control performance results of the traditional scheme and the scheme proposed in the present invention, it can be seen that compared with the traditional scheme, the scheme proposed in this embodiment has smaller power ripple and smaller total harmonic distortion of the AC voltage without the need for modeling, and maintains a faster dynamic response speed, that is, it has a higher-quality multi-objective control effect.
[0093] Although the above-described illustrative embodiments of the present invention have been described to facilitate understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A control method for a multi-port inverter based on the PPO reinforcement learning algorithm, characterized in that It includes the following steps: (1) Build a multi-port inverter and a grid-side load, and design the action space of the controller according to the control objective of the multi-port inverter: A = [k, e refd , e refq , where k represents the power distribution factor, and e refd , e refq represent the dq-axis components obtained after the CLARK transformation of the modulation wave e ref . (2) Calculate the duty cycle of the switching tube based on the equivalent two-level model; (3) Design the state observer of the controller: Among them, P h , are respectively the high-voltage port output power and the reference output power of the multi-port inverter, V h , V l are respectively the high-voltage port and low-voltage port voltages of the multi-port inverter, i g is the grid-side current of the multi-port inverter, u gq , u gd are respectively the dq-axis components of the grid-side voltage u g after CLARK transformation, are respectively the dq-axis components of the grid-side reference voltage after CLARK transformation; (4) Establish an Actor-Critic network; The Actor-Critic network includes an Actor network and a Critic network; The Actor network has a mean path branch and a standard deviation path branch. The mean path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a tanh layer, a scaling layer, a mean layer, and a merged output layer connected in series. The standard deviation path branch consists of a state input layer, a fully connected layer, a relu layer, a fully connected layer, a softplus layer, a standard deviation layer, and a merged output layer connected in series; The Critic network consists of a state input layer, a fully connected layer, a normalization layer, a relu layer, a fully connected layer, a normalization layer, a relu layer, and an output layer connected in series; (5) Design the reward function for guided training: where ω1, ω2, and ω3 are weight parameters, γ is a reward coefficient; min and max are 1% and 5% of the reference power P * respectively; P h and are the output power of the high-voltage port and the reference output power of the multi-port inverter, respectively; (6) Train the Actor-Critic network based on the PPO reinforcement learning algorithm; (7) Select the optimal Actor-Critic network; Suppose a total of H Actor-Critic networks are trained in step (6); Given and the grid-side load value load, and then collect the outputs of the multi-port inverter to form the test state S * , and input the state S to be measured * into H Actor networks simultaneously. Obtain the action A through the Actor networks, and then substitute the action A * into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tubes. Control the output of the multi-port inverter through the duty cycle of the switching tubes, and then select the Actor network corresponding to the minimum output power ripple and the minimum total voltage harmonic distortion of the multi-port inverter; (8), Given and the grid-side load value load, then collect the output of the multi-port inverter to form the state S to be measured, input the state S to the trained Actor network, obtain the action A through the Actor network, and then substitute the action A into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, so as to realize the power control of the multi-port inverter.
2. The multi-port inverter control method based on the PPO reinforcement learning algorithm according to claim 1, wherein The calculation method of the duty cycle of the switching tube is as follows: (2.1) Build an equivalent two-level model for the modulation link of the multi-port inverter. The equivalent two-level model contains 6 non-zero vectors and 2 zero vectors; (2.2) Perform a CLARK transformation on the modulation wave e ref to determine the sector, and then calculate the dwell times T1 and T2 corresponding to the two non-zero vectors in each sector, as well as the zero vector dwell time T0; ( 2.3), Calculate the duty cycle d of each phase switch tube in the equivalent two-level model x , where x = a, b, c represents three phases; Among them, T s is the switching period, and k is the zero-vector adjustment parameter; (2.4) Calculate the duty cycles d of the switching transistors S x1 , S x2 in the multi-port inverter x1 , d x2 ; where V h , V l are the voltages of the high-voltage port and the low-voltage port of the multi-port inverter, respectively.
3. The multi-port inverter control method based on the PPO reinforcement learning algorithm according to claim 1, characterized in that The process of training the Actor-Critic network based on the PPO reinforcement learning algorithm is as follows: (3.1) Set the maximum number of training rounds T of PPO reinforcement learning. Each training round includes N sampling operations. Set the variable i = 0, 2,..., N - 1; Initialize the Actor-Critic network; Initialize the weight parameters ω1, ω2, ω3; (3.2), Randomly initialize and give and the grid-side load value load, and initialize the initial state S0 of the state observer; (3.3) Substitute the initial state S0 into Actor: π(A|S) and Critic: V(S) to obtain the action A0 and the output of the Critic, and calculate the reward value R1; Substitute the action A0 into the equivalent two-level model of the modulation link to obtain the duty cycle of the switching tube, so as to control the output of the multi-port inverter; Randomly given again and the grid-side load value load, and combine the output of the multi-port inverter to obtain the next state S1 (3.4) Execute step (3.3) according to the state S1, and so on until N sampling operations are completed; (3.5), Calculate the average reward value of N sampling operations (3.6) If the average reward value or the PPO reinforcement learning reaches the maximum number of training rounds T, if either condition is met, the training stops, and the trained Actor-Critic network is obtained, then jump to step (3.10); otherwise, enter step (3.7); (3.7), calculate the advantage function D after each sampling operation i and the return value G i ; Among them, V(s i ) represents the output of the Critic network in state S i ; δ k is the time difference error; b is a constant. If S i+N is the final state, then b = 0, otherwise b = 1; λ is the smoothing factor, is the discount factor; Calculate the return G i : G i = D i + V(S i ) (3.8) Update the Actor-Critic network; (3.8.1), randomly select M values from N groups of D i to calculate the loss value L of the Actor network actor ; c i (θ) = max(min(r i (θ), 1 + ε), 1 - ε) where, π(A i |S i ; θ) represents the probability that the Actor network executes action A i in state S i with the current network parameter θ; π(A i |S i ; θ old ) represents the probability that the Actor network executes action A i in state S old with the previous network parameter θ i ; ε is the set clipping factor; (3.8.2), randomly extract M values from N groups of G i to calculate the loss value L of the Critic network critic ; (3.8.3), using the gradient descent method, through the loss value L actor Update the network parameters of the Actor network, through the loss value L critic Update the weight parameters of the Critic network; (3.9) Increment the number of training rounds of PPO reinforcement learning by 1, and then return to step (6.2) for the next round of training; (3.10) Update the weight parameters ω1, ω2, ω3 in the reward function for guided training, and then return to step (3.2). Retrain an Actor-Critic network under the updated weight parameters, and so on until the required number of networks to be trained is reached.
Citation Information
Cited By
Integrated energy storage wind power generation system and control method
CN121440776A