Multi-region energy interconnection distributed automatic power generation control method, system and equipment

By using Markov decision-making process and improved dual-deep Q network DDQN algorithm in a distributed power generation control system with multi-regional energy interconnection, the problem of traditional automatic power generation control strategies being difficult to cope with random disturbances of distributed power supplies is solved, and more efficient system frequency balance and stable operation are achieved.

CN120016447APending Publication Date: 2025-05-16STATE GRID HUBEI ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093281.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional automatic power generation control strategies are difficult to effectively solve the random disturbance caused by large-scale distributed power supply and flexible load access to the power grid, which affects the safety and economic operation of the power grid.

Method used

By building a distributed power generation control system architecture with multi-regional energy interconnection, the Markov decision-making process is used to obtain the state quantity and decision reward function, and a dual experience pool and curiosity network are added to the dual-deep Q network DDQN to optimize the controller to improve the dynamic response adjustment performance of the system.

Benefits of technology

The improved DDQN algorithm can quickly obtain better dynamic optimization control solutions, improve the system's convergence speed and response speed, effectively suppress frequency fluctuations, and improve the safe and stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016447A_ABST
    Figure CN120016447A_ABST
Patent Text Reader

Abstract

The invention provides a multi-region energy interconnection distributed automatic power generation control method, system and device, and belongs to the field of automatic power generation control, and the method comprises the following steps: constructing a multi-region energy interconnection distributed power generation control system architecture; using a Markov decision process to obtain a state quantity and a decision reward function of the distributed power generation control system architecture; according to the method, the double experience pools are added in the double deep Q network DDQN, so that the current valuable experience information in the state quantity can be obtained, outdated experience information is forgotten, the convergence speed of the DDQN is improved, the curiosity network is added in the double deep Q network DDQN, the internal reward of the decision reward function can be improved, and the decision reward efficiency is improved. The convergence speed of the improved DDQN algorithm is greatly improved; a controller of each area of a power generation control system architecture is used as an intelligent agent, and an improved DDQN algorithm is used for training the intelligent agent, so that a better dynamic optimization control scheme can be quickly obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of automatic power generation control, and in particular relates to a multi-region energy interconnected distributed automatic power generation control method, system and equipment. Background Art

[0002] With the global energy transformation, the proportion of renewable energy power generation in multi-regional energy interconnection systems is also increasing. The large-scale, high-penetration intermittent distributed power sources and flexible loads connected to the power grid have led to extremely strong disturbance, randomness and intermittency in the frequency and power generation of multi-regional energy interconnection power grids, which seriously affects the safety and economic operation of the power grid. Automatic generation control (AGC), as an important part of the power system automation process, is the key to monitoring and adjusting power generation output, maintaining system frequency and interconnection line exchange power within a predetermined range. However, traditional automatic generation control strategies are difficult to solve the random disturbance problem caused by the access of large-scale distributed power sources and flexible loads to the power grid. Therefore, from the perspective of the AGC strategy, it is of great significance to find a control strategy that effectively improves the safe and stable operation of the system in the context of multi-regional energy interconnection.

[0003] The prior art proposes a PID control method based on predictive optimization to effectively deal with the uncertainty brought about by wind power access, thereby obtaining the optimal control performance of AGC. However, the PID control method often requires manual adjustment of parameters, has poor adaptability, and responds slowly when the system changes rapidly. Wang Haowei et al. proposed a multi-source automatic power generation control system and operation optimization method based on model predictive control. They built a multi-source automatic power generation control system model based on model predictive control, and used a particle swarm optimization algorithm to optimize the dynamic response regulation performance of the system in the region. However, the method has high computational complexity, and the response to uncertainty and emergencies may not be fast enough. Summary of the invention

[0004] In order to overcome the disadvantage of low control efficiency of the above-mentioned existing power generation control method, the present invention provides a multi-region energy interconnection distributed power generation control method, comprising the following steps:

[0005] Construct a distributed power generation control system architecture for multi-regional energy interconnection;

[0006] Use Markov decision process to obtain the state quantity and decision reward function of distributed generation control system architecture;

[0007] Adding dual experience pools and curiosity networks to the dual deep Q network DDQN; storing the state quantity in the dual experience pools, using the dual experience pools to adjust the weight of the transition state quantity, obtaining the current valuable experience information in the state quantity, while selectively forgetting the outdated experience information, and updating the weight of the state quantity; using the curiosity network to determine the intrinsic reward of the decision reward function, and updating the curiosity network; obtaining an improved DDQN algorithm through the updated state quantity weights and the updated curiosity network;

[0008] The controllers of each area of ​​the distributed power generation control system architecture are taken as intelligent agents, and the improved DDQN algorithm is used to train the intelligent agents to obtain optimized controllers, and the optimized controllers are used to control the power generation parameters.

[0009] Preferably, the state quantity includes a state set representing the environment, an action set representing the actions of the agent, a set representing the probability of transition of the environment state, and a reward function and a discount factor representing the reward for the agent.

[0010] Preferably, the set representing the probability of environmental state transition includes multiple state transition functions, specifically:

[0011] For the generator set that participates in both primary and secondary frequency regulation, its state transfer function is as follows:

[0012]

[0013] For a generator set that only participates in frequency regulation once, its state transfer function is as follows:

[0014] P Gi,t+1 =P Gi,t -K Gi (Δf t+1 -Δf t ),

[0015] For the generator set that only participates in secondary frequency regulation, its state transfer function is as follows:

[0016]

[0017] For controllable load, the state transfer function is as follows:

[0018]

[0019] In the formula, For the unit output change, For controllable load power adjustment, P Gi,t+1 represents the power generation of unit i at time t+1, P Gi,t represents the power generation of unit i at time t, K Gi represents the frequency modulation coefficient of unit i, Δft+1 represents the system frequency deviation at time t+1, Δf t represents the system frequency deviation at time t, represents the power demand of the ith controllable load at time t+1, represents the power demand of the i-th controllable load at time t.

[0020] Preferably, the decision reward function consists of an external reward function and an internal reward function.

[0021] Preferably, the dual experience pools include a source experience pool and a target experience pool, the source experience pool represents a data set of state quantities of the distributed power generation control system architecture at previous moments; the target experience pool represents a data set of state quantities of the distributed power generation control system architecture at current moments.

[0022] The present invention also provides a multi-region energy interconnection distributed power generation control system, comprising:

[0023] System building modules for building a multi-regional energy interconnection distributed generation control system architecture;

[0024] A state quantity acquisition module is used to obtain the state quantity and decision reward function of the distributed generation control system architecture using a Markov decision process;

[0025] An algorithm improvement module is used to add a dual experience pool and a curiosity network to a dual deep Q network DDQN; store the state quantity in the dual experience pool, use the dual experience pool to adjust the weight of the transition state quantity, obtain the current valuable experience information in the state quantity, and selectively forget the outdated experience information, and update the weight of the state quantity; use the curiosity network to determine the intrinsic reward of the decision reward function, and update the curiosity network; obtain an improved DDQN algorithm through the updated state quantity weight and the updated curiosity network;

[0026] The power generation control module is used to use the controllers of each area of ​​the distributed power generation control system architecture as intelligent agents, use the improved DDQN algorithm to train the intelligent agents, obtain optimized controllers, and use the optimized controllers to control power generation parameters.

[0027] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the multi-regional energy interconnection distributed power generation control method.

[0028] The multi-region energy interconnection distributed automatic power generation control method provided by the present invention has the following beneficial effects:

[0029] The present invention increases the convergence speed of DDQN by adding a dual experience pool in a dual-depth Q network DDQN. The dual experience pool can adjust the weight of the transition state quantity, obtain the current valuable experience information in the state quantity, selectively forget the outdated experience information, and update the weight of the state quantity. By adding a curiosity network in the dual-depth Q network DDQN, the intrinsic reward of the decision reward function of the decision reward mechanism can be improved, and the convergence speed of the improved DDQN algorithm can be greatly improved. By taking the controllers of each area of ​​the distributed power generation control system architecture as intelligent agents and using the improved DDQN algorithm to train the intelligent agents, a better dynamic optimization control solution can be quickly obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the embodiment of the present invention and its design scheme, the following briefly introduces the drawings required for this embodiment. The drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0031] Figure 1 It is a flow chart of a multi-regional energy interconnection distributed power generation control method according to an embodiment of the present invention;

[0032] Figure 2 It is a comparison chart of frequency change curve;

[0033] Figure 3 This is a comparison chart of |ACE| performance of different algorithms. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the technical solution of the present invention and implement it, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the scope of protection of the present invention.

[0035] In the description of the present invention, it is to be understood that the terms “center”, “longitudinal”, “lateral”, “length”, “width”, “thickness”, “up”, “down”, “front”, “back”, “left”, “right”, “vertical”, “horizontal”, “top”, “bottom”, “inside”, “outside”, “axial”, “radial”, “circumferential”, etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the technical solutions of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0036] In addition, the terms "first", "second", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance. In the description of the present invention, it should be noted that, unless otherwise clearly specified or limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances. In the description of the present invention, unless otherwise specified, "plurality" means two or more, which will not be described in detail here.

[0037] Example

[0038] The present invention provides a multi-region energy interconnection distributed power generation control method, specifically as follows Figure 1 As shown, the following steps are included:

[0039] Step 1: Build a multi-regional energy interconnection distributed power generation control system architecture.

[0040] The power generation control system architecture constructed by the present invention is a distributed multi-region architecture, in which wind power, photovoltaics, gas turbines, energy storage systems and controllable loads all participate in the system operation of each region, aiming to realize real-time monitoring, frequency adjustment and unit output distribution of the multi-region energy interconnection system. The architecture uses interconnection lines to connect the regions physically and informationally, and can share information and transmit power. Each region transmits real-time data to the data monitoring system and the long-term historical database, and updates the system state through the reward value of the deep reinforcement learning controller and feeds back to each region, so as to realize the unit output adjustment, interconnection line power transmission adjustment and controllable load adjustment, so that the system frequency remains basically constant, so as to realize the automatic power generation control function of the system and ensure the safe and stable operation of the system.

[0041] The regional control error ACE of each area is calculated by the following formula:

[0042] ACE=βΔf t +ΔP ij,t .

[0043] Where β is the frequency deviation coefficient; Δf t is the system frequency deviation; ΔP ij,t is the contact line deviation between regions i and j.

[0044] Step 2: Use the Markov decision process to obtain the state quantity and decision reward function of the distributed generation control system architecture.

[0045] The components of the Markov decision model are {S, A, P, R, γ}, where S represents the state set of the environment, A represents the action set of the agent (controller of each area), P represents the set of environmental state transition probabilities, R represents the reward for the agent, and γ represents the discount factor, which is used to weigh short-term rewards and long-term rewards.

[0046] 1) State Space

[0047] The state space (i.e., state set) S includes the wind turbine power generation P at time t of the system W,t , photovoltaic unit power generation power P PV,t , gas turbine power generation P GT,t , Energy storage system status ESS t , load demand P L,t , tie line power deviation ΔP T,t , system frequency deviation Δf t , area control error ACE t ; The predicted value of wind turbine power generation at time t+1 of the system P W,t+1 , the predicted value of photovoltaic power generation P PV,t+1 , gas turbine output power prediction value P GT,t+1 , energy storage system state prediction value ESS t+1 , load demand forecast value P L,t+1 , tie line power deviation prediction value ΔP Tt+1 , system frequency deviation prediction value Δf t+1 , regional control error prediction value ACE t+1 ,Right now:

[0048]

[0049] 2) Action Space

[0050] After the agent obtains state information from the environment, it will select an action in the action space (i.e., action set) according to the ε-greedy strategy. The action space A includes the output change of the AGC unit. Controllable load power adjustment Right now:

[0051]

[0052] 3) State transfer function

[0053] For the generator set that participates in both primary and secondary frequency regulation, its state transfer function is as follows:

[0054]

[0055] For a generator set that only participates in frequency regulation once, its state transfer function is as follows:

[0056] P Gi,t+1 =P Gi,t -K Gi (Δf t+1 -Δf t ).

[0057] For the generator set that only participates in secondary frequency regulation, its state transfer function is as follows:

[0058]

[0059] For controllable load, the state transfer function is as follows:

[0060]

[0061] In the formula, For the unit output change, For controllable load power adjustment, P Gi,t+1 represents the power generation of unit i at time t+1, P Gi,t represents the power generation of unit i at time t, K Gi represents the frequency modulation coefficient of unit i, Δf t+1 represents the system frequency deviation at time t+1, Δf t represents the system frequency deviation at time t, represents the power demand of the ith controllable load at time t+1, represents the power demand of the i-th controllable load at time t.

[0062] 4) Reward Function

[0063] The present invention uses the instantaneous value of regional control error ACE t (i.e., the regional control error at time t) and the absolute value of the system frequency deviation |Δf t | is used as the external reward function to obtain the optimal AGC unit output in the region and achieve system power balance under the AGC strategy. t With Δf t The dimension of is normalized to obtain the external reward function, which is expressed as follows:

[0064] r t =-λ1|ACE t |-λ2100|Δf t |.

[0065] In the formula, ACE t is the regional control error at time t; Δf t is the frequency deviation value at time t; λ1 and λ2 are both weight coefficients, and here λ1=λ2=0.5 is selected.

[0066] In order to enhance the learning effect of the intelligent agent in a complex environment, this paper introduces an intrinsic reward function The calculation formula is as follows:

[0067]

[0068] Where η is the scale factor, and η>0, is the state prediction value for the next time, s t+1 The actual value of the state at the next time.

[0069] Finally, the decision reward function is obtained, and its formula is as follows:

[0070]

[0071] In the formula, r total is the reward value, r t is the external reward function, is the intrinsic reward function.

[0072] 5) Discount Factor

[0073] The goal of deep reinforcement learning is to maximize the total reward in the entire decision cycle, and the discount factor γ is the coefficient for converting future rewards to the current moment. The choice of the discount factor usually needs to be as large as possible under the premise that the algorithm can effectively converge, so the invention sets the discount factor to 0.95.

[0074] Step 3: Add dual experience pools and curiosity networks to the dual deep Q network DDQN; store the state quantity in the dual experience pool, use the dual experience pool to adjust the weight of the transition state quantity, obtain the current valuable experience information in the state quantity, selectively forget the outdated experience information, and update the weight of the state quantity; use the curiosity network to determine the intrinsic reward of the decision reward function, and update the curiosity network; obtain the improved DDQN algorithm through the updated state quantity weight and the updated curiosity network.

[0075] The experience pool is used to store the state (data in the state set), action (data in the action set), reward value and data information of the next state in each time period. When all the data information is stored in the same experience pool, some outdated experience may interfere with the learning of the current strategy, or it may be difficult to distinguish which experience is more important during the learning process. In order to solve these problems, the present invention improves the experience playback mechanism based on the idea of ​​transfer learning, and divides the experience pool into the source experience pool e S With the target experience pool e T There are two parts, the source experience pool is used to store previous data information (ie, source experience), and the target experience pool is used to store the most recently recorded data that conforms to the current AGC (ie, target experience).

[0076] e S ={e1,e2,…,e n}.

[0077] e T ={e1,e2,…,e m}.

[0078] In the formula, n is the number of samples in the source experience pool, m is the number of samples in the target test pool, and e S ∪e T Use e total Represents the experience pool.

[0079] The present invention enhances the importance of current valuable experience by adjusting the weight of transition, while selectively forgetting outdated experience information.

[0080] When changes in environmental patterns are detected, historical experience can be filtered through transfer learning. In the initial stage, it is impossible to determine which experience is worth retaining, so the weight values ​​are randomly distributed. Weight values ​​are assigned to the experience in the source experience pool. The experience distribution weight value in the target experience pool Among them, n+ λm is the normalization factor, λ>1, and λ represents the importance coefficient of the target experience relative to the source experience.

[0081] From the experience pool total Chinese j Draw a small batch of samples, where the TD error of each sample is calculated as follows:

[0082]

[0083] In the formula, r total (s t ,a t ) represents the reward value, Indicates the next state s t+1 Next, take the best action a t+1 The expected return at time , Q(s t ,a t ) means in state s t Take action a t The estimated return is , and γ represents the discount factor.

[0084] Adjust the weights of all experience points as follows:

[0085]

[0086] Where β is a hyperparameter that controls the update rate. If a source experience has a high TD error, it can be reasonably inferred that its characteristics are different from the current environment, resulting in outdated experience, so its weight should be reduced; if a target experience has a large TD error, the experience will be highlighted.

[0087] In order to help the intelligent agent fully explore and better interact with the environment, and enable the intelligent agent to learn the optimal strategy faster, the present invention introduces a curiosity network to improve the reward mechanism of the DDQN algorithm. The curiosity network is an intrinsic reward generator, which includes a forward model for predicting the result after performing action a in state s; and a subtractor for generating intrinsic rewards.

[0088] First, the forward model inputs the current state s t and action a t To predict the state at the next time:

[0089]

[0090] Where F is the learning function; δ t is the forward model parameter, by minimizing the loss function L(δ t-1 ) to update.

[0091]

[0092] In the formula, L(δ t ) represents the loss function, δ t is the forward model parameter, σ c is the learning rate of the curiosity network, Indicates that in L(δ t ) in δ t Find the gradient, is the gradient operator.

[0093] At the same time, using the predicted value The difference between the actual value of the next state is used to calculate the intrinsic reward It is expressed as:

[0094]

[0095] Where η is the scale factor, and η>0, is the state prediction value for the next time, s t+1 The actual value of the state at the next time.

[0096] By introducing the curiosity network to drive DDQN, the total reward of DDQN is composed of two parts. One is the external reward function r provided by the environment. t, and the other is the curiosity intrinsic reward function output by the curiosity network This curiosity rewards It can effectively stimulate the curiosity of the agent and encourage it to explore the entire environment more fully, further improving the AGC control effect.

[0097] The acquisition process of the improved DDQN algorithm is as follows:

[0098] (1) Initialize the DDQN network and curiosity network, and set hyperparameters such as learning rate and discount factor;

[0099] (2) Collect environmental status s t , and select action a based on the current training network and ε-greedy strategy t ;

[0100] (3) Execute the selected action a t , and observe the next state s t+1 and instant rewards t ;

[0101] (4) t ,a t ,r t ,s t+1 ) is stored in the AGC experience pool D for updating;

[0102] (5) Probabilistically extract samples from the experience pool to calculate the TD error and update the weights;

[0103] (6) Using the Curiosity Network to Derive Intrinsic Rewards So we get the total reward value r total , Update the curiosity network;

[0104] (7) Calculate the estimated network Q value and the target network value, and update the estimated network parameters θ;

[0105] (8) Update the target network parameters by copying the estimated network parameters θ at a fixed frequency

[0106] (9) Repeat steps (2) to (8) until all time periods are trained.

[0107] The estimated network Q value in the above steps refers to the valuation of possible actions based on the current observed state at each time step, and the difference between the estimated Q value and the actual reward is reduced as the parameters are updated. The Q value is the expected return after taking a specific action in a given state, which is used to evaluate the value of state-action. The Q value represents the expected total return of taking a certain action in a certain state.

[0108] Step 4: Take the controllers of each area of ​​the distributed power generation control system architecture as intelligent agents, use the improved DDQN algorithm to train the intelligent agents, obtain optimized controllers, and use the optimized controllers to control power generation parameters.

[0109] In order to verify the control effect of the control method of the present invention, the present invention incorporates wind power, photovoltaic, energy storage system, and controllable load on the basis of the IEEE standard load frequency control system model, and introduces a step disturbance with an amplitude of 800MW. At the same time, the traditional DDQN control and the improved DDQN of the present invention are used to control the frequency of the system to evaluate the reliability and robustness of the improved DDQN algorithm in AGC control. The test results are as follows Figure 2 As shown in the figure, the curve is the fluctuation diagram of Δf curve in a certain area. Figure 2 It can be seen that the Δf peak value of the improved DDQN algorithm of the present invention in this area is significantly smaller than that of the traditional DDQN algorithm, and the fluctuation frequency is smaller, and the speed of restoring the system frequency is faster. This shows that the improved DDQN algorithm of the present invention can effectively suppress frequency fluctuations, that is, the system can effectively adjust the unit output when the load changes or there is a large interference, so as to achieve a balanced and stable system frequency.

[0110] Figure 3 The |ACE| performance comparison chart of different algorithms is shown in Figure 2. Similarly, a step disturbance with an amplitude of 800MW is introduced into the above-mentioned IEEE load frequency control system model, and the |ACE| performance test of the system is performed using PID control, DDQN control and the improved DDQN control of the present invention to evaluate the performance of the improved DDQN algorithm in processing complex and dynamic systems. The test results are shown in Figure 2. Figure 3 As shown in the figure, it can be seen that the |ACE| (average of the absolute value of ACE) of PID control, DDQN control and improved DDQN control are 23.43kW, 19.71kW and 16.28kW respectively, among which the |ACE| of the method of the present invention is the smallest, and its |ACE| is reduced by 30.52% and 21.07% compared with the PID control method and traditional DDQN control respectively. This shows that under the control of the improved DDQN algorithm of the present invention, the system is least affected by the sudden change of disturbance, and the control method has better control performance.

[0111] The present invention also provides a multi-region energy interconnection distributed power generation control system, comprising:

[0112] System building modules for building a multi-regional energy interconnection distributed generation control system architecture;

[0113] A state quantity acquisition module is used to obtain the state quantity and decision reward function of the distributed generation control system architecture using a Markov decision process;

[0114] An algorithm improvement module is used to add a dual experience pool and a curiosity network to a dual deep Q network DDQN; store the state quantity in the dual experience pool, use the dual experience pool to adjust the weight of the transition state quantity, obtain the current valuable experience information in the state quantity, and selectively forget the outdated experience information, and update the weight of the state quantity; use the curiosity network to determine the intrinsic reward of the decision reward function, and update the curiosity network; obtain an improved DDQN algorithm through the updated state quantity weight and the updated curiosity network;

[0115] The power generation control module is used to use the controllers of each area of ​​the distributed power generation control system architecture as intelligent agents, use the improved DDQN algorithm to train the intelligent agents, obtain optimized controllers, and use the optimized controllers to control power generation parameters.

[0116] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute a multi-regional energy interconnection distributed power generation control method.

[0117] The embodiments described above are only preferred specific implementation modes of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of the technical solutions that can be obviously obtained by any technician familiar with the field within the technical scope disclosed in the present invention belong to the protection scope of the present invention.

Claims

1. A multi-region energy interconnection distributed power generation control method, characterized in that: The steps include: Construct a distributed power generation control system architecture for multi-regional energy interconnection; Use Markov decision process to obtain the state quantity and decision reward function of distributed generation control system architecture; Adding dual experience pools and curiosity networks to the dual deep Q network DDQN; storing the state quantity in the dual experience pools, using the dual experience pools to adjust the weight of the transition state quantity, obtaining the current valuable experience information in the state quantity, while selectively forgetting the outdated experience information, and updating the weight of the state quantity; using the curiosity network to determine the intrinsic reward of the decision reward function, and updating the curiosity network; obtaining an improved DDQN algorithm through the updated state quantity weights and the updated curiosity network; The controllers of each area in the distributed power generation control system architecture are taken as intelligent agents, and the improved DDQN algorithm is used to train the intelligent agents to obtain optimized controllers, and the optimized controllers are used to control the power generation parameters.

2. The multi-regional energy interconnection distributed power generation control method according to claim 1 is characterized in that: The state quantity includes a state set representing the environment, an action set representing the action of the agent, a set representing the probability of transition of the environment state, and a reward function and a discount factor representing the reward for the agent.

3. The multi-regional energy interconnection distributed power generation control method according to claim 2 is characterized in that: The set representing the probability of environmental state transition includes multiple state transition functions, specifically: For the generator set that participates in both primary and secondary frequency regulation, its state transfer function is as follows: For a generator set that only participates in frequency regulation once, its state transfer function is as follows: P Gi,t+1 =P Gi,t -K Gi (Δf t+1 -Δf t ), For the generator set that only participates in secondary frequency regulation, its state transfer function is as follows: For controllable load, the state transfer function is as follows: In the formula, For the unit output change, For controllable load power adjustment, P Gi,t+1 represents the power generation of unit i at time t+1, P Gi,t represents the power generation of unit i at time t, K Gi represents the frequency modulation coefficient of unit i, Δf t+1 Represents the system frequency deviation at time t+1, Δf t represents the system frequency deviation at time t, represents the power demand of the ith controllable load at time t+1, represents the power demand of the i-th controllable load at time t.

4. The multi-regional energy interconnection distributed power generation control method according to claim 1 is characterized in that: The decision reward function is composed of an external reward function and an internal reward function.

5. The multi-regional energy interconnection distributed power generation control method according to claim 1, characterized in that: The dual experience pools include a source experience pool and a target experience pool. The source experience pool represents a data set of the state quantities of the distributed power generation control system architecture at previous moments; the target experience pool represents a data set of the state quantities of the distributed power generation control system architecture at current moments.

6. A multi-region energy interconnection distributed power generation control system, characterized in that: include: System building modules for building a multi-regional energy interconnection distributed generation control system architecture; A state quantity acquisition module is used to obtain the state quantity and decision reward function of the distributed generation control system architecture using a Markov decision process; An algorithm improvement module is used to add a dual experience pool and a curiosity network to a dual deep Q network DDQN; store the state quantity in the dual experience pool, use the dual experience pool to adjust the weight of the transition state quantity, obtain the current valuable experience information in the state quantity, and selectively forget the outdated experience information, and update the weight of the state quantity; use the curiosity network to determine the intrinsic reward of the decision reward function, and update the curiosity network; obtain an improved DDQN algorithm through the updated state quantity weight and the updated curiosity network; The power generation control module is used to use the controllers of each area of ​​the distributed power generation control system architecture as intelligent agents, use the improved DDQN algorithm to train the intelligent agents, obtain optimized controllers, and use the optimized controllers to control power generation parameters.

7. A computer device, characterized in that: It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the multi-regional energy interconnection distributed power generation control method as described in any one of claims 1-5.