Method and device for determining parameters of a virtual synchronous generator

By using deep reinforcement learning algorithms to optimize the rotational inertia and damping coefficient of a virtual synchronous generator in real time, the problem that parameter adjustment in existing technologies cannot adapt to complex grid operating conditions is solved, thereby improving the frequency stability of the power grid.

CN120749807BActive Publication Date: 2025-11-07YUNNAN POWER GRID CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511240518.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-07
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing virtual synchronous generator (VSG) parameter adjustment methods cannot adapt to complex and ever-changing power grid conditions, leading to frequency stability issues, especially in the case of high renewable energy penetration rates, making it difficult to cope with power surges and load fluctuations.

Method used

The rotational inertia and damping coefficient of the VSG are adjusted in real time by using a deep reinforcement learning algorithm. By collecting the grid frequency difference, power difference and frequency change rate in real time, the parameters are optimized by using a reinforcement learning agent action network and evaluation network to form a closed-loop control and achieve dynamic adjustment.

Benefits of technology

It improves the accuracy and response speed of virtual synchronous generator parameters, enabling timely responses to complex grid conditions and enhancing grid frequency stability and response accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120749807B_ABST
    Figure CN120749807B_ABST
Patent Text Reader

Abstract

The application relates to a virtual synchronous generator parameter determination method and device, and relates to the field of electric power equipment; the moment of inertia and the damping coefficient can be dynamically determined, and the accuracy of the parameters is improved. The virtual synchronous generator parameter determination method comprises the following steps: collecting the frequency difference, the power difference and the frequency change rate of an electric network in real time, and determining the state vector at each moment to input into a reinforcement learning intelligent agent; the reinforcement learning intelligent agent comprises an action network and an evaluation network; the action network outputs corresponding action parameters, and the evaluation network outputs Q values; the reward value is calculated according to the state vector and the action parameters; when the reward value meets a preset condition, an experience tuple is extracted; the target Q value of the experience tuple is obtained, the time difference error of the Q value and the target Q value is calculated, and the evaluation network is updated; the strategy gradient of the action parameters is calculated, and the action network is updated; the action parameters corresponding to the current state vector are determined through the optimized reinforcement learning intelligent agent; and the action parameters are the moment of inertia and the damping coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power equipment, in particular to a virtual synchronous generator parameter determination method and device. BACKGROUND

[0002] With the increasing penetration of new energy, the inertia of the power system gradually decreases, leading to increasingly prominent power grid frequency stability problems. The virtual synchronous generator (VSG) technology, by simulating the inertia and damping characteristics of synchronous machines, has become an important means to improve the stability of new energy grid-connected systems.

[0003] The existing VSG parameter adjustment method is mainly the fixed parameter method, which determines the moment of inertia J and the damping coefficient D through offline simulation. For example, patent CN202410953321.7 "Rotational Inertia and Damping Coefficient Adaptive Control Technology of Light Storage VSG System" proposes to dynamically adjust J and D through a PID controller. Its technical solution includes: selecting a power-frequency curve simulated in a fixed parameter control mode, dividing the oscillation period into four intervals, ①t1-t2, ②t2-t3, ③t3-t4, ④t4-t5, analyzing the parameter changes in each interval, and adjusting the moment of inertia and the damping coefficient according to the analysis results and the strategy of adjusting the control parameters. However, power grid disturbances are random and diverse, such as load sudden changes and intermittent new energy output. The pre-set fixed intervals cannot cover all dynamic scenarios, which may cause parameter adjustment lag or failure. Moreover, the adjustment rule based on the arctan function and fixed threshold can only handle disturbances within a specific range, and cannot adapt to complex nonlinear coupled systems. Therefore, the VSG control using fixed parameters J and D is difficult to cope with complex and variable power grid conditions (such as power surges and load fluctuations), which may cause frequency overshoot and long adjustment time. SUMMARY

[0004] The present application provides a virtual synchronous generator parameter determination method and device, which dynamically adjusts the moment of inertia J and the damping coefficient D of the VSG through a deep reinforcement learning algorithm, making the parameters more accurate.

[0005] In a first aspect, the present application provides a virtual synchronous generator parameter determination method, comprising:

[0006] Real-time acquisition of the frequency difference, power difference and frequency change rate of the power grid, and determination of the state vector at each time according to the frequency difference, power difference and frequency change rate at each time;

[0007] Inputting the state vector at each time into a reinforcement learning agent, the reinforcement learning agent comprising an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network outputting Q values based on the state vector and the action parameters;

[0008] The reward value of each moment is calculated according to the state vector and the action parameter of each moment, and when the reward value meets a preset condition, the state vector, the action parameter and the reward value of a target moment and the state vector of a next moment are extracted as an experience tuple, the target moment being the moment when the reward value meets the preset condition;

[0009] The target Q value of the experience tuple is obtained through the evaluation network corresponding to the target network, a time difference error between the Q value of the experience tuple and the target Q value is calculated, and the evaluation network is updated based on the time difference error; wherein the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated;

[0010] The policy gradient of the action parameter output by the action network for the experience tuple is calculated, and the action network is updated through the policy gradient to obtain an optimized reinforcement learning agent;

[0011] The action parameter corresponding to the current state vector is determined by the optimized reinforcement learning agent, and the action parameter is the moment of inertia and the damping coefficient.

[0012] According to the virtual synchronous generator parameter determination method in the embodiment, reinforcement learning is combined with virtual synchronous generator control technology, and the moment of inertia and the damping coefficient under different states of the power grid can be determined in real time by the reinforcement learning agent, so that the parameters are more accurate and truly reflect the characteristics of the generator, which is beneficial to dynamically adjust the moment of inertia and the damping coefficient of the virtual synchronous generator and improve the response accuracy of the virtual synchronous generator.

[0013] In the above scheme, the reward value of each moment is calculated according to the state vector and the action parameter of each moment, comprising:

[0014] The reward value corresponding to the state vector and the action parameter is calculated through a reward function;

[0015] The reward function is:

[0016]

[0017] Wherein, k1, k2, k3 are respectively the penalty coefficients of the frequency difference, the power difference and the frequency change rate, k4 is the penalty coefficient of the action parameter change, J is the moment of inertia, D is the damping coefficient, and t is the time.

[0018] In the above scheme, the state vector of each moment is determined according to the frequency difference, the power difference and the frequency change rate of each moment, comprising:

[0019] The frequency difference, power difference and frequency change rate at each time are normalized, and the normalized values are converted into vectors to obtain corresponding state vectors.

[0020] The above scheme further includes:

[0021] According to the preset soft update coefficient, the parameters of the target network corresponding to the evaluation network and the action network are updated respectively; the action network and the corresponding target network are initially the same.

[0022] In the above scheme, the time difference error is:

[0023]

[0024] Wherein is the evaluation network, is the parameter of the evaluation network; is the target Q value, is the i-th state vector, is the i-th action parameter; wherein,

[0025]

[0026] Wherein, is the i-th reward value, and gamma is the discount factor, is the target network corresponding to the evaluation network, is the parameter of the target network corresponding to the evaluation network, is the target network corresponding to the action network, is the parameter of the target network corresponding to the action network.

[0027] The above scheme further includes:

[0028] The parameter gradient is determined by minimizing the time difference error, and the parameters of the evaluation network are updated by using the gradient descent method based on the parameter gradient;

[0029] The parameters of the action network are updated by using the gradient ascent method based on the policy gradient.

[0030] The above scheme further includes: adjusting the parameters of the virtual synchronous generator by the moment of inertia and the damping coefficient.

[0031] In a second aspect, the application provides a virtual synchronous generator parameter determination device, comprising:

[0032] A state acquisition module is configured to acquire the frequency difference, power difference and frequency change rate of the power grid in real time, and determine the state vector at each time according to the frequency difference, power difference and frequency change rate at each time.

[0033] The parameter output module is used to input the state vector at each time step into the reinforcement learning agent. The reinforcement learning agent includes an action network and an evaluation network. The action network outputs corresponding action parameters based on the state vector, and the evaluation network determines the Q-value of the state vector and the action parameters.

[0034] The reward calculation module is used to extract the state vector, action parameters, and reward value at the target time, as well as the state vector at the next time, as an experience tuple when the reward value meets the preset conditions. The target time is the time when the reward value meets the preset conditions.

[0035] The evaluation network online optimization module is used to obtain the target Q value of the empirical tuple through the target network corresponding to the evaluation network, calculate the time difference error between the Q value of the empirical tuple and the target Q value, and update the evaluation network based on the time difference error; wherein, the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated;

[0036] An online optimization module for the action network is used to calculate the policy gradient of the action network for the action parameters output by the experience tuple, and to update the action network using the policy gradient to obtain an optimized reinforcement learning agent.

[0037] The parameter determination module is used to determine the action parameters corresponding to the current state vector through the optimized reinforcement learning agent. The action parameters are the moment of inertia and the damping coefficient.

[0038] Thirdly, this application provides an electronic device including a memory and one or more processors. The memory stores one or more computer programs, each including instructions that, when executed by the processor, cause the electronic device to perform the virtual synchronous generator parameter determination method as described in the first aspect.

[0039] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the virtual synchronous generator parameter determination method as described in the first aspect.

[0040] Fifthly, this application provides a computer program product that, when run on an electronic device, causes the electronic device to execute the virtual synchronous generator parameter determination method as described in the first aspect.

[0041] Understandably, the beneficial effects that the virtual synchronous generator parameter determination device, electronic device, computer-readable storage medium, and computer program product provided above can be referred to the beneficial effects in the first aspect, and will not be repeated here. Attached Figure Description

[0042] Figure 1 A flowchart of a virtual synchronous generator parameter determination method provided in an embodiment of the present application is shown in FIG. 1.

[0043] Figure 2 A control principle structure diagram of a virtual synchronous generator in the virtual synchronous generator parameter determination method provided in an embodiment of the present application is shown in FIG. 2.

[0044] Figure 3 A structure diagram of a virtual synchronous generator parameter determination apparatus provided in an embodiment of the present application is shown in FIG. 3.

[0045] Figure 4 A structure diagram of an electronic device provided in an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION

[0046] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, the same items or similar items with basically the same functions and effects are distinguished by using “first”, “second”, etc. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit the sequence. Those skilled in the art can understand that “first”, “second”, etc. do not limit the quantity and execution sequence, and “first”, “second”, etc. also do not necessarily mean different. It should be noted that in the embodiments of the present application, “exemplary” or “for example” means to take an example, illustration or description. Any embodiment or design scheme described as “exemplary” or “for example” in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. In fact, the use of “exemplary” or “for example” is intended to present the relevant concept in a specific manner. In the embodiments of the present application, “at least one” means one or more, and “multiple” means two or more.

[0047] It should be noted that “at the time of” in the embodiments of the present application can be at the moment when a certain condition occurs, or can be within a period of time after a certain condition occurs, which is not limited in the embodiments of the present application.

[0048] Next, some technical terms related to the present application are explained:

[0049] Virtual synchronous generator (VSG): through power electronic converters, the external characteristics (such as inertia, damping) of synchronous generators are simulated to realize the frequency and voltage support functions of grid-connected devices.

[0050] Deep Deterministic Policy Gradient (DDPG): A reinforcement learning algorithm that combines deep neural networks with deterministic policies, suitable for optimization problems in continuous action spaces.

[0051] Moment of inertia (J): A parameter representing the inertia of the VSG-simulated synchronous generator rotor, affecting the rate of change of system frequency.

[0052] Damping coefficient (D): A parameter representing the damping effect of the VSG-simulated synchronous generator, affecting the decay rate of system oscillations.

[0053] The implementation of the present embodiment will be described in detail below with reference to the accompanying drawings.

[0054] The present embodiment provides a virtual synchronous generator parameter determination method. Exemplarily, the virtual synchronous generator parameter determination method can be applied to various electronic devices such as computers (PCs), tablets, virtual reality / augmented reality devices, wearable devices, industrial computers, car computers, etc.; it can also be applied to servers, clouds, server clusters, etc., and the present embodiment does not make special limitations on this.

[0055] Figure 1 A flowchart of the virtual synchronous generator parameter determination method provided by the present embodiment is shown.

[0056] As shown in Figure 1 , the virtual synchronous generator parameter determination method can include the following steps:

[0057] Step 101: Real-time acquisition of the frequency difference, power difference, and frequency change rate of the power grid, and determination of the state vector at each time point according to the frequency difference, power difference, and frequency change rate at each time point.

[0058] Specifically, the frequency difference, power difference, and frequency change rate at each time point are normalized, and the normalized values are converted into vectors to obtain the corresponding state vectors. For example, the real-time acquisition of the frequency difference , power difference , and frequency change rate of the power grid is normalized into a state vector:

[0059]

[0060] The normalization process can be represented as:

[0061] (1)

[0062] After normalization, the numerical range of each parameter is [-1, 1], which can avoid the influence of dimensional differences on the convergence of the algorithm.

[0063] Step 102: input the state vector of each time into a reinforcement learning agent, the reinforcement learning agent comprising an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network determining the Q value of the state vector and action parameters.

[0064] The reinforcement learning agent refers to a reinforcement learning model using DDPG decision-making. The reinforcement learning agent comprises an actor network and a critic network. The actor network and the critic network are online networks, and the reinforcement learning agent further comprises corresponding target networks. When the model is initially constructed, the online network and the target network are the same, i.e., the actor network and the corresponding target actor network are the same, and the critic network and the corresponding target critic network are the same. The online network is used to process the online collected power grid data, i.e., the state vector, and perform online optimization according to the processing result.

[0065] Exemplarily, the state vector s t is input into the actor network to obtain the continuous action a generated by the network output, and the action generation formula is:

[0066] (2)

[0067] wherein is an actor policy function, is an actor network parameter, is an exploration noise (such as an OU noise). The adjustment range of the action parameter is constrained .

[0068] The critic network determines the Q value according to the state vector and the action parameter, and is used to evaluate the revenue of the action output by the actor network.

[0069] The actor network is equivalent to an “intelligent decision maker”, which generates the adjustment strategy of J and D in real time according to the current power grid state (s , , ), and the output action affects the electromechanical characteristics of the VSG.

[0070] The critic network is equivalent to a “strategy evaluator”, which evaluates the long-term revenue of the current action through the Q value function to guide the optimization strategy of the actor.

[0071] Step 103: calculate the reward value of each time according to the state vector and the action parameter of each time, extract the state vector, the action parameter and the reward value of the target time, and the state vector of the next time as an experience tuple when the reward value meets a preset condition, wherein the target time is the time when the reward value meets the preset condition.

[0072] The reward value corresponding to the state vector and the action parameter is calculated by a reward function ; for example, the reward function is:

[0073] (3)

[0074] wherein, k 1, k 2, k 3 are the penalty coefficients of the frequency difference, the power difference and the frequency change rate respectively, k 4 is the penalty coefficient of the action parameter change, J is the moment of inertia, D is the damping coefficient, t is the time.

[0075] When the reward value of a certain moment t satisfies a preset condition, for example, is greater than a preset threshold, the moment t is the target moment, and the state vector t of the moment s t, The action parameter a t and the state vector of the next moment s t+1 are taken as a set of experience values, i.e. an experience tuple. For example, the experience tuple can also include the reward value of the moment t , i.e. the experience tuple is represented as: s t , a t , r t , s t+1 The experience tuple can be saved into an experience replay pool for optimizing the reinforcement learning agent.

[0076] Step 104: obtaining the target Q value of the experience tuple by the evaluation network corresponding to the target network, calculating the time difference error between the Q value of the experience tuple and the target Q value, and updating the evaluation network based on the time difference error; wherein the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated.

[0077] Step 105: calculating the policy gradient of the action parameter output by the action network for the experience tuple, updating the action network by the policy gradient, and obtaining the optimized reinforcement learning agent.

[0078] The process of acquiring the power grid frequency difference, power difference and frequency change rate is carried out in real time. The application determines the corresponding action parameters, i.e. the moment of inertia and the damping coefficient, through the real-time state vector to adjust the virtual synchronous generator, and updates the reinforcement learning agent through the accumulated data (i.e. the experience tuple), and determines the action of the next time through the updated reinforcement learning agent.

[0079] Specifically, the target network corresponding to the evaluation network is used to calculate the target Q value of the experience tuple, and the time difference error between the Q value of the experience tuple and the target Q value is calculated; wherein the target network is initially the same as the evaluation network and the action network; the evaluation network is updated based on the time difference error, and the target network is updated synchronously. The specific way to update the target network is: updating the parameters of the target network according to the preset soft update coefficient. For the action network, the policy gradient of the action network is calculated, and the action network is updated through the policy gradient.

[0080] Wherein, the goal of Critic network update is to minimize the error (TD error) between the predicted Q value and the target Q value. The time difference error is used as the loss function, and the expression is:

[0081] (4)

[0082] Wherein is the evaluation network, is the parameter of the evaluation network; is the target Q value, is the i-th state vector, is the i-th action parameter; wherein,

[0083] (5)

[0084] Wherein, is the i-th reward value, and γ is the discount factor, is the target network corresponding to the evaluation network, is the parameter of the target network corresponding to the evaluation network, is the target network corresponding to the action network, is the target network parameter corresponding to the action network.

[0085] In this embodiment, the parameter gradient is determined by minimizing the time difference error, and the parameters of the evaluation network are updated by gradient descent. Minimizing the time difference error means that the target Q value calculated is substituted into formula (4) to solve the parameter gradient of the evaluation network that minimizes the time difference error. The update process is as follows: when the amount of experience pool data is sufficient, a batch of data is randomly sampled Update the network, N is the batch size (number of samples). Sample from the experience tuples, get the sampled batch, calculate the target Q value through the sampled batch, minimize L θ Q Determine the parameter gradient, update the evaluation network parameters using gradient descent θ Q′ .

[0086] The goal of actor network update is to maximize the Q value evaluated by critic, and the parameters are updated in the direction of policy gradient, and the policy gradient is:

[0087] (6)

[0088] In the formula, is the policy gradient, is the gradient of the critic output to the action, is the gradient of the actor output to the parameter.

[0089] When updating the actor network, the gradient of the critic network output needs to be determined, so the update process is: fix the critic parameter , calculate the policy gradient according to the critic parameter, and update the parameters of the actor network based on the policy gradient using the gradient ascent method.

[0090] The target network corresponding to the actor network and the critic network is also updated synchronously every time it is updated. The update method of the target network is as follows:

[0091] (7)

[0092] Where, is the soft update coefficient, , to ensure the stability of training. is the updated online critic corresponding target network parameters, is the updated online actor network corresponding target network parameters, is the critic network corresponding target network parameters before updating, is the actor network corresponding target network parameters before updating.

[0093] The initial parameters of the online network and the target network are the same. When updating the target network, part of the parameters (the proportion is τ ) is retained, and the remaining part (i.e. 1- τ ) is randomly updated to form the parameters of the updated target network. That is, the network parameters of the target network corresponding to the evaluation network after updating is determined by and The network parameter of the target network corresponding to the action network is updated By And .

[0094] Step 106: determining the action parameter corresponding to the current state vector by the optimized reinforcement learning agent, the action parameter being the moment of inertia and the damping coefficient.

[0095] The state vector of the power grid is collected, and the corresponding action parameter is output by the action network, and the action parameter is taken as the moment of inertia and the damping coefficient of the virtual synchronous generator, so as to realize the frequency control of the virtual synchronous generator. The process can be performed in real time, so that after the parameters of the virtual synchronous generator are adjusted, the state vector of the next moment can be used to verify whether the parameters of the last moment are accurate, thereby forming a closed loop. The experience tuple corresponding to the accurate parameters is used to update and optimize the reinforcement learning model, and the updated and optimized model can be used to determine the action parameter again, realize real-time and dynamic updating of the action parameter, and make the action parameter reflect the real characteristics to the greatest extent and improve the accuracy of the action parameter.

[0096] The embodiment further includes adjusting the parameters of the virtual synchronous generator by the moment of inertia and the damping coefficient. The moment of inertia and the damping coefficient of the virtual synchronous generator are adjusted to the values output by the optimized reinforcement learning agent, so as to more accurately respond to power grid fluctuations and adjust the stability of the power grid system.

[0097] VSG control can make the converter have voltage and frequency regulation function by simulating the external characteristics of the synchronous generator. The system has damping and inertia support and can realize adjustable damping and inertia. VSG control mainly includes active frequency control, reactive voltage control and virtual impedance control.

[0098] Active frequency control mainly considers the rotor motion equation and primary frequency characteristics of the synchronous generator.

[0099] The rotor motion equation of the synchronous generator is:

[0100] (8)

[0101] wherein, J is the moment of inertia; D is the damping coefficient; ω is the synchronous angular velocity, ω 0 is the angular velocity reference value; T m mechanical torque; T e is the electromagnetic torque; T d is the damping torque; Pe is the electromagnetic power; P m is the mechanical power.

[0102] The frequency modulation characteristic of the synchronous generator is:

[0103] (9)

[0104] wherein, P ref is the active power reference value; k p is the droop coefficient.

[0105] Substituting equation (9) into equation (8) gives:

[0106] (10)

[0107] The active frequency control of the virtual synchronous generator can be obtained by "digitally zeroing" equation (10) and performing Laplace transform:

[0108] (11)

[0109] After determining the moment of inertia and the damping coefficient, the active frequency control of the virtual synchronous generator can be determined according to the rotor motion equation and the primary frequency modulation characteristic of the virtual synchronous generator, and the structure diagram of the active frequency control of the virtual synchronous generator is shown in FIG. 1. Figure 2

[0110] For example, when the power fluctuation occurs in the power grid (such as sudden load increase), the frequency deviation and the change rate increase, the Critic network feeds back the punishment signal through the reward function, and drives the Actor to adjust J and D : increasing the moment of inertia J , delaying the frequency change speed, and suppressing overshoot; increasing the damping coefficient D , accelerating the oscillation decay, and shortening the adjustment time. In the steady state condition, the unnecessary parameter adjustment is suppressed through the parameter change punishment term (k4) in the reward function, and the system stability is maintained.

[0111] The present embodiment acquires the power grid data in real time, outputs the action parameters of the DDPG agent, executes the parameters, and adjusts J and D ​Real-time optimization and reward feedback are used to refine the DDPG agent, forming a closed-loop control chain of "perception-decision-execution-feedback." This enables real-time updates to the policy network, allowing for more timely and effective responses to complex grid conditions, such as load fluctuations and new energy processing volatility. Compared to statically adjusting parameters using preset rules, this implementation utilizes deep reinforcement learning algorithms for real-time optimization, overcoming the limitations of static adjustments and providing efficient and adaptive support for frequency stability of the power factor. Especially for complex scenarios such as high-proportion renewable energy grid-connected systems and isolated microgrids, it can promptly capture the dynamic coupling characteristics of the grid, improving the response speed and accuracy to grid changes.

[0112] Furthermore, this embodiment also provides a virtual synchronous generator parameter determination device, which can be used to execute the above-described virtual synchronous generator parameter determination method.

[0113] like Figure 3 As shown, the virtual synchronous generator parameter determination device 300 specifically includes: a state acquisition module 301, used to acquire the frequency difference, power difference, and frequency change rate of the power grid in real time, and determine the state vector at each moment based on the frequency difference, power difference, and frequency change rate at each moment; a parameter output module 302, used to input the state vector at each moment into a reinforcement learning agent, the reinforcement learning agent including an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network determining the Q value of the state vector and the action parameters; and a reward calculation module 303, used to extract the state vector, action parameters, and reward value at the target moment, as well as the state vector at the next moment, as an experience tuple when the reward value meets a preset condition, the target moment being when the reward value meets the preset condition. The evaluation network online optimization module 304 is used to obtain the target Q value of the empirical tuple through the target network corresponding to the evaluation network, calculate the time difference error between the Q value of the empirical tuple and the target Q value, and update the evaluation network based on the time difference error; wherein, the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated; the action network online optimization module 305 is used to calculate the policy gradient of the action network for the action parameters output by the empirical tuple, update the action network through the policy gradient, and obtain the optimized reinforcement learning agent; the parameter determination module 306 is used to determine the action parameters corresponding to the current state vector through the optimized reinforcement learning agent, wherein the action parameters are the moment of inertia and the damping coefficient.

[0114] The specific details of each module or unit in the above-mentioned virtual synchronous generator parameter determination device have been described in detail in the corresponding virtual synchronous generator parameter determination method, so they will not be repeated here.

[0115] This application also provides an electronic device. Figure 3 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 3 The electronic device 600 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0116] like Figure 3 As shown, the electronic device 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0117] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.

[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs the functions defined in the embodiments of this application.

[0119] For example, when the computer program is executed by the central processing unit (CPU) 601, the following can be performed: real-time acquisition of the frequency difference, power difference and frequency change rate of the power grid, and determination of the state vector at each time point according to the frequency difference, power difference and frequency change rate at each time point; input of the state vector at each time point into a reinforcement learning agent, the reinforcement learning agent including an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network outputting a Q value based on the state vector and the action parameters; calculation of the reward value at each time point according to the state vector and the action parameters at each time point, extraction of the state vector, action parameters and reward value at a target time point, and the state vector at a next time point as an experience tuple when the reward value meets a preset condition, the target time point being the time point at which the reward value meets the preset condition; obtaining of a target Q value of the experience tuple by a target network corresponding to the evaluation network, calculation of a time difference error between the Q value of the experience tuple and the target Q value, and updating of the evaluation network based on the time difference error; wherein the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated; calculation of a policy gradient of the action parameters output by the action network for the experience tuple, updating of the action network by the policy gradient, and obtaining of an optimized reinforcement learning agent; and determination of the action parameters corresponding to the current state vector by the optimized reinforcement learning agent, the action parameters being the moment of inertia and the damping coefficient.

[0120] Note that the computer-readable medium shown in the disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wireless, wireline, optical fiber, RF, etc., or any suitable combination of the above.

[0121] The flow diagrams and block diagrams in the drawings are illustrations of possible architectures, functions, and operations of systems, methods, and computer program products in accordance with various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code that comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0122] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0123] As another aspect, the present application also provides a computer readable medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The computer readable medium carries one or more programs, which include instructions that, when executed by the electronic device, cause the electronic device to implement the method described in the above embodiments.

[0124] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.

[0125] The above merely describes the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of determining parameters of a virtual synchronous generator, characterized in that, The method comprises: real-time acquisition of frequency difference, power difference and frequency change rate of the power grid, and determination of a state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment; input of the state vector at each moment into a reinforcement learning agent, the reinforcement learning agent comprising an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network outputting a Q value based on the state vector and the action parameters; calculation of a reward value at each moment according to the state vector and the action parameters at each moment, extraction of a state vector, action parameters and reward value at a target moment, and a next moment state vector as an experience tuple when the reward value meets a preset condition, the target moment being the moment when the reward value meets the preset condition; the calculation of the reward value at each moment according to the state vector and the action parameters comprises: calculation of the reward value corresponding to the state vector and the action parameters by a reward function; the reward function being: wherein, k 1, k 2, k 3 are a frequency difference, a power difference, a frequency change rate penalty coefficient, respectively, k 4 is a motion parameter change penalty coefficient, J is a moment of inertia, D is a damping coefficient, t is time; acquisition of a target Q value of the experience tuple by a target network corresponding to the evaluation network, calculation of a time difference error between the Q value of the experience tuple and the target Q value, and update of the evaluation network based on the time difference error; wherein the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated; calculation of a policy gradient of the action parameters output by the action network for the experience tuple, update of the action network by the policy gradient, and obtaining of an optimized reinforcement learning agent; determination of action parameters corresponding to the current state vector by the optimized reinforcement learning agent, the action parameters being moment of inertia and damping coefficient; further comprising: update of parameters of target networks corresponding to the evaluation network and the action network respectively according to a preset soft update coefficient; the action network and the corresponding target network are initially the same; the time difference error being: wherein to evaluate a network, to evaluate parameters of a network; to target Q-values, to an i-th state vector, to an i-th action parameter; wherein, wherein, is the ith reward value, and γ is a discount factor, is a target network corresponding to the evaluation network, is a parameter of the target network corresponding to the evaluation network, is a target network corresponding to the action network, is a parameter of the target network corresponding to the action network.

2. The virtual synchronous generator parameter determination method according to claim 1, characterized in that, the determination of the state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment comprises: normalization of the frequency difference, power difference and frequency change rate at each moment, conversion of the normalized values into vectors, and obtaining of corresponding state vectors.

3. The virtual synchronous generator parameter determination method of claim 1, wherein further comprising: determination of a parameter gradient by minimizing the time difference error, and update of parameters of the evaluation network by a gradient descent method based on the parameter gradient; update of parameters of the action network by a gradient ascent method based on the policy gradient.

4. The virtual synchronous generator parameter determination method of claim 1, wherein further comprising: adjustment of parameters of the virtual synchronous generator by the moment of inertia and the damping coefficient.

5. A virtual synchronous generator parameter determination apparatus, characterized by, application to the virtual synchronous generator parameter determination method of any one of claims 1-4, comprising: a state acquisition module for real-time acquisition of frequency difference, power difference and frequency change rate of the power grid, and determination of a state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment; a parameter output module for input of the state vector at each moment into a reinforcement learning agent, the reinforcement learning agent comprising an action network and an evaluation network, the action network outputting corresponding action parameters based on the state vector, and the evaluation network determining a Q value of the state vector and the action parameters; The reward calculation module is configured to calculate a reward value of each time instant according to the state vector and the action parameter of each time instant, extract the state vector, the action parameter and the reward value of a target time instant and a next time instant state vector as an experience tuple when the reward value meets a preset condition, and the target time instant is a time instant when the reward value meets the preset condition; The reward value of each time instant is calculated according to the state vector and the action parameter of each time instant, including: calculating the reward value corresponding to the state vector and the action parameter through a reward function; the reward function is: wherein, k 1, k 2, k 3 are a frequency difference, a power difference, a frequency change rate penalty coefficient, respectively, k 4 is a motion parameter change penalty coefficient, J is a moment of inertia, D is a damping coefficient, t is time; The evaluation network online optimization module is configured to obtain a target Q value of the experience tuple through a target network corresponding to the evaluation network, calculate a time difference error between the Q value of the experience tuple and the target Q value, and update the evaluation network based on the time difference error; wherein the evaluation network and the corresponding target network are initially the same, and the target network is also updated when the evaluation network is updated; The action network online optimization module is configured to calculate a policy gradient of the action parameter output by the action network for the experience tuple, update the action network through the policy gradient, and obtain an optimized reinforcement learning intelligent agent; The parameter determination module is configured to determine the action parameter corresponding to the current state vector through the optimized reinforcement learning intelligent agent, and the action parameter is the moment of inertia and the damping coefficient.

6. An electronic device, comprising: The electronic device includes a processor and a memory, and the memory stores one or more computer programs including instructions that, when executed by the electronic device, cause the electronic device to perform the virtual synchronous generator parameter determination method of any one of claims 1-4.

7. A computer program product, characterised in that, When the computer program product is running on the electronic device, the electronic device performs the virtual synchronous generator parameter determination method of any one of claims 1-4.

Citation Information

Patent Citations

  • Self-adaptive control technology for rotational inertia and damping coefficient of optical storage VSG system

    CN119134537A

  • Virtual synchronous generator parameter adaptive control method based on DDPG algorithm

    CN115276093A

  • Virtual power system stabilizer parameter setting method

    CN115800320A