Active frequency control method and device of virtual synchronous generator
By dynamically adjusting the rotational inertia and damping coefficient of the virtual synchronous generator using deep reinforcement learning algorithms, the problem of fixed parameters of the virtual synchronous generator in grid frequency regulation is solved, thereby improving the stability and regulation speed of the grid.
Patent Information
- Application Number
- CN202511240523.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-18
AI Technical Summary
Existing virtual synchronous generators have fixed parameters in grid frequency regulation, which makes it difficult to cope with complex and ever-changing grid conditions, resulting in problems such as frequency overshoot and excessively long regulation time.
A deep reinforcement learning algorithm is used to dynamically adjust the rotational inertia and damping coefficient of a virtual synchronous generator. By collecting grid data in real time, the rotational inertia and damping coefficient are optimized using a reinforcement learning agent to cope with grid power fluctuations.
It enables virtual synchronous generators to respond to grid changes in real time, improving the stability and regulation speed of the grid system and adapting to frequency stability under complex operating conditions.
Smart Images

Figure CN120978804A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power equipment, in particular to an active frequency control method and device of a virtual synchronous generator. BACKGROUND
[0002] With the increasing penetration of new energy, the inertia of the power system gradually decreases, leading to increasingly prominent grid frequency stability problems. The virtual synchronous generator (VSG) technology outputs power by simulating the inertia and damping characteristics of synchronous machines, responds to the changes in the frequency of the new energy grid-connected system, and thus maintains the stable operation of the power grid.
[0003] The existing virtual synchronous machine adopts an offline simulation method to identify the corresponding inertia and damping characteristic parameters, which are fixed and unchanged during active frequency regulation, and thus it is difficult to cope with complex and variable grid operating conditions (such as power surges and load fluctuations), which can easily cause frequency overshoot and excessively long regulation time. SUMMARY
[0004] The present application provides an active frequency control method and device of a virtual synchronous generator, which dynamically adjusts the rotational inertia J and damping coefficient D of the VSG through a deep reinforcement learning algorithm, and more timely responds to power fluctuations in the power grid.
[0005] In a first aspect, the present application provides an active frequency control method of a virtual synchronous generator, comprising:
[0006] Real-time acquisition of the frequency difference, power difference and frequency change rate of the power grid, and determination of the state vector at each time according to the frequency difference, power difference and frequency change rate at each time;
[0007] Inputting the state vector at each time into a reinforcement learning agent, wherein the reinforcement learning agent comprises an action network, and the action network outputs corresponding action parameters based on the state vector;
[0008] Optimizing the reinforcement learning agent according to the state vector and action parameters at each time to obtain an optimized reinforcement learning agent, and determining the action parameters corresponding to the current state vector through the optimized reinforcement learning agent, wherein the action parameters are the rotational inertia and the damping coefficient;
[0009] Determining the rotational inertia and the damping coefficient, and determining the active frequency control according to the rotor motion equation and the primary frequency characteristics of the virtual synchronous generator.
[0010] According to the active frequency control method of the virtual synchronous generator in the embodiment, reinforcement learning is combined with virtual synchronous generator control technology, and a reinforcement learning agent can determine the moment of inertia and the damping coefficient in different states of the power grid in real time, so as to dynamically adjust the moment of inertia and the damping coefficient of the virtual synchronous generator, determine the active frequency in real time, and make the virtual synchronous generator respond to changes in the power grid in real time, thereby helping to maintain the stability of the power grid system.
[0011] In the above scheme, the rotor motion equation of the virtual synchronous generator is:
[0012]
[0013] The primary frequency characteristic is:
[0014] P m =P ref +k p (ω0-ω)
[0015] According to the rotor motion equation and the primary frequency characteristic of the virtual synchronous generator, the active frequency control is determined as:
[0016]
[0017] wherein J is the moment of inertia; D is the damping coefficient; ω is the synchronous angular velocity, ω0 is the angular velocity reference value, T m is the mechanical torque; T e is the electromagnetic torque; T d is the damping torque, P e is the electromagnetic power, P ref is the active power reference value, P m is the mechanical power, and k p is the droop coefficient.
[0018] In the above scheme, the reinforcement learning agent is optimized according to the state vector and the action parameter at each moment, including:
[0019] The reward value at each moment is calculated according to the state vector and the action parameter at each moment, and when the reward value meets a preset condition, the state vector, the action parameter and the reward value at a target moment, and the state vector at a next moment are extracted as experience tuples, and the target moment is the moment when the reward value meets the preset condition;
[0020] The reinforcement learning agent is optimized through multiple experience tuples.
[0021] In the above scheme, the reinforcement learning agent further includes an evaluation network, and the reinforcement learning agent is optimized through multiple experience tuples, including:
[0022] obtaining a Q value output by the evaluation network for the experience tuple, and calculating a target Q value of the experience tuple by a target network corresponding to the evaluation network;
[0023] calculating a time difference error between the Q value and the target Q value of the experience tuple; wherein the evaluation network and the corresponding target network are initially identical;
[0024] updating the evaluation network based on the time difference error, and synchronously updating the target network.
[0025] The above scheme further comprises:
[0026] calculating a policy gradient of the action parameter output by the action network for a plurality of experience tuples, updating the action network by the policy gradient, and synchronously updating a target network corresponding to the action network, wherein the action network and the corresponding target network are initially identical.
[0027] The above scheme further comprises:
[0028] updating the parameters of the target networks corresponding to the action network and the evaluation network respectively according to a preset soft update coefficient.
[0029] In the above scheme, the reward value at each time is calculated according to the state vector and the action parameter at each time, comprising:
[0030] calculating the reward value corresponding to the state vector and the action parameter by a reward function;
[0031] The reward function is:
[0032]
[0033] wherein k1, k2, k3 are respectively the penalty coefficients of the frequency difference, the power difference and the frequency change rate, k4 is the penalty coefficient of the action parameter change, J is the moment of inertia, D is the damping coefficient, and t is the time.
[0034] In a second aspect, the present application provides an active frequency control device of a virtual synchronous generator, comprising:
[0035] a state acquisition module, configured to acquire the frequency difference, the power difference and the frequency change rate of the power grid in real time, and determine the state vector at each time according to the frequency difference, the power difference and the frequency change rate at each time;
[0036] a parameter output module, configured to input the state vector at each time into a reinforcement learning intelligent agent, wherein the reinforcement learning intelligent agent comprises an action network, and the action network outputs the corresponding action parameter based on the state vector;
[0037] an online optimization module, configured to optimize the reinforcement learning agent according to the state vector and the action parameter at each time, to obtain an optimized reinforcement learning agent, and determine the action parameter corresponding to the current state vector through the optimized reinforcement learning agent, the action parameter being the moment of inertia and the damping coefficient;
[0038] a frequency control module, configured to determine the moment of inertia and the damping coefficient, and determine the active frequency control according to a rotor motion equation of the virtual synchronous generator and a primary frequency characteristic.
[0039] In a third aspect, the present application provides an electronic device, comprising a memory and one or more processors. The memory stores one or more computer programs, and the computer programs comprise instructions which, when executed by the processor, cause the electronic device to perform the active frequency control method of the virtual synchronous generator in the first aspect.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are run on an electronic device, the electronic device performs the active frequency control method of the virtual synchronous generator in the first aspect.
[0041] In a fifth aspect, the present application provides a computer program product, which, when run on an electronic device, causes the electronic device to perform the active frequency control method of the virtual synchronous generator in the first aspect.
[0042] It can be understood that the active frequency control device of the virtual synchronous generator, the electronic device, the computer-readable storage medium, and the computer program product provided in the above embodiments can achieve the beneficial effects as described in the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 a flowchart of the active frequency control method of the virtual synchronous generator provided in the embodiments of the present application;
[0044] Figure 2 a frequency control structure diagram in the active frequency control method of the virtual synchronous generator provided in the embodiments of the present application;
[0045] Figure 3 a structural diagram of the active frequency control device of the virtual synchronous generator provided in the embodiments of the present application;
[0046] Figure 4 a structural diagram of the electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION
[0047] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", and the like are used to distinguish the same or similar items with basically the same function and role. For example, the first chip and the second chip are merely used to distinguish different chips, and do not limit the sequence. Those skilled in the art can understand that the terms "first", "second", and the like do not limit the quantity and execution sequence, and the terms "first", "second", and the like do not necessarily mean different. It should be noted that in the embodiments of the present application, the words "exemplary" or "for example" are used to represent an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words "exemplary" or "for example" are intended to present the relevant concept in a specific manner. In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more.
[0048] It should be noted that "at" in the embodiments of the present application can be at the moment when a certain condition occurs, or within a period of time after a certain condition occurs, which is not limited in the embodiments of the present application.
[0049] Next, some technical terms related to the present application will be explained:
[0050] Virtual synchronous generator (VSG): Through power electronic converters, the external characteristics (such as inertia and damping) of synchronous generators are simulated to realize the frequency and voltage support functions of grid-connected devices.
[0051] Deep Deterministic Policy Gradient (DDPG): A reinforcement learning algorithm combining deep neural networks and deterministic policy, suitable for optimization problems in continuous action space.
[0052] Moment of inertia (J): A parameter representing the inertia of the rotor of the VSG simulated synchronous generator, which affects the rate of change of system frequency.
[0053] Damping coefficient (D): A parameter representing the damping effect of the VSG simulated synchronous generator, which affects the decay rate of system oscillation.
[0054] The implementation of the embodiments will be described in detail below with reference to the accompanying drawings.
[0055] The embodiment provides an active frequency control method of a virtual synchronous generator. The active frequency control method of the virtual synchronous generator can be applied to various electronic devices such as a computer (PC), a tablet computer, a virtual reality / augmented reality device, a wearable device, an industrial computer, a car machine, and the like; and can also be applied to a server, a cloud, a server cluster, and the like. The embodiment is not specially limited in this regard.
[0056] Figure 1 A flowchart of the active frequency control method of the virtual synchronous generator is shown.
[0057] As shown in the figure, the active frequency control method of the virtual synchronous generator can include the following steps: Figure 1
[0058] Step 101: Real-time collection of frequency difference, power difference and frequency change rate of the power grid, and determination of state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment.
[0059] Specifically, the frequency difference, power difference and frequency change rate at each moment are normalized, and the normalized values are converted into a vector to obtain the corresponding state vector. For example, the frequency difference Δω, the power difference ΔP and the frequency change rate dω / dt of the power grid are collected in real time, and are normalized into a state vector:
[0060] s t =[Δω norm ,ΔP norm ,(dω / dt) norm ]
[0061] The normalization process can be represented as:
[0062]
[0063] After normalization, the numerical range of each parameter is [-1, 1], which can avoid the influence of dimensional difference on the convergence of the algorithm.
[0064] Step 102: Inputting the state vector at each moment into a reinforcement learning agent, wherein the reinforcement learning agent includes an action network, and the action network outputs corresponding action parameters based on the state vector.
[0065] The reinforcement learning agent includes an action network and an evaluation network, the action network outputs corresponding action parameters based on the state vector, and the evaluation network determines the Q value of the state vector and the action parameters.
[0066] The reinforcement learning agent refers to a reinforcement learning model using DDPG decision-making. The reinforcement learning agent includes an actor network and a critic network. The actor network and the critic network are online networks, and the online network further includes a corresponding target network. When the model is initially constructed, the online network and the target network are the same, that is, the actor network and the corresponding target network thereof are the same, and the critic network and the corresponding target network thereof are the same. The online network is used to process the online collected power grid data, that is, the state vector, and perform online optimization according to the processing result.
[0067] Exemplarily, the state vector s t The input actor network to obtain the continuous action a t =[J, D], and the action generation formula is:
[0068] a t = μ(s t | θ μ ) + N t (2)
[0069] Wherein μ(·) is an actor policy function, θ μ is an actor network parameter, and N t is an exploration noise (such as an OU noise). The adjustment range of the action parameter is constrained to J: [0.05, 5] and D: [0, 70].
[0070] The critic network determines the Q value according to the state vector and the action parameter, and is used to evaluate the income of the action output by the actor network.
[0071] The actor network is equivalent to an “intelligent decision maker”, which generates an adjustment strategy of J and D according to the current power grid state (Δω, ΔP, dω / dt) in real time, and the output action affects the electromechanical characteristics of the VSG.
[0072] The critic network is equivalent to a “strategy evaluator”, which evaluates the long-term income of the current action through the Q value function, and guides the actor optimization strategy.
[0073] Step 103: optimizing the reinforcement learning agent according to the state vector and the action parameter at each time to obtain an optimized reinforcement learning agent, determining the action parameter corresponding to the current state vector through the optimized reinforcement learning agent, and the action parameter is the moment of inertia and the damping coefficient.
[0074] The reinforcement learning agent is optimized according to the state vector and the action parameter at each moment, to obtain an optimized reinforcement learning agent, including: calculating the reward value at each moment according to the state vector and the action parameter at each moment, when the reward value meets a preset condition, extracting the state vector, the action parameter and the next moment state vector at a target moment as an experience tuple, the target moment is the moment when the reward value meets the preset condition, and the reinforcement learning agent is optimized through multiple experience tuples.
[0075] The reward value r corresponding to the state vector and the action parameter is calculated by a reward function t ; for example, the reward function is:
[0076]
[0077] Wherein, k1, k2, k3 are respectively the penalty coefficients of the frequency difference, the power difference and the frequency change rate, k4 is the penalty coefficient of the action parameter change, J is the moment of inertia, D is the damping coefficient, and t is the time.
[0078] When the reward value at a moment t meets a preset condition, for example, is greater than a preset threshold, then the moment t is the target moment, and the state vector s t , the action parameter a t and the state vector s t+1 at the next moment are taken as a set of experience values, that is, an experience tuple. For example, the experience tuple can also include the reward value at the moment t, that is, the experience tuple is represented as: (s t , a t , r t , s t+1 ). The experience tuple can be saved in an experience replay pool, and used for optimizing the reinforcement learning agent.
[0079] The process of obtaining the frequency difference, the power difference and the frequency change rate is real-time, and the corresponding action parameter, that is, the moment of inertia and the damping coefficient, is determined by the real-time state vector, the virtual synchronous generator is adjusted, and the accumulated data (that is, the experience tuple) is used to update the reinforcement learning agent, and the action at the next time is determined by the updated reinforcement learning agent.
[0080] Specifically, a Q value output by the evaluation network for the experience tuple is obtained, a target Q value of the experience tuple is calculated by a target network corresponding to the evaluation network, a time difference error between the Q value and the target Q value of the experience tuple is calculated, the evaluation network and the corresponding target network are initially identical, the evaluation network is updated based on the time difference error, and the target network is synchronously updated. The specific way of updating the target network is that the parameters of the target network are updated according to a preset soft update coefficient. For the action network, a policy gradient of action parameters output by the action network for multiple experience tuples is calculated, the action network is updated through the policy gradient, and a target network corresponding to the action network is synchronously updated, the action network and the corresponding target network are initially identical. The method of updating the target network corresponding to the action network is the same as that of updating the target network corresponding to the evaluation network.
[0081] The target of the Critic network update is to minimize the error (TD error) between the predicted Q value and the target Q value. The time difference error is used as the loss function, and the expression is as follows:
[0082]
[0083] wherein Q(·) is the evaluation network, θ Q is a parameter of the evaluation network; y i is the target Q value, s i is the i th state vector, a i is the i th action parameter; wherein,
[0084] y i = r i + γQ'(s i+1 , μ'(s i+1 θ μ′ )| θ Q′ (5)
[0085] wherein r i is the i th reward value, γ is a discount factor, Q'(·) is a target network of the evaluation network, θ Q′ is a target network parameter of the evaluation network, μ'(·) is a target network of the action network, θ μ′ is a target network parameter of the action network.
[0086] In the embodiment, the parameter gradient is determined by minimizing the time difference error, and the parameters of the evaluation network are updated by gradient descent. Minimizing the time difference error means that the target Q value calculated is substituted into formula (4) to solve the parameter gradient of the evaluation network that minimizes the time difference error. The update process is as follows: when the amount of data in the experience pool is sufficient, a batch of data is randomly sampled Update the network, N is the batch size (number of samples). Sample from the experience tuples, get the sampled batch, calculate the target Q value through the sampled batch, minimize L(θ Q ) to determine the parameter gradient, update the evaluation network parameter θ Q′ using gradient descent.
[0087] The goal of actor network update is to maximize the Q value evaluated by the critic, and the parameters are updated in the direction of the policy gradient, and the policy gradient is:
[0088]
[0089] In the formula, is the policy gradient, is the gradient of the critic output to the action, is the gradient of the actor output to the parameter.
[0090] When updating the actor network, the gradient of the critic network output needs to be determined, so the update process is: fix the critic parameter θ Q , calculate the policy gradient based on the critic parameter, and update the parameters of the actor network based on the policy gradient using gradient ascent.
[0091] The target network corresponding to the actor network and the critic network is also updated synchronously every time it is updated. The update method of the target network is as follows:
[0092] θ Q′ ←τθ Q +(1-τ)θ Q′ ,θ μ′ ←τθ μ +(1-τ)θ μ′ (7)
[0093] Where τ is a soft update coefficient, τ << 1, to ensure training stability. θ Q is the updated online critic corresponding target network parameter, θ μ is the updated online actor network corresponding target network parameter, θ Q′ is the critic network corresponding target network parameter before updating, θ μ′ is the actor network corresponding target network parameter before updating.
[0094] The parameters of the online network and the target network are initially the same. When the target network is updated, part of the parameters (the proportion is τ) are retained, and the remaining parameters (i.e. 1-τ) are randomly updated to form the parameters of the updated target network. That is, the network parameter θ Q′ of the target network corresponding to the evaluation network after updating is τθQ and (1-t) theta Q′ , respectively, the action network updates the network parameters theta of the target network μ′ by t theta μ and (1-t) theta μ′ .
[0095] Step 104: Determine the moment of inertia and damping coefficient, and determine the active frequency control according to the rotor motion equation and primary frequency modulation characteristic of the virtual synchronous generator.
[0096] The moment of inertia and damping coefficient of the virtual synchronous generator are adjusted to the values output by the optimized reinforcement learning agent, so as to more accurately respond to power grid fluctuations and adjust the stability of the power grid system.
[0097] The VSG control simulates the external characteristics of the synchronous generator, so that the converter has voltage and frequency regulation functions, the system has damping and inertia support, and the damping and inertia can be adjusted. The VSG control mainly includes active frequency control, reactive voltage control and virtual impedance control.
[0098] The active frequency control mainly considers the rotor motion equation and primary frequency modulation characteristic of the synchronous generator.
[0099] The rotor motion equation of the synchronous generator is:
[0100]
[0101] where J is the moment of inertia, D is the damping coefficient, omega is the synchronous angular velocity, omega0 is the angular velocity reference value, T m is the mechanical torque, T e is the electromagnetic torque, T d is the damping torque, P e is the electromagnetic power, and P m is the mechanical power.
[0102] The frequency modulation characteristic of the synchronous generator is:
[0103] P m = P ref + k p (omega0-omega) (9)
[0104] where P ref is the active power reference value, and k p is the droop coefficient.
[0105] Substituting equation (9) into equation (8) gives:
[0106]
[0107] By performing a Laplace transform on equation (10) to "digitally generate zeros", the active frequency control of the virtual synchronous generator can be obtained:
[0108]
[0109] Once the moment of inertia and damping coefficient are determined, the active frequency control of the virtual synchronous generator can be determined based on its rotor motion equations and primary frequency regulation characteristics. The active frequency control structure diagram of this virtual synchronous generator is shown below. Figure 2 As shown.
[0110] For example, when power fluctuations occur in the power grid (such as a sudden load increase), the frequency deviation Δω and the rate of change dω / dt increase. The Critic network feeds back a penalty signal through the reward function, driving the Actor to adjust J and D: increasing the moment of inertia J to slow down the rate of frequency change and suppress overshoot; and increasing the damping coefficient D to accelerate oscillation decay and shorten the settling time. Under steady-state conditions, unnecessary parameter adjustments are suppressed through the parameter change penalty term (k4) in the reward function, maintaining system stability.
[0111] This implementation method acquires grid data in real time, the DDPG agent outputs action parameters, executes the parameters, optimizes J and D in real time, and further optimizes the DDPG agent through reward feedback, forming a closed-loop control chain of "perception-decision-execution-feedback". This enables real-time updates to the policy network, allowing for more timely and effective responses to complex grid conditions, such as load surges and fluctuations in renewable energy processing. Compared to adjusting parameters using static preset rules, this implementation method uses deep reinforcement learning algorithms for real-time optimization, overcoming the limitations of static adjustment and providing efficient and adaptive support for the frequency stability of the power coefficient. Especially for complex scenarios such as high-proportion renewable energy grid-connected systems and isolated microgrids, it can promptly capture the dynamic coupling characteristics of the grid, improving the response speed and accuracy to grid changes.
[0112] Furthermore, this embodiment also provides an active frequency control device for a virtual synchronous generator, which can be used to execute the above-described active frequency control method for a virtual synchronous generator.
[0113] like Figure 3As shown, the active frequency control device 300 of the virtual synchronous generator specifically comprises: a state acquisition module 301, configured to acquire a frequency difference, a power difference and a frequency change rate of a power grid in real time, and determine a state vector at each moment according to the frequency difference, the power difference and the frequency change rate at each moment; a parameter output module 302, configured to input the state vector at each moment into a reinforcement learning agent, the reinforcement learning agent comprising an action network, the action network outputting corresponding action parameters based on the state vector; an online optimization module 303, configured to optimize the reinforcement learning agent according to the state vector at each moment and the action parameters, to obtain an optimized reinforcement learning agent, and determine the action parameters corresponding to the current state vector through the optimized reinforcement learning agent, the action parameters being moment of inertia and damping coefficient; and a frequency control module 304, configured to determine the moment of inertia and the damping coefficient, and determine the active frequency control according to a rotor motion equation of the virtual synchronous generator and a primary frequency characteristic.
[0114] The specific details of each module or unit in the active frequency control device of the virtual synchronous generator have been described in detail in the corresponding active frequency control method of the virtual synchronous generator, and thus will not be described here again.
[0115] The embodiment of the application further provides an electronic device, Figure 3 A structural schematic diagram of an electronic device suitable for implementing the embodiment of the disclosure is shown. Figure 3 The electronic device 600 shown is only an example, and should not bring any limitation to the functions and use range of the embodiment of the disclosure.
[0116] As Figure 3 As shown, the electronic device 600 comprises a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded into a random access memory (RAM) 603 from a storage portion 608. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, the ROM 602 and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0117] The following components are connected to the I / O interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as necessary. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage part 608 as necessary.
[0118] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-described functions defined in the embodiments of the present application are executed.
[0119] For example, when the computer program is executed by the central processing unit (CPU) 601, the following can be performed: real-time acquisition of a frequency difference, a power difference, and a frequency change rate of a power grid, and determination of a state vector at each time according to the frequency difference, the power difference, and the frequency change rate at each time; input of the state vector at each time into a reinforcement learning agent, the reinforcement learning agent including an action network, the action network outputting a corresponding action parameter based on the state vector; optimization of the reinforcement learning agent according to the state vector at each time and the action parameter, to obtain an optimized reinforcement learning agent, determination of an action parameter corresponding to a current state vector by the optimized reinforcement learning agent, the action parameter being a moment of inertia and a damping coefficient; determination of the moment of inertia and the damping coefficient, and determination of an active frequency control according to a rotor motion equation of a virtual synchronous generator and a primary frequency characteristic.
[0120] Note that the computer-readable medium shown in the disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the disclosure, the computer-readable signal medium can include a data signal that propagates in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including, but not limited to, wireless, wireline, optical fiber, RF, etc., or any suitable combination of the above.
[0121] The flow diagrams and block diagrams in the drawings are illustrations of possible architectures, functions, and operations of systems, methods, and computer program products in accordance with various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.
[0122] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0123] As another aspect, the present application also provides a computer readable medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The computer readable medium carries one or more programs, which include instructions that, when executed by the electronic device, cause the electronic device to implement the method described in the above embodiments.
[0124] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.
[0125] The above is merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto, any changes or replacements within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of active frequency control of a virtual synchronous generator, characterized in that, The method comprises the steps of: real-time acquisition of frequency difference, power difference and frequency change rate of the power grid, and determination of a state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment; inputting the state vector at each moment into a reinforcement learning agent, wherein the reinforcement learning agent comprises an action network, and the action network outputs corresponding action parameters based on the state vector; optimization of the reinforcement learning agent according to the state vector and the action parameters at each moment to obtain an optimized reinforcement learning agent, and determination of the action parameters corresponding to the current state vector by the optimized reinforcement learning agent, wherein the action parameters are moment of inertia and damping coefficient; determination of the moment of inertia and the damping coefficient, and determination of the active frequency control according to the rotor motion equation and the primary frequency modulation characteristic of the virtual synchronous generator.
2. The method of active frequency control of a virtual synchronous generator according to claim 1, characterized in that, The rotor motion equation of the virtual synchronous generator is: The primary frequency modulation characteristic is: P m = P ref + k p (ω0- ω) The active frequency control is determined according to the rotor motion equation and the primary frequency modulation characteristic of the virtual synchronous generator. where J is the moment of inertia; D is the damping coefficient; ω is the synchronous angular velocity, ω0 is the angular velocity reference value, T m mechanical torque; T e is the electromagnetic torque; T d is the damping torque, P e is the electromagnetic power, P ref is the active power reference value, P m is the mechanical power, k p is the droop coefficient.
3. The method of active frequency control of a virtual synchronous generator according to claim 1, characterized in that, The optimization of the reinforcement learning agent according to the state vector and the action parameters at each moment comprises: calculation of a reward value at each moment according to the state vector and the action parameters, extraction of a state vector, action parameters and reward value at a target moment and a next moment state vector as an experience tuple when the reward value meets a preset condition, and the target moment is the moment when the reward value meets the preset condition; optimization of the reinforcement learning agent by the multiple experience tuples.
4. The method for active frequency control of a virtual synchronous generator according to claim 3, characterized in that, The reinforcement learning agent further comprises an evaluation network, and the optimization of the reinforcement learning agent by the multiple experience tuples comprises: obtaining a Q value output by the evaluation network for the experience tuple, and calculating a target Q value of the experience tuple by a target network corresponding to the evaluation network; calculation of a time difference error between the Q value of the experience tuple and the target Q value; wherein the evaluation network and the corresponding target network are initially the same; updating the evaluation network based on the time difference error, and synchronously updating the target network.
5. The method of active frequency control of a virtual synchronous generator according to claim 4, characterized in that, Further comprising: calculation of a policy gradient of the action parameters output by the action network for the multiple experience tuples, updating of the action network by the policy gradient, and synchronous updating of a target network corresponding to the action network, wherein the action network and the corresponding target network are initially the same.
6. The method of active frequency control of a virtual synchronous generator according to claim 5, characterized in that, Further comprising: updating of parameters of the target networks corresponding to the action network and the evaluation network respectively according to a preset soft update coefficient.
7. The method of active frequency control of a virtual synchronous generator according to claim 1, characterized in that, The calculation of the reward value at each moment according to the state vector and the action parameters comprises: calculation of the reward value corresponding to the state vector and the action parameters by a reward function; The reward function is: wherein k1, k2 and k3 are penalty coefficients of the frequency difference, the power difference and the frequency change rate respectively, k4 is a penalty coefficient of the action parameter change, J is the moment of inertia, D is the damping coefficient, and t is the time.
8. An active frequency control device of a virtual synchronous generator, characterized in that, The method comprises the steps of: a state acquisition module for real-time acquisition of frequency difference, power difference and frequency change rate of the power grid, and determination of a state vector at each moment according to the frequency difference, power difference and frequency change rate at each moment; a parameter output module, configured to input a state vector at each moment into a reinforcement learning agent, the reinforcement learning agent comprising an action network, the action network outputting corresponding action parameters based on the state vector; an online optimization module, configured to optimize the reinforcement learning agent according to the state vector and the action parameters at each moment, to obtain an optimized reinforcement learning agent, and to determine the action parameters corresponding to the current state vector through the optimized reinforcement learning agent, the action parameters being the moment of inertia and the damping coefficient; a frequency control module, configured to determine the moment of inertia and the damping coefficient, and to determine active frequency control according to a rotor motion equation of the virtual synchronous generator and a primary frequency characteristic.
9. An electronic device, comprising: An electronic device comprising a processor and a memory, the memory storing one or more computer programs comprising instructions which, when executed by the electronic device, cause the electronic device to perform the active frequency control method of the virtual synchronous generator according to any one of claims 1-7.
10. A computer program product, characterised in that, The computer program product, when running on the electronic device, causes the electronic device to perform the active frequency control method of the virtual synchronous generator according to any one of claims 1-7.