Adaptive control device and method for virtual synchronous generator under network attack

By converting the control problem of virtual synchronous generators into reinforcement learning problems, and using the DDPG control algorithm, the problem of insufficient defense capabilities of virtual synchronous generators in the prior art when facing network attacks is solved, and efficient modelless control of inverters and effective defense of network attacks is achieved.

CN114928099BActive Publication Date: 2025-05-23NORTHEASTERN UNIV CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210466755.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-23
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

In the face of cyber attacks, existing virtual synchronous generators (VSGs) are difficult to effectively defend against coordinated and non-coordinated attacks, and the reliability and universality of the robust method are insufficient, making it difficult to adapt to different power system models.

Method used

The adaptive and optimized control problems of virtual synchronous generators are transformed into reinforcement learning problems, and the optimal control strategy based on DDPG control algorithm is adopted to embed it into the virtual synchronous generator controller, and the model-free performance of the inverter is achieved through a distributed reinforcement learning algorithm.

Benefits of technology

It realizes efficient control of the inverter, adapts to various power system models, improves the ability to defend against network attacks, and ensures the smooth and safe operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114928099B_ABST
    Figure CN114928099B_ABST
Patent Text Reader

Abstract

The present invention provides an adaptive control device and method for a virtual synchronous generator under network attacks, and relates to the technical field of generator control. The adaptive and optimal control problem of a virtual synchronous generator (VSG) is converted into a reinforcement learning problem, and system interferences such as collaborative and non-cooperative attacks in network attacks are considered. An optimal control strategy based on a DDPG control algorithm is designed and embedded into a virtual synchronous generator (VSG) controller. Through virtual synchronous generator simulation, the expected performance of an inverter (IBDG) is achieved in a long-term model-free manner. The present invention isolates the attacked computing unit or multiple computing units under collusion attacks against collaborative attacks and non-cooperative attacks from a public network, so as not to affect the overall parameter output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of generator control, and in particular to a device and method for adaptively controlling a virtual synchronous generator under network attacks. Background Art

[0002] The growing pressure of environmental protection makes it urgent to develop adaptive and highly permeable renewable energy. In the face of the energy crisis, the integration of energy technology and advanced information technology has been widely recognized in the world. Distributed information energy systems such as smart grids, energy Internet and integrated energy systems have developed rapidly. The main purpose is to achieve energy utilization efficiency and maximize economic development through the coordinated transformation of multiple energy sources; at the same time, adapt to the high penetration of renewable energy and integrate new energy into the traditional energy Internet, so as to achieve the important strategic goal of carbon peak and carbon neutrality. As the dominant power system in the energy system, in recent years, people have become more and more interested in using optimization-based methods to study the parameter setting of virtual synchronous generators (VSGs), in which the adjustment of parameters is driven by the optimal solution. There are mainly two methods: rule-based methods and optimization-based methods. In addition, for the coordinated attacks and coordinated attacks from the public network, a robust method based on decentralized reputation gradient management has been proposed recently to detect and eliminate coordinated attacks. As one of the core issues, the optimal solution of relevant parameters in the power system and the corresponding anti-interference capability play a very important role in maximizing energy benefits and social and economic benefits.

[0003] As the main controller for important parameters such as angular frequency, input power, system inertia and damping factor in power systems, the virtual synchronous generator (VSG) has variable rotational inertia compared to the traditional SGS controller while maintaining the original advantage of slow frequency drop. However, the existing optimization-based methods of virtual synchronous generators (VSGs) ignore the interaction between the virtual synchronous generator (VSG) and its working environment, and it is very difficult to establish a corresponding accurate model in the interconnected structure of a complex power system. Although such a model can be established under certain special circumstances, it generally has high cost, nonlinear and strong coupling properties. Therefore, it is also difficult for engineers to analyze the impact of virtual synchronous generators (VSGs) on the stability of power systems and design corresponding optimal control strategies under various uncertain system disturbances. In addition, different power systems require different models to be established for control, which limits the universality of this method and has a small scope of application. As for coping with coordinated and non-coordinated attacks from public networks, this robust method is developed based on the leader-follower algorithm and cannot be presented in a distributed manner. Secondly, the privacy-preserving aggregation algorithm used is not reliable enough, and only protects some information transmitted between followers, but not the global information provided by the leader. Again, in this algorithm, an isolation process for misbehaving units is established based on the "confidence level". However, the definition of this "confidence level" is unclear, and its value range is unknown; finally, this algorithm cannot handle scenarios where the measurement of load data is attacked. Summary of the invention

[0004] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned prior art and to provide an adaptive control device and method for a virtual synchronous generator under network attacks. The adaptive and optimal control problem of the virtual synchronous generator (VSG) is converted into a reinforcement learning problem, and system interferences such as collaborative and non-cooperative attacks in network attacks are considered. An optimal control strategy based on the DDPG control algorithm is designed and embedded into the virtual synchronous generator (VSG) controller. Through virtual synchronous generator simulation, the expected performance of the inverter (IBDG) is achieved in a long-term model-free manner.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0006] On the one hand, the present invention provides an adaptive control device for a virtual synchronous generator under network attacks, including a signal acquisition unit, a data preprocessing and conversion unit, a distributed microprocessor, a display unit and a communication unit.

[0007] The signal acquisition unit includes a three-phase voltage sensor, a three-phase current sensor and a voltage and current phase angle sensor, through which the three-phase voltage Vin, three-phase current Iin and voltage and current phase angle theta (in) signals of the power grid are collected respectively. The signal acquisition unit performs simple processing on the collected signals, calculates the rate of change of the circuit active output power and angular frequency, and transmits them to the data preprocessing and conversion unit as its input signal in combination with the angular frequency.

[0008] The data preprocessing and conversion unit includes a noise reduction processing module, a smoothing processing module and an A / D conversion module. The input signal of the data preprocessing and conversion unit is subjected to noise reduction by the noise reduction processing module, and the smoothing processing module outputs an output signal corresponding to the input signal after smoothing. These output signals are used as inputs of the A / D conversion module to complete the conversion from analog signals to digital signals, and then output to the input end of the distributed microprocessor.

[0009] The distributed microprocessor includes an Actor network module, a Critic network module, a target network module, a data storage module and a network attack detection module; the Actor network module is an actor neural network in a reinforcement learning algorithm, which mainly selects actions in a continuous action space according to the state at each iteration. Because the Actor is updated based on rounds, the learning efficiency is relatively slow. At this time, we found that a value-based algorithm can be used as a Critic to achieve single-step updates, namely the Critic network module. The Actor network neurons select behaviors based on probability, the Critic network neurons judge the scores of behaviors based on the behaviors of the Actor network neurons, and the Actor network neurons modify the probability of selecting behaviors according to the scores of the Critic network neurons. Combined with the existing deep policy gradient and Q-learning learning algorithms, the problem of convergence difficulties in the Actor-Critic model is solved; the target network module is used to update parameters one step slower than the Actor and Critic neural network modules to prevent security issues caused by single-step parameter updates. Based on the signal output by the data preprocessing and conversion unit, a distributed reinforcement learning algorithm is executed. The state signal of each iteration cycle is obtained by the information acquisition module. The corresponding action is selected in the action space according to the parameterized control function. The reward value function is calculated for both the state and action of this iteration cycle, and the parameters of the virtual synchronous generator control function are updated. The inverter control is achieved by imitating this behavior; the network attack detection module is mainly used to judge whether the data behavior of the calculation unit is abnormal. The communication unit transmits the energy mismatch and IC information to the calculation unit network in the distributed microprocessor through the sending module to calculate the reasonable limit of energy mismatch and IC. By comparing the sent data with the reasonable limit obtained in the previous iteration process, the abnormal unit sends the abnormal unit information to the network attack isolation module of the communication unit for isolation operation, and the data of the communication unit with good behavior is normally controlled; at the same time, the distributed microprocessor outputs a signal to the display unit for display after the calculation is completed. The distributed microprocessor also has the function of data storage.

[0010] The display unit will display the current system operation mode, the inverter's active output power, the angular frequency, and the rate of change of the angular frequency.

[0011] The communication unit includes a receiving module, a network attack isolation module and a sending module; the receiving module includes a decryption module and a receiving serial interface, which is used to receive the updated parameter data obtained by the distributed reinforcement learning algorithm through the receiving serial interface, and then the decryption module decrypts the data; the decrypted data is used to control the virtual synchronous generator, and the inverter is controlled by simulating the control of the virtual synchronous generator. The network attack isolation module is used to disconnect the abnormal behavior unit from other computing units and update the connection matrix; the sending module includes an encryption module and a sending serial interface, which is used to encrypt the energy mismatch and IC data of this iteration process and send the encrypted data to the distributed microprocessor through the sending serial interface, and encrypt the updated parameters and send them to the adjacent computing unit to update the parameters of the entire virtual synchronous generator. The serial interface is used to communicate with the neighboring node, thereby generating information interaction. The communication unit will send local information to the neighboring node or receive neighboring node information from the neighboring node at the time of dynamic event triggering.

[0012] On the other hand, the present invention also provides a method for adaptively controlling a virtual synchronous generator under network attacks, which is implemented by using the above-mentioned adaptive control device for a virtual synchronous generator under network attacks. The method comprises the following steps:

[0013] S1. Divide the global computing process into k computing units. The information sharing between the computing units is described by the graph G = (V, E, W), where V = {v ij |i,j=1,2,......,k} is a set of nodes representing computing units; Indicates available public links; Represents the adjacency matrix, and its specific form is:

[0014]

[0015] S2. Each inverter (IBDG) unit is equipped with a virtual synchronous generator (VSG) controller to observe the active output power, angular frequency and the derivative of the angular frequency of the inverter. Specifically, the active output power and angular frequency are collected by the signal acquisition unit and the angular frequency derivative is calculated. The active output power, angular frequency and the derivative of the angular frequency are expressed as: P out.t ,ω t ,dω t / dt, the observation set at time t is defined as s t : The three data in the observation set are processed by the data preprocessing and conversion unit, and the digital-to-analog conversion is performed and transmitted to the distributed microprocessor.

[0016] S3, the adjustment of adaptive parameters is to adjust the parameters in the control function u(·) and First, based on the observed system state s t , combined with the parameters in the control function u(·) of each actor neural network calculation unit at time t select a suitable action a at time t from the action space t : For the parameterization of the actor network function, u represents the actor network, and j represents the j-th calculation unit. The action value a t is composed of , and are the virtual inertia and damping factors respectively, and the corresponding input power is obtained by adjusting the virtual inertia and the damping factor

[0017] S4. Send s t , a t as inputs to the corresponding j-th critic neural network, design a relevant eigenvalue function for the observed values in s t , and calculate the corresponding reward value

[0018] S5. The reward value obtained from the state s t will be further calculated and defined as the cumulative reward R t : It is expressed as the reward value accumulated in the period from 0 to T, where T is the total time and γ is the attenuation coefficient

[0019] After observing s t and the action are input into the critic neural network, define and as the expectations of the cumulative reward values. Among them, is the parameterization of the critic neural network function, Q represents the critic neural network, and j represents the j-th calculation unit

[0020] S6. After the actor neural network selects relevant actions through , generate the next state s t+1 , and input it into the target network, and the update of the parameters in the target network follows the update of the parameters in the critic neural network and the actor neural network, defined as and And repeat steps S3 - S5 to obtain the expectation

[0021] S7. Calculate the relevant loss functions for updating the parameters of the critic and actor neural networks:

[0022] S8. Perform a detection process of coordinated attacks and non-coordinated attacks in defense against network attacks on each computing unit, and isolate abnormal units;

[0023] S9. For the critic and actor neural networks with good behavior, update the parameters of the critic and actor neural networks and further update them through local calculations.

[0024] S10. Parameters in the target network and Update as follows: τ is the update coefficient.

[0025] S11, will (s t ,a t ,r t ,s t+1 ) is included in the memory pool as a step in the learning process of the reinforcement learning algorithm.

[0026] S12. After a large number of training sets, the optimization learning of the virtual synchronous generator adaptive control strategy (i.e., u(·) mentioned above) is completed. The two parameters of virtual inertia and damping factor are adjusted according to different output power requirements. Based on the working principle of the basic closed-loop phase-locked loop, the existing virtual synchronous generator improved power decoupling control algorithm is used to control the grid-connected inverter, so as to control the output of the inverter through the operating behavior of the virtual synchronous generator.

[0027] The beneficial effect of adopting the above technical solution is that the adaptive control device and method of the virtual synchronous generator under network attack provided by the present invention converts the optimal parameter design of the inverter into a deep reinforcement learning task by collecting inverter data without establishing a corresponding accurate model, so that this optimization control method is suitable for various power system models, and can be applied to other energy systems after slight changes. The reward value index introduced in the method of the present invention enables the critic network to assign corresponding reward values ​​to the state return value and the corresponding action value at each moment, and selects the action at the next moment from the action space by accumulating the reward value. The adaptive control device of the virtual synchronous generator under network attack of the present invention has fast calculation speed, high precision, simple operation and wide application range. The present invention isolates the computing unit that receives the attack or multiple computing units that receive the collusion attack for coordinated attacks and non-cooperative attacks from the public network, so as not to affect the overall parameter output. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of the structure of a virtual synchronous generator adaptive control device resistant to network attacks provided by an embodiment of the present invention;

[0029] Figure 2 A circuit diagram of a DC-to-AC inverter provided in an embodiment of the present invention;

[0030] Figure 3 A schematic diagram of a signal acquisition circuit for the inverter output circuit related status provided by an embodiment of the present invention;

[0031] Figure 4 A circuit schematic diagram of a data smoothing processing and conversion unit provided in an embodiment of the present invention;

[0032] Figure 5 A schematic diagram of a circuit of a distributed microprocessor provided by an embodiment of the present invention;

[0033] Figure 6 A circuit schematic diagram of a communication unit provided by an embodiment of the present invention;

[0034] Figure 7 A circuit schematic diagram of a display unit provided by an embodiment of the present invention;

[0035] Figure 8 A general flow chart of a method for adaptively controlling a virtual synchronous generator under network attacks provided by an embodiment of the present invention;

[0036] Fig. 9 A flowchart of a process for detecting and isolating abnormal units when responding to public network attacks provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0038] like Figure 1 As shown, the virtual synchronous generator adaptive control device for network attacks provided in this embodiment includes a signal acquisition unit, a data preprocessing and conversion unit, a distributed microprocessor, a data storage unit, a display unit and a communication unit.

[0039] The inverter is the control terminal of the adaptive control device of this embodiment. Figure 2 As shown in the figure, the 12V DC voltage is converted into a 220V AC voltage signal through an error amplifier, a regulator, an oscillator, a PWM generator with dead-zone control, a low-voltage protection circuit and a short-circuit protection circuit. The updated parameters of the virtual synchronous generator obtained by the reinforcement learning algorithm are used to control the inverter by simulating the behavior mode of the virtual synchronous generator, so that its angular frequency and output active power change more smoothly and slowly, so that it can be more safely connected to the large power grid.

[0040] like Figure 3 As shown, the signal acquisition unit collects the three-phase voltage Vin, three-phase current Iin and voltage and current phase angle theta (in) signals of the power grid through a three-phase voltage sensor, a three-phase current sensor and a voltage and current phase angle sensor, performs simple processing, calculates the circuit active output power and the rate of change of the angular frequency, and sends them to the data preprocessing and conversion unit in combination with the angular frequency. The voltage and current phase angle sensor is composed of LM393 and 74HK36 and their additional circuits. The signal acquisition unit transmits the collected signal to the data preprocessing and conversion unit as its input signal.

[0041] Data preprocessing and conversion units such as Figure 4 As shown, it includes a noise reduction processing module, a smoothing processing module and an A / D conversion module. The noise reduction processing module is composed of MAX4104 and its corresponding circuits, the smoothing processing module is composed of LTC1562IG and its corresponding circuits, and the A / D conversion module uses a TLC2543 converter. The input signals (Vin, Iin, theta (in)) of the data preprocessing and conversion unit are subjected to noise reduction by the noise reduction processing module, and the smoothing processing module outputs output signals (Vout, Iout, theta (out)) corresponding to the input after smoothing. These output signals are used as the input of TLC2543, and the conversion from analog signals to digital signals is completed and then output. The output signals of the data preprocessing and conversion unit are D0~D3.

[0042] The distributed microprocessor includes an Actor network module, a Critic network module, a target network module, a data storage module and a network attack detection module; the Actor network module is an actor neural network in a reinforcement learning algorithm, which mainly selects actions in a continuous action space according to the state at each iteration. Since the Actor is updated based on rounds, the learning efficiency is relatively slow. A value-based algorithm can be used as a Critic to achieve single-step updates, namely the Critic network module. The Actor network neurons select behaviors based on probability, the Critic network neurons judge the scores of behaviors based on the behaviors of the Actor network neurons, and the Actor network neurons modify the probability of selecting behaviors according to the scores of the Critic network neurons. Combining the existing deep policy gradient and Q-learning learning algorithms, the problem of the difficulty of convergence of the Actor-Critic model is solved; the target network module is used to update parameters one step slower than the Actor and Critic neural network modules to prevent security issues caused by single-step parameter updates. The data storage module is a memory built into the distributed microprocessor, which is mainly used as a memory pool in the reinforcement learning algorithm. In this embodiment, the distributed microprocessor adopts the STC89C52RC microprocessor, such as Figure 5 As shown, it also has the function of data storage. The output signals D0~D3 after preprocessing and conversion by the data preprocessing and conversion unit are transmitted to the P3 port of the distributed microprocessor STC89C52RC. The distributed microprocessor STC89C52RC executes the distributed reinforcement learning algorithm based on this signal. The state signal of each iteration cycle obtained by the information acquisition module selects the corresponding action in the action space according to the parameterized control function, calculates the reward value function for both the state and action of this iteration cycle, and updates the parameters of the virtual synchronous generator control function. The control of the inverter is achieved by imitating this behavior. The network attack detection module is mainly used to judge whether the data behavior of the computing unit is abnormal. The communication unit transmits the energy mismatch and IC information to the computing unit network in the distributed microprocessor through the sending module to calculate the reasonable limit of energy mismatch and IC. By comparing the sent data with the reasonable limit obtained in the previous iteration process, the abnormal unit sends the abnormal information of the unit to the network attack isolation module of the communication unit for isolation operation, and the data of the communication unit with good behavior is normally controlled. At the same time, the distributed microprocessor AT89C52 will output signals P0~P7 to the display unit for display after the calculation is completed.

[0043] The display unit uses HDG12864L-6 for display, such as Figure 6The distributed microprocessor STC89C52RC sends the calculated output signals P0-P7 to the display unit, and the display unit HDG12864L-6 displays the current system operation mode, the inverter's active output power, the angular frequency, and the rate of change of the angular frequency.

[0044] The communication unit includes a receiving module, a network attack isolation module and a sending module; the receiving module includes a decryption module and a receiving serial interface, which is used to receive the updated parameter data obtained through the distributed reinforcement learning algorithm through the receiving serial interface, and then the decryption module decrypts the data; the decrypted data is used to control the virtual synchronous generator, and the inverter is controlled by simulating the control of the virtual synchronous generator. The network attack isolation module is used to disconnect the unit with abnormal behavior from other computing units and update the connection matrix; the sending module includes an encryption module and a sending serial interface, which is used to encrypt the energy mismatch and IC data of this iteration process and send the encrypted data to the distributed microprocessor through the sending serial interface, and encrypt the updated parameters and send them to the adjacent computing unit to update the parameters of the entire virtual synchronous generator. The communication unit uses serial interfaces COMPIM and SP232EEN to communicate with neighboring nodes, thereby generating information interaction, such as Figure 7 As shown. The communication unit will send local information to the neighboring node or receive neighboring node information from the neighboring node when the dynamic event is triggered. A network attack defense module is added to the communication module, which detects network attacks and transmits relevant information to the computing unit network of the distributed microprocessor for calculation, isolates abnormal units, and transmits data of well-behaved communication units to the inverter.

[0045] A method for adaptive control of a virtual synchronous generator under network attack is implemented by using the above-mentioned adaptive control device for a virtual synchronous generator under network attack. The overall flow chart is as follows: Figure 8 As shown, the specific method is described as follows.

[0046] S1. Divide the global computing process into k computing units, thereby accelerating the convergence speed. The information sharing between computer computing units is described by the graph G = (V, E, W), where V = {v ij |i,j=1,2,......,k} is a set of nodes representing computing units; Indicates available public links; Represents the adjacency matrix, and its specific form is:

[0047]

[0048] S2. In order to show the dynamic characteristics, each inverter (IBDG) unit is equipped with a virtual synchronous generator (VSG) controller to observe the active output power, angular frequency and the derivative of the angular frequency of the inverter. Specifically, the active output power and angular frequency are collected by the signal acquisition unit and the angular frequency derivative is calculated. The active output power, angular frequency and the derivative of the angular frequency are expressed as: P out.t ,ω t ,dω t / dt, the observation set at time t is defined as s t : The three data in the observation set are processed by the data preprocessing and conversion unit, and the digital-to-analog conversion is performed and transmitted to the distributed microprocessor.

[0049] S3, the adjustment of adaptive parameters is to adjust the parameters in the control function u(·) and First, based on the observed system state s t , combined with the parameters in the control function u(·) of each actor neural network computing unit at time t Select the appropriate action a at time t from the action space t : is the parameterization of the actor network function, where u represents the actor network and j represents the jth computational unit. t Depend on constitute, and are virtual inertia and damping factor respectively. The corresponding input power is obtained by adjusting the virtual inertia and damping factor.

[0050] S4, will s t 、a t As input, it is sent to the jth corresponding critic neural network. t Design relevant eigenvalue functions based on the observed values ​​and calculate the corresponding reward values.

[0051] S4.1. Design the reward value characteristic function for angular frequency, let ψ ω =|ω t -ω n | is the absolute value of the angular frequency deviation, then the angular frequency reward value characteristic function is:

[0052]

[0053] Among them, ω n is the standard angular frequency, ρ ω is a small penalty coefficient, σ ω is a large penalty coefficient, The maximum absolute value of the acceptable angular frequency deviation.

[0054] S4.2. The reward value characteristic function for the angular frequency change rate is designed as follows:

[0055]

[0056] Among them, ρ dω It is a small penalty coefficient. The specific setting of the penalty coefficient needs to be combined with the requirements of each power plant for angular frequency stability.

[0057] S4.3. For the active output power design reward value characteristic function, let ψ P =|P out.t -P ref | is the absolute value of the deviation between the current active output power and the actual required output power, where P ref is the reference active power, then the active output power reward value characteristic function is:

[0058]

[0059] Among them, ρ P is the corresponding penalty coefficient and is greater than 1.

[0060] S4.4. Based on the expected performance and the characteristic function defined above, the reward at time t is expressed as:

[0061]

[0062] Among them, b ω ,b dω ,b P >0, is the weight coefficient. Different output characteristics can be obtained by designing different weight coefficients. t The reward value obtained is defined as r t .

[0063] S5. The dynamic performance of active power and angular frequency regulation is measured by a relatively long time reward. Whether the dynamic performance will be better depends on the accumulation over a period of time.

[0064] To this end, from state s t The reward value obtained will be further calculated and defined as the cumulative reward R t : It is expressed as the reward value accumulated from 0 to T, where T is the total time and γ is the decay coefficient.

[0065] In observations t And actions After inputting into the critic neural network, define as well as is the expected cumulative reward value. is the parameterization of the critic neural network function (Q represents the critic neural network, j represents the jth computing unit).

[0066] S6, in the actor neural network through After selecting the relevant actions, the next state s is generated t+1 , and input into the target network. A separate target network is widely used in stable reinforcement algorithms, and the update of the parameters in the target network will slowly follow the parameter update in the critic neural network and the actor neural network, defined as and Repeat steps S3-S5 to get the desired

[0067] S7. Calculate the relevant loss function used to update the critic and actor neural network parameters:

[0068] S7.1. Calculate the loss function used to update the critic neural network parameters:

[0069]

[0070] in, is the loss function of the critic network for controlling parameters, Indicates that the expected sampling space for subsequent expressions is the data in memory pool D, Represents the corresponding expectation in the target space. And calculates the relevant gradient in Indicates the gradient calculation for the control function parameters in the critic network.

[0071] S7.2. Calculate the loss function used to update the actor neural network parameters:

[0072]

[0073] Where J is the approximate value function, Indicates that the sampling space for the expected solution of the subsequent expression is the state at time t, and Indicates that the gradient of the control function parameters in the actor network is calculated by the chain rule. and It is further updated through local calculations based on its own and neighbors’ information.

[0074] S8, for each computing unit, a detection process of coordinated and non-coordinated attacks in defense network attacks is performed, such as Fig. 9 shown.

[0075] S8.1: Introduce consensus variables y and λ; define λ as the first consensus agreement, λ(h+1)=Mλ(h)+ηy(h), where λ represents the column stack vector of IC; η is the convergence point coefficient, which is a sufficiently small positive constant; M is the updated consensus matrix for the first consensus agreement. y is the second consensus agreement, y(h+1)=Ny(h)-(x(h+1)-x(h)), where y is the column stack vector of the local estimated energy mismatch; x is the column stack vector of the energy output; N is the updated consensus matrix for the second consensus agreement.

[0076] S8.2. Calculate the relevant judgment information. Calculate the reasonable bounds of the relevant information by using the computing units other than the suspected attacked units.

[0077] S8.2.1, for the computation of the second consensus y: In the next iteration, a reasonable bound on the mismatch of the local estimated energy of the vth computational unit is

[0078]

[0079]

[0080] Where Y(h,R) is the detection threshold function determined by iterating h and R, R is a common parameter representing the maximum value of the ramp rate limit of various energy devices, and D + , D - Represent positive and negative value functions respectively.

[0081] S8.2.2. For the calculation of the first consensus agreement λ: In the next iteration, the rational bound on the IC of the vth computational unit is:

[0082]

[0083]

[0084] In the formula

[0085]

[0086]

[0087]

[0088]

[0089] Where l represents the physical transmission line, They represent the transmission lines when the first consensus agreement takes the upper and lower bounds, respectively; is the set of computing units connected to the data output of the vth computing unit, λl (h) represents the actual IC value obtained in the hth iteration of the unit connected by the physical transmission line of this unit, and σ is the first consensus agreement limit coefficient.

[0090] S8.3, the received local energy mismatch actual value y v (h) Compare with the reasonable bound of the local estimated energy mismatch of computing unit v; if the actual value y v (h) Satisfaction Then go to step S8.4; otherwise, the computing unit v is judged to be misbehaving, and go to step S8.5. Going to S8.5 here is to update the detection result variable, and the detected result variable needs to be transmitted to the adjacent computing unit. Isolation is an operation that needs to be performed on both sides of the transmission line, that is, the computing units at both ends of the transmission line need to change their connection matrices, so it is necessary to go to S8.5.

[0091] S8.4. The actual incremental electricity cost λ v (h) Compare with the theoretical estimated range; if Then the computing unit v is judged to have good behavior; otherwise, the computing unit v is judged to have bad behavior; If the unit does not exist, it is directly judged that the computing unit v behaves improperly.

[0092] S8.5. Update the detection result variable d of the calculation unit v according to the following formula v (h):

[0093]

[0094] The updated detection result variables are sent to the adjacent computing units. v (h), so that the neighboring units of the abnormal unit do not participate in the communication detection, but the misbehaving unit can still be identified and the isolation process can be activated.

[0095] S8.6. Isolation process under non-coordinated attack: The network attack only targets a certain computing unit.

[0096] S8.6.1. The counting rules are defined as:

[0097]

[0098] Among them, d v (h) is the detection result variable of the h-th iteration updating calculation unit v, T iv (h) is the number of well-behaved computational units after the hth iteration.

[0099] S8.6.2. The reputation value is designed as a threshold to determine whether the computing unit is isolated, as follows:

[0100]

[0101] Among them, ω h > 0 is a time-varying parameter called reputation coefficient, which is used to dynamically adjust the adaptation rate. When all units behave well, rep iv (h) will always be e; when a misbehaving unit appears, rep iv (h) < e. When the reputation values ​​of all communication lines connected to the abnormal unit fall below a threshold, the abnormal unit is disconnected from the communication network.

[0102] S8.7. Isolation process under coordinated attack: The public network launches a collusion attack on two or more computing units at the same time.

[0103] S8.7.1. After the abnormal unit with bad behavior has been isolated through the process of step S8.6, the connection information is recorded in the observation subnet and the link matrix is ​​updated as follows:

[0104]

[0105] Among them, f iv (h) represents the value of the connection matrix between the ith computing unit and the vth computing unit after the hth iteration process, E' is the edge set of the observed network topology, is the set of computing units connected to the data output of the vth computing unit, N v - The set of computational units connected to the data input of the vth computational unit.

[0106] S8.7.2, link matrix F(h) = [f ij ] n×n If a non-zero element appears in the vth row of the link matrix, it means that the abnormal computing unit v is not completely isolated, that is, there is a collusion attack in the communication network. The cth computing unit adjacent to the vth computing unit is easily judged as a collusion unit of the computing unit v and is considered to be a misbehaving computing unit. Then the computing unit c repeats the detection process in steps S8.2-S8.5, and uploads the detection results for the relevant isolation process in step S8.6.

[0107] S8.7.3. Repeat steps S8.7.1-S8.7.2 until all abnormal units are isolated.

[0108] S9. For the critic and actor neural networks with good behavior, update the parameters of the critic and actor neural networks and further update them through local calculations.

[0109] S9.1. Update the relevant parameters of the critic neural network: And calculate the average value of the parameters of a total of k computing units: where ζ Q is the learning rate of the critic neural network.

[0110] S9.2, update the actor neural network related parameters: And calculate the average value of the parameters of a total of k computing units: where ζ u is the learning rate of the actor neural network.

[0111] S10. Parameters in the target network and To update: τ is the update coefficient, which is determined by the update efficiency required in the specific practical application. In summary, a reinforcement learning process is completed.

[0112] S11, will (s t ,a t ,r t ,s t+1 ) is included in the memory pool as a step in the learning process of the reinforcement learning algorithm.

[0113] S12. After a large number of training sets, the optimization and learning of the virtual synchronous generator adaptive control strategy (i.e., u(·) mentioned above) is completed. The two parameters of virtual inertia and damping factor can be adjusted according to different output power requirements. Based on the working principle of the basic closed-loop phase-locked loop, the existing virtual synchronous generator improved power decoupling control algorithm is used to control the grid-connected inverter, reduce power oscillation and impact current, and control the output of the inverter through the operating behavior of the virtual synchronous generator.

[0114] The present invention transforms the control algorithm of the virtual synchronous generator into a reinforcement learning problem, introduces two sets of neural networks, observes the continuous state of the inverter (IBDG), and rewards the continuous action of the controller. In a data-driven manner, the design of the corresponding parameters is adjusted with the real-time observed data, so that the inverter can obtain the expected performance in a model-free manner. When using data-driven, the corresponding reward value index is introduced, and the corresponding reward value is fed back for the control strategy of each state and the observed corresponding state value, and the control action of the virtual synchronous generator controller at the next moment is adjusted according to the reward value index. The data-driven virtual synchronous generator adaptive optimization control strategy based on the distributed deep reinforcement learning algorithm is embedded in the virtual synchronous generator (VSG) controller, and assembled into a corresponding virtual synchronous generator adaptive control device facing network attacks. By topologically configuring the communication network of the computing unit, and by designing the reputation value and the link matrix, defense is carried out against the coordinated attack and non-cooperative attack from the public network to ensure the smooth and safe operation of the system.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. An adaptive control device for virtual synchronous generators under network attacks, Features: The device includes a signal acquisition unit, a data preprocessing and conversion unit, a distributed microprocessor, a display unit and a communication unit; The signal acquisition unit includes a three-phase voltage sensor, a three-phase current sensor and a voltage and current phase angle sensor, through which the three-phase voltage Vin, three-phase current Iin and voltage and current phase angle theta in signals of the power grid are respectively collected; the signal acquisition unit performs simple processing on the collected signals, calculates the rate of change of the active output power of the circuit and the angular frequency, and transmits them to the data preprocessing and conversion unit as its input signal in combination with the angular frequency; Data preprocessing and conversion unit, including noise reduction processing module, smoothing processing module and A / D conversion module; The input signal of the data preprocessing and conversion unit is subjected to noise reduction by the noise reduction processing module, and the smoothing processing module outputs an output signal corresponding to the input signal after smoothing. These output signals are used as inputs of the A / D conversion module to complete the conversion of analog signals to digital signals, and then output to the input end of the distributed microprocessor; The distributed microprocessor includes an actor network module, a critic network module, a target network module, a data storage module and a network attack detection module; the actor network module is an actor neural network in a distributed reinforcement learning algorithm, which selects actions in a continuous action space according to the state at each iteration; the critic network module is used to implement single-step updates as a critic through a value-based algorithm; the actor network neurons select behaviors based on probability, the critic network neurons judge the scores of behaviors based on the behaviors of the actor network neurons, and the actor network neurons modify the probability of selecting behaviors according to the scores of the critic network neurons; the target network module updates parameters one step slower than the actor and critic neural network modules to prevent security issues caused by single-step parameter updates; through information collection The state signal of each iteration cycle obtained by the collection module is selected in the action space according to the parameterized control function, and the reward value function is calculated for both the state and action of this iteration cycle, and the parameters of the virtual synchronous generator control function are updated, and the control of the inverter is achieved by imitating this behavior; the network attack detection module is mainly used to judge whether the data behavior of the computing unit is abnormal, and the communication unit transmits the energy mismatch and IC information to the computing unit network in the distributed microprocessor through the sending module to calculate the reasonable limit of energy mismatch and IC, and compares the sent data with the reasonable limit obtained in the previous iteration process, and sends the abnormal unit abnormal information to the network attack isolation module of the communication unit for isolation operation, and the data of the communication unit with good behavior is normally controlled; at the same time, the distributed microprocessor outputs a signal to the display unit for display after the calculation is completed; The data storage module is used to store data during the calculation process; The display unit is used to display the current system operation mode, the active output power of the inverter, the angular frequency and the rate of change of the angular frequency; The communication unit uses a serial interface to communicate with neighbor nodes, thereby generating information interaction, and sends local information to neighbor nodes or receives neighbor node information from neighbor nodes when a dynamic event is triggered; the communication unit includes a receiving module, a network attack isolation module and a sending module; the receiving module includes a decryption module and a receiving serial interface, which is used to receive the updated parameter data obtained through the distributed reinforcement learning algorithm through the receiving serial interface, and then the decryption module decrypts the data; the decrypted data is used to control the virtual synchronous generator, and the inverter is controlled by simulating the control of the virtual synchronous generator; the network attack isolation module is used to disconnect the unit with abnormal behavior from other computing units and update the connection matrix; the sending module includes an encryption module and a sending serial interface, which is used to encrypt the energy mismatch and IC data of this iterative process and then send the encrypted data to the distributed microprocessor through the sending serial interface, and encrypt the updated parameters and send them to the adjacent computing units to update the parameters of the entire virtual synchronous generator.

2. The adaptive control device for virtual synchronous generators under network attacks according to claim 1, Features: The voltage and current phase angle sensor includes LM393 and 74HK36 and additional circuits thereof; The noise reduction processing module includes MAX4104 and its corresponding additional circuits, the smoothing processing module includes LTC1562IG and its corresponding additional circuits, and the A / D conversion module uses a TLC2543 converter; the output signal of the data preprocessing and conversion unit is D0~D3; The distributed microprocessor adopts the STC89C52RC microprocessor; the output signals D0-D3 preprocessed and converted by the data preprocessing and conversion unit are transmitted to the P3 port of the distributed microprocessor STC89C52RC, and the distributed microprocessor STC89C52RC executes the distributed reinforcement learning algorithm based on the signal; after the calculation is completed, the distributed microprocessor AT89C52 will output signals P0-P7 to the display unit for display; The display unit uses HDG12864L-6 for display; The communication unit uses serial interfaces COMPIM and SP232EEN to communicate with neighboring nodes.

3. An adaptive control method for virtual synchronous generators under network attacks, Features: The method is implemented by using the virtual synchronous generator adaptive control device for network attacks as described in claim 1, and comprises the following steps: S1. Divide the global computing process into k computing units. The information sharing between the computing units is described by the graph G = (V, E, W), where V = {v ij |i,j=1,2,......,k} is a set of nodes representing computing units; Indicates available public links; Represents the adjacency matrix, and its specific form is: S2. Each inverter unit is equipped with a virtual synchronous generator controller to observe the active output power, angular frequency, and the derivative of the angular frequency of the inverter. Specifically, the signal acquisition unit collects the active output power and angular frequency and calculates the derivative of the angular frequency. The active output power, angular frequency, and the derivative of the angular frequency are respectively represented as: P out.t , ω t , dω t / dt. The observation set at time t is defined as s t : The data preprocessing and conversion unit processes the three data within the observation set and performs digital-to-analog conversion and transmits them into the distributed microprocessor; S3, the adjustment of adaptive parameters is to adjust the parameters in the control function u(·) and First, based on the observed system state s t , combined with the parameters in the control function u(·) of each actor neural network computing unit at time t Select the appropriate action a at time t from the action space t : is the parameterization of the actor network function, where u represents the actor network, j represents the jth computing unit, and the action value a t Depend on constitute, and are virtual inertia and damping factor respectively, and the corresponding input power is obtained by adjusting the virtual inertia and damping factor; S4, will s t 、a t As input, it is sent to the jth corresponding critic neural network. t Design relevant eigenvalue functions based on the observed values ​​and calculate the corresponding reward values; S5, from state s t The reward value obtained will be further calculated and defined as the cumulative reward R t : It is expressed as the reward value accumulated during the period from 0 to T, where T is the total time and γ is the decay coefficient; During observation s t and action After being input into the critic neural network, define and as the expectation of the cumulative reward value; where is the parameterization of the critic neural network function, Q represents the critic neural network, and j represents the j-th computing unit; S6, in the actor neural network through After selecting the relevant actions, the next state s is generated t+1 , and input into the target network, and the update of the parameters in the target network follows the update of the parameters in the critic neural network and the actor neural network, defined as and Repeat steps S3-S5 to get the desired S7, calculate the relevant loss function for updating the critic and actor neural network parameters; S8. Perform a detection process of coordinated attacks and non-coordinated attacks in defense against network attacks on each computing unit, and isolate abnormal units; S9, for the critic and actor neural networks with good behavior, update the parameters of the critic and actor neural networks, and further update them through local calculations; S10. Parameters in the target network and Update as follows: τ is the update coefficient; S11, will (s t ,a t ,r t ,s t+1 ) is included in the memory pool as a learning process of the one-step reinforcement learning algorithm; S12. After a large number of training sets, the optimization learning of the virtual synchronous generator adaptive control strategy is completed. The two parameters of virtual inertia and damping factor are adjusted according to different output power requirements. Based on the working principle of the basic closed-loop phase-locked loop, the existing virtual synchronous generator improved power decoupling control algorithm is used to control the grid-connected inverter, so as to control the output of the inverter through the operating behavior of the virtual synchronous generator.

4. The method for adaptive control of a virtual synchronous generator under network attacks according to claim 3, Features: The specific method of step S4 is: S4.

1. Design the reward value characteristic function for angular frequency, let ψ ω =|ω t -ω n | is the absolute value of the angular frequency deviation, then the angular frequency reward value characteristic function is: Among them, ω n is the standard angular frequency, ρ ω is the small penalty coefficient, σ ω is the large penalty coefficient, is the maximum value of the absolute value of the acceptable angular frequency deviation; S4.

2. The reward value characteristic function for the angular frequency change rate is designed as follows: Among them, ρ dω is a small penalty coefficient. The specific penalty coefficient should be set in combination with the requirements of each power plant for angular frequency stability. S4.

3. For the active output power design reward value characteristic function, let ψ P =|P out.t -P ref | is the absolute value of the deviation between the current active output power and the actual required output power, where P ref is the reference active power, then the active output power reward value characteristic function is: C(P out.t )=e ρPψP Among them, ρ P is the corresponding penalty coefficient and is greater than 1; S4.

4. Based on the expected performance and the characteristic function defined above, the reward at time t is expressed as: Among them, b ω ,b dω ,b P >0, is the weight coefficient, different weight coefficients are designed to obtain different output characteristics, then by inputting s t The reward value obtained is defined as r t .

5. The method for adaptive control of a virtual synchronous generator under network attacks according to claim 4, Features: The specific method of step S7 is: S7.

1. Calculate the loss function used to update the critic neural network parameters: in, is the loss function of the critic network for controlling parameters, Indicates that the expected sampling space for subsequent expressions is the data in memory pool D, Represents the corresponding expectation in the target space; and calculates the relevant gradient in Indicates the gradient calculation for the control function parameters in the critic network; S7.

2. Calculate the loss function used to update the actor neural network parameters: where J is the approximation function, E st represents that the sampling space for which the subsequent expression is expected to be solved is the state at time t, and represents the gradient calculation for the control function parameters in the actor network through the chain rule, and the parameters and are further updated through local calculation based on their own and neighbor information.

6. The method for adaptively controlling a virtual synchronous generator under network attacks according to claim 5, Features: The specific method of step S8 is: S8.1: Introduce consensus variables y and λ; define λ as the first consensus variable, λ(h+1)=Mλ(h)+ηy(h), where λ represents the column stack vector of IC; η is the convergence point coefficient, which is a sufficiently small positive constant; M is the updated consensus matrix for the first consensus agreement; y is the second consensus agreement, y(h+1)=Ny(h)-(x(h+1)-x(h)), where y is the column stack vector of the local estimated energy mismatch; x is the column stack vector of the energy output; N is the updated consensus matrix for the second consensus agreement; S8.

2. Calculate the relevant judgment information; calculate the reasonable bounds of the relevant information by using the calculation units other than the suspected attacked units; the specific method is: S8.2.1, for the computation of the second consensus y: In the next iteration, a reasonable bound on the mismatch of the local estimated energy of the vth computational unit is: Where Y(h,R) is the detection threshold function determined by iterating h and R, R is a common parameter representing the maximum value of the ramp rate limit of various energy devices, and D + , D - Represent positive and negative value functions respectively; S8.2.

2. For the calculation of the first consensus agreement λ: In the next iteration, the rational bound on the IC of the vth computational unit is: In the formula Where l represents the physical transmission line, They represent the transmission lines when the first consensus agreement takes the upper and lower bounds, respectively; is the set of computing units connected to the data output of the vth computing unit, λ l (h) represents the actual IC value obtained in the hth iteration of the unit connected by the physical transmission line of this unit, and σ is the first consensus agreement limit coefficient; S8.3, the received local energy mismatch actual value y v (h) Compare with the reasonable bound of the local estimated energy mismatch of computing unit v; if the actual value y v (h) Satisfaction Then go to step S8.4; otherwise, the computing unit v is judged to be misbehaving, and go to step S8.5; S8.

4. The actual incremental electricity cost λ v (h) Compare with the theoretical estimated range; if Then the computing unit v is judged to have good behavior; otherwise, the computing unit v is judged to have bad behavior; If the unit does not exist, the computation unit v is directly judged to be misbehaving; S8.

5. Update the detection result variable d of the calculation unit v according to the following formula v (h): Send the updated detection result variables to adjacent computing units; S8.6, Isolation process under non-coordinated attack: The network attack only targets a certain computing unit; the details are as follows: S8.6.

1. The counting rules are defined as: Among them, d v (h) is the detection result variable of the h-th iteration updating calculation unit v, T iv (h) is the number of well-behaved computational units after the hth iteration; S8.6.

2. The reputation value is designed as a threshold to determine whether the computing unit is isolated, as follows: Among them, ω h > 0 is a time-varying parameter called reputation coefficient, which is used to dynamically adjust the adaptation rate; when all units behave well, rep iv (h) will always be e; when a misbehaving unit appears, rep iv (h)<e; when the reputation value of all communication lines connected to the abnormal unit drops below the threshold, the abnormal unit is disconnected from the communication network; S8.

7. Isolation process under coordinated attack: The public network launches a collusion attack on two or more computing units at the same time, as follows: S8.7.

1. After the abnormal unit with bad behavior has been isolated through the process of step S8.6, the connection information is recorded in the observation subnet and the link matrix is ​​updated as follows: Among them, f iv (h) represents the value of the connection matrix between the ith computing unit and the vth computing unit after the hth iteration process, E' is the edge set of the observed network topology, is the set of computing units connected to the data output of the vth computing unit, The set of computing units connected to the data input of the vth computing unit; S8.7.2, link matrix F(h) = [f ij ] n×n , if a non-zero element appears in the vth row of the link matrix, it means that the abnormally behaving computing unit v is not completely isolated, that is, there is a collusion attack in the communication network; the cth computing unit adjacent to the vth computing unit is easily judged as a collusion unit of the computing unit v and is considered to be a misbehaving computing unit, then the computing unit c repeats the detection process in steps S8.2-S8.5, and uploads the detection results to perform the relevant isolation process in step S8.6; S8.7.

3. Repeat steps S8.7.1-S8.7.2 until all abnormal units are isolated.

7. The method for adaptively controlling a virtual synchronous generator under network attacks according to claim 6, Features: The specific method of step S9 is: S9.

1. Update the relevant parameters of the critic neural network: And calculate the average value of the parameters of a total of k computing units: where ζ Q is the learning rate of the critic neural network; S9.2, update the actor neural network related parameters: And calculate the average value of the parameters of a total of k computing units: where ζ u is the learning rate of the actor neural network.

Citation Information

Patent Citations

  • Adaptive adjustment inverter controller

    CN110880774A

  • Improved virtual synchronous generator model prediction control method and system

    CN113179059A