A Microgrid Frequency Control Method Based on Quantum Reinforcement Learning

CN120300828BActive Publication Date: 2026-09-01NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510352927.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2026-09-01
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

[0007]针对背景技术中现有技术的不足,本发明提出了一种基于量子强化学习的微电网频率控制方法,利用量子强化学习对光储一体机的输出频率进行控制,能有效解决分布式电源在孤岛微电网下频率响应速度慢的问题

Benefits of technology

[0076] (1) This invention uses a quantum reinforcement learning algorithm to perform secondary control of the microgrid frequency, which can significantly accelerate the frequency recovery speed. Moreover, compared with the traditional virtual synchronous machine control method, this invention significantly improves the dynamic response capability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120300828B_ABST
    Figure CN120300828B_ABST
Patent Text Reader

Abstract

This invention discloses a microgrid frequency control method based on quantum reinforcement learning. The method utilizes the output characteristics of photovoltaic (PV) and energy storage systems, setting the PV to operate at its maximum power output point and amplifying it via a Boost circuit, while the energy storage is charged and discharged using a Buck-Boost circuit, constructing a PV-energy storage integrated machine model based on coordinated PV-energy storage output. Secondly, a primary control model based on a virtual synchronous machine is constructed, and a secondary frequency control model is established after analyzing external disturbances. The obtained secondary frequency control signal is analyzed, establishing an action set and state level. A reward function is designed using quantum reinforcement learning, and the secondary frequency control is achieved by updating the Q-value, quickly restoring the frequency to 50Hz. This invention utilizes quantum reinforcement learning to control the output frequency of the PV-energy storage integrated machine, effectively solving the problem of slow frequency response speed of distributed power sources in isolated microgrids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of microgrid frequency control technology, and specifically to a microgrid frequency control method based on quantum reinforcement learning. Background Technology

[0002] Currently, distributed power sources in isolated microgrids mainly consist of photovoltaics (PV) and energy storage. However, PV output is intermittent and random, increasing frequency uncertainty and reducing the anti-interference capability of the traditional power grid, thus impacting its stability and security. When a microgrid operates in islanded mode, without the support of a large power grid for frequency stability, PV must rely on energy storage to coordinate and support frequency operation. Therefore, considering the uncertainty of PV output, it is necessary to utilize integrated PV and energy storage systems to control the active power of PV and energy storage to achieve frequency stability, enabling more stable system operation, regulation, and management.

[0003] Currently, the coordinated control model for integrated photovoltaic and energy storage systems models photovoltaics and energy storage separately, which can be equivalent to a single DC voltage source connected to the inverter. The paper "Research on Grid-connected Control Technology of Photovoltaic Energy Storage Based on Virtual Synchronous Machine" proposes a primary and secondary frequency regulation control strategy based on the charging and discharging characteristics of the energy storage battery. The paper "Research on Multi-level Voltage Control and Islanding Frequency Stability of Distribution Networks with Distributed Photovoltaics" uses VSG control in islanding mode to prevent power frequency fluctuations caused by large disturbances.

[0004] Experts and scholars in the field of power systems have also conducted extensive research on secondary control. The paper "Distributed Secondary Optimization Control of Multiple Microgrids Based on Reinforcement Learning" addresses the static frequency and voltage deviation problem caused by the drooping primary control of distributed power sources in microgrids, proposing a distributed secondary optimization control method based on reinforcement learning local feedback. The paper "Frequency Coordination Control Strategy of Multiple Photovoltaic-Storage Virtual Synchronizers Based on Reinforcement Learning" (Electrical Drive) proposes a frequency coordination control strategy for multiple photovoltaic-storage virtual synchronizers based on reinforcement learning.

[0005] The aforementioned literature has the following drawbacks: When a microgrid operates isolated from the main grid, its frequency is determined by its own distributed power sources. However, photovoltaic output is intermittent, requiring energy storage to immediately generate or absorb active power to achieve frequency stability. However, the control method based on reinforcement learning is slow in terms of training data, resulting in unreliable power supply.

[0006] Therefore, this invention patent utilizes quantum reinforcement learning theory to construct a correlation model of the photovoltaic-storage integrated machine and a secondary frequency control model of the photovoltaic-storage integrated machine. It proposes a microgrid frequency control method based on quantum reinforcement learning. By connecting the photovoltaic and energy storage in parallel as equivalent DC voltage sources and connecting them to the inverter's virtual synchronous machine frequency primary control model, the frequency deviation of the inverter output is then subjected to secondary control. The control process is exponentially accelerated using the quantum reinforcement learning method. Summary of the Invention

[0007] To address the shortcomings of existing technologies in the background section, this invention proposes a microgrid frequency control method based on quantum reinforcement learning. By utilizing quantum reinforcement learning to control the output frequency of the photovoltaic-storage integrated machine, this method can effectively solve the problem of slow frequency response speed of distributed power sources in isolated microgrids.

[0008] To achieve the technical objective of this invention, the following technical solution is adopted:

[0009] A microgrid frequency control method based on quantum reinforcement learning includes the following steps:

[0010] Step S1: Based on the output characteristics of photovoltaic and energy storage, define that the photovoltaic works at the maximum power output point and is amplified by the Boost circuit, and that the energy storage is charged and discharged through the Buck-Boost circuit, and construct a photovoltaic-energy storage integrated machine model based on the coordinated output of photovoltaic and energy storage.

[0011] Step S2: Based on the integrated photovoltaic and energy storage model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a secondary frequency control model for external disturbance analysis.

[0012] Step S3: Analyze the frequency secondary control model established in step S2, establish the state set, the frequency action set and frequency state set of the integrated photovoltaic energy storage secondary control, design the reward function, and use the quantum Q learning method to update and obtain the optimal action.

[0013] Furthermore, the method for defining the photovoltaic system's operation at its maximum power output point and amplifying it via a Boost circuit in step S1 is as follows:

[0014] Step S1-1-1: Define the output current of the photovoltaic cell as follows:

[0015]

[0016] In the formula, I is the output current of the photovoltaic cell, I ph I0 is the photogenerated equivalent current, I0 is the PN junction current, ... shq is the leakage current of the photovoltaic cell, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, and R... sh This is the equivalent parallel resistance value of the actual photovoltaic cell;

[0017] Step S1-1-2: Obtain its short-circuit current I SC Open circuit voltage U OC Output P at maximum power m Voltage U at maximum power m and the current I at maximum power m Equation (1) is transformed into an equivalent form, and the maximum current I at this time is... m for:

[0018]

[0019] The open-circuit current is:

[0020]

[0021] Steps S1-1-3, combined with Equations 2 and 3, yield the undetermined coefficients of capacitors C1 and C2:

[0022]

[0023] Step S1-1-4: By determining the numerical values ​​of Equation 4, the operating conditions of the photovoltaic equivalent circuit are obtained. The maximum power point tracking algorithm is then used to obtain the specific value of the capacitor at maximum power.

[0024] Step S1-1-5: Based on the photovoltaic port characteristic curve, the photovoltaic cell is equivalent to a DC voltage source; the front-end uses a Boost converter circuit to boost the photovoltaic array voltage first.

[0025] Step S1-1-6: By adjusting the duty cycle to control the switching transistor S, the DC voltage on the photovoltaic side is regulated. The input-output relationship of this circuit is as follows:

[0026]

[0027] In the above formula, t on t off T d These represent the on and off times and periods of the switching transistor S, respectively, and E is the output voltage of the photovoltaic cell. The output voltage E of the photovoltaic cell rises to U0 and is then fed into the inverter.

[0028] Furthermore, in step S1, the energy storage circuit is connected in parallel with the photovoltaic Boost circuit via a DC-DC circuit, and a Buck-Boost circuit is used for energy storage. The method for charging and discharging the energy storage through the Buck-Boost circuit is as follows:

[0029] Step S1-2-1: PWM technology can control the on / off state of switching transistors S1 and S2 to achieve bidirectional energy flow. The DC / DC converter has two operating modes: when the energy storage battery needs to output power externally, i.e., when discharging externally, the converter operates in Boost mode, and energy flows from the lithium battery to the DC bus; when the lithium battery needs to absorb power, i.e., to charge itself, the converter operates in Buck mode, and energy flows from the DC bus to the energy storage battery.

[0030] Step S1-2-2: When the converter operates in Buck mode, switch S2 is turned on and switch S1 is turned off, and current flows from the bus to the battery, satisfying the following:

[0031] U bat =DU dc (6)

[0032] In the above formula, U bat Where is the energy storage voltage, D is the duty cycle, and U is the energy storage voltage. dc This is the DC bus voltage.

[0033] Step S1-2-3: When the converter operates in Boost mode, switch S1 is turned on and switch S2 is turned off, and current flows from the battery to the bus, satisfying the following:

[0034] U bat =(1-D)U dc (7)

[0035] Among them, the switching transistors S1 and S2 are complementary, with one being on and the other off. Only one control signal is needed to control the switching devices of the bidirectional DC / DC converter, enabling the battery to be charged and discharged reasonably as needed, and adjusting the duty cycle during the charging and discharging process to change the DC bus voltage.

[0036] Furthermore, the primary control model based on the virtual synchronizer described in step S2 achieves active power distribution through droop control; the expression for the primary frequency control is as follows:

[0037] f = f ref +K f (P ref -P) (8)

[0038] In the above formula, f ref K f P refP and P represent the frequency reference value, frequency droop control coefficient, active power reference value, and actual active power value, respectively.

[0039] Furthermore, the frequency secondary control model described in step S2 collects the real-time frequency magnitudes of distributed power sources and uses the central controller to issue control commands to each distributed power source; the secondary control compensates for the frequency of the microgrid, restoring it to its rated value; the frequency secondary control expression is as follows:

[0040] f = f ref +K f (P ref -P)+Δf * (9)

[0041] In the above formula, Δf * This is a secondary frequency control command.

[0042] Furthermore, in step S3, quantum reinforcement learning is used to perform secondary control of the frequency; the quantum reinforcement learning algorithm uses quantum states to represent the data:

[0043] Step S3-1-1: Encode the traditional state set and action set, convert them into quantum states, and limit the frequency to within 50±0.5Hz;

[0044] Step S3-1-2: Retain the frequency deviation to two decimal places and set the state set to: S f A set of 101 elements of the expression {-0.5,-0.49,-0.48,-0.47,-0.46,-0.45,……,-0.05,-0.04,-0.03,-0.02,-0.01,0,0.01,0.02,0.03,0.04,0.05,……,0.45,0.46,0.47,0.48,0.49,0.50}Hz.

[0045] Step S3-1-3, the frequency action set of the secondary control of the designed integrated photovoltaic and energy storage unit is: A f ={-0.2,-0.15,-0.1,-0.06,-0.01,-0.006,-0.002,0,0.002,0.006,0.01,0.06,0.1,0.15,0.2}Hz.

[0046]

[0047] In the above formula, c f Here, Δf represents the reward function coefficients, and Δf represents the frequency deviation.

[0048] Since quantum gates do not directly handle multiplication operations, and constructing quantum multiplication circuits using quantum gates is quite complex, a combined quantum-classical approach is employed in reward function calculation. In quantum reinforcement learning, the reward function coefficient c is represented by 4 bits. f The reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} are represented as 0000, 0001, ..., 1001 and assigned to the quantum states of the state set. Then, the quantum states are sent to a classical computer to assign the corresponding reward function coefficients to 0000, 0001, ..., 1001 and perform multiplication calculations.

[0049] Given a set of states, actions, and a reward function, we select the optimal action and update the Q-value table to store the expected cumulative reward obtained by the agent in different states by taking different actions. The frequency-based secondary control command, the Q-value update, and the optimal action formula are shown below:

[0050]

[0051] Where, Δf * This is a secondary frequency control command, Q(s) t ,a t ) represents the Q-value table in classical computers; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9, s t Let s be the state at time t. t+1 Let a be the state at time t+1. t Let t be the action at time t, and a be the possible actions to be taken.

[0052] Given a set of states, actions, and a reward function, we select the optimal action and update the Q-value table to store the expected cumulative reward obtained by the agent in different states by taking different actions. The frequency-based secondary control command, the Q-value update, and the optimal action formula are shown below:

[0053]

[0054] Where, Δf * This is a secondary frequency control command, Q(s) t ,a t ) represents the Q-value table in a classic computer; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9. Further, the frequency action set for the secondary control of the integrated optical storage machine designed in step S3 is: A f ={-0.2,-0.15,-0.1,-0.06,-0.01,-0.006,-0.002,0,0.002,0.006,0.01,0.06,0.1,0.15,0.2}Hz.

[0055] Furthermore, in step S3, the frequency action set of the state set and the secondary control of the integrated optical storage machine are encoded using amplitude encoding, using 7 bits and 4 bits respectively to represent the quantum state set |s> and the quantum action set |a>.

[0056] The quantum state set is represented as:

[0057]

[0058] The quantum action set is represented as:

[0059]

[0060] in,

[0061]

[0062] Where, |ρ n | and |β n | represents the probability of the element corresponding to the quantum state and quantum action, respectively;

[0063] Furthermore, in step S3, the Grover algorithm is used to update the probability amplitude of the quantum action set |a>; during amplitude update, the Grover operation is repeated to reinforce actions that yield high rewards as output; the specific method is as follows:

[0064] Step S3-2-1: Initialize the behavior. The initialization formula is shown below:

[0065]

[0066] Among them, |a s (n) > is the initialization action superposition state.

[0067] Step S3-2-2: Construct a U-transformation, defined as Ua = 2|a s (n) > s (n) |-I, keep|a s (n) >, but |a s (n) >It will flip the sign of any orthogonal basis vector; when Ua acts on any vector, it preserves |a s ( n) The structure >, but it will flip |a s (n) > orthogonal basis symbols;

[0068] Step S3-2-3: Use Grover operation, with the i-th state |a i >Replace |a​s (n) >, construct transformation U ai =I-2|a i > i | Define a unitary matrix U Grov =U a U ai Through |a s (n)> Repeated application of U Grov The i-th action |a i The amplitude probability of > is increased, while the amplitude of other behaviors is reduced;

[0069] Step S3-2-4, the initial action f(s) can be expressed as:

[0070]

[0071] In Formula 16, an angle θ is defined that satisfies Established at this time Similarly, perform the Grover operation on |as(n)> again. This is represented as:

[0072]

[0073] Step S3-2-5: By repeating the Grover operation, the probability amplitude of the corresponding action is enhanced according to the reward value, so that any state can be transformed into another specified ground state with a high probability.

[0074] Perform an action from the quantum action set |a>, observe all new states, update the policy and Q value using formula (10), explore the next action, and repeat the Grover iteration operation L times to update the probability amplitude until all states approach 0 Hz.

[0075] Compared with the prior art, the present invention has the following beneficial effects:

[0076] (1) This invention uses a quantum reinforcement learning algorithm to perform secondary control of the microgrid frequency, which can significantly accelerate the frequency recovery speed. Moreover, compared with the traditional virtual synchronous machine control method, this invention significantly improves the dynamic response capability of the system.

[0077] (2) Through quantum state encoding and Grover's algorithm, this invention can achieve high-precision adjustment of frequency deviation. The design of the state set and action set enables the system to achieve precise control within a range of ±0.5Hz, effectively avoiding system instability caused by excessive frequency fluctuations.

[0078] ​(3) The photovoltaic-storage integrated machine model proposed in this invention can make full use of the maximum power output point algorithm of photovoltaic and the fast charging and discharging characteristics of energy storage through the coordinated output of photovoltaic and energy storage, and automatically realize the charging and discharging of energy storage, which significantly improves the energy utilization efficiency and frequency stability of microgrid.

[0079] (4) This invention utilizes quantum reinforcement learning to accelerate the secondary frequency control, enabling the frequency of the isolated microgrid to recover rapidly.

[0080] (5) This invention utilizes a quantum reinforcement learning algorithm, leveraging the parallelism of quantum computing and the search acceleration characteristics of Grover's algorithm to significantly improve the computational efficiency of frequency control. Compared to traditional reinforcement learning methods, quantum reinforcement learning can find the optimal control strategy in a shorter time, improving the real-time performance and reliability of the system. Attached Figure Description

[0081] Figure 1 This is a flowchart of a microgrid frequency control method based on quantum reinforcement learning according to the present invention;

[0082] Figure 2 This is a diagram of an integrated photovoltaic and energy storage system based on a microgrid frequency control method using quantum reinforcement learning, as described in this invention.

[0083] Figure 3 This is a Buck-Boost charge-discharge diagram of energy storage based on a microgrid frequency control method using quantum reinforcement learning, as described in this invention.

[0084] Figure 4 This is a diagram illustrating the effect of the primary voltage control of the present invention;

[0085] Figure 5 This is a diagram illustrating the effect of the secondary frequency control of the present invention. Detailed Implementation

[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0087] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0088] Example 1

[0089] like Figure 1 As shown, a microgrid frequency control method based on quantum reinforcement learning includes the following steps:

[0090] Step S1: Based on the output characteristics of photovoltaic and energy storage, define that the photovoltaic works at the maximum power output point and is amplified by the Boost circuit, and that the energy storage is charged and discharged through the Buck-Boost circuit, and construct a photovoltaic-energy storage integrated machine model based on the coordinated output of photovoltaic and energy storage.

[0091] Step S2: Based on the integrated photovoltaic and energy storage model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a secondary frequency control model for external disturbance analysis.

[0092] Step S3: Analyze the frequency secondary control model established in step S2, establish the state set, the frequency action set and frequency state set of the integrated photovoltaic energy storage secondary control, design the reward function, and use the quantum Q learning method to update and obtain the optimal action.

[0093] Furthermore, the method for defining the photovoltaic system's operation at its maximum power output point and amplifying it via a Boost circuit in step S1 is as follows:

[0094] Step S1-1-1: Define the output current of the photovoltaic cell as follows:

[0095]

[0096] In the formula, I is the output current of the photovoltaic cell, I ph I0 is the photogenerated equivalent current, I0 is the PN junction current, ... sh q is the leakage current of the photovoltaic cell, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, and R... sh This is the equivalent parallel resistance value of the actual photovoltaic cell;

[0097] Step S1-1-2: Obtain its short-circuit current I SC Open circuit voltage U OC Output P at maximum power m Voltage U at maximum power m and the current I at maximum power m Equation (1) is transformed into an equivalent form, and the maximum current I at this time is... m for:

[0098]

[0099] The open-circuit current is:

[0100]

[0101] Steps S1-1-3, combined with Equations 2 and 3, yield the undetermined coefficients of capacitors C1 and C2:

[0102]

[0103] Step S1-1-4: By determining the numerical values ​​of Equation 4, the operating conditions of the photovoltaic equivalent circuit are obtained. The maximum power point tracking algorithm is then used to obtain the specific value of the capacitor at maximum power.

[0104] Step S1-1-5: Based on the photovoltaic port characteristic curve, the photovoltaic cell is equivalent to a DC voltage source; the front-end uses a Boost converter circuit to boost the photovoltaic array voltage first.

[0105] Step S1-1-6: By adjusting the duty cycle to control the switching transistor S, the DC voltage on the photovoltaic side is regulated. The input-output relationship of this circuit is as follows:

[0106]

[0107] In the above formula, t on t off T d These represent the on and off times and periods of the switching transistor S, respectively, and E is the output voltage of the photovoltaic cell. The output voltage E of the photovoltaic cell rises to U0 and is then fed into the inverter.

[0108] Furthermore, in step S1, the energy storage circuit is connected in parallel with the photovoltaic Boost circuit via a DC-DC circuit, and a Buck-Boost circuit is used for energy storage. The method for charging and discharging the energy storage through the Buck-Boost circuit is as follows:

[0109] Step S1-2-1: PWM technology can control the on / off state of switching transistors S1 and S2 to achieve bidirectional energy flow. The DC / DC converter has two operating modes: when the energy storage battery needs to output power externally, i.e., when discharging externally, the converter operates in Boost mode, and energy flows from the lithium battery to the DC bus; when the lithium battery needs to absorb power, i.e., to charge itself, the converter operates in Buck mode, and energy flows from the DC bus to the energy storage battery.

[0110] Step S1-2-2: When the converter operates in Buck mode, switch S2 is turned on and switch S1 is turned off, and current flows from the bus to the battery, satisfying the following:

[0111] U bat =DU dc (twenty three)

[0112] In the above formula, U bat Where is the energy storage voltage, D is the duty cycle, and U is the energy storage voltage. dcThis is the DC bus voltage.

[0113] Step S1-2-3: When the converter operates in Boost mode, switch S1 is turned on and switch S2 is turned off, and current flows from the battery to the bus, satisfying the following:

[0114] U bat =(1-D)U dc (twenty four)

[0115] Among them, the switching transistors S1 and S2 are complementary, with one being on and the other off. Only one control signal is needed to control the switching devices of the bidirectional DC / DC converter, enabling the battery to be charged and discharged reasonably as needed, and adjusting the duty cycle during the charging and discharging process to change the DC bus voltage.

[0116] Furthermore, the primary control model based on the virtual synchronizer described in step S2 achieves active power distribution through droop control; the expression for the primary frequency control is as follows:

[0117] f = f ref +K f (P ref -P) (25)

[0118] In the above formula, f ref K f P ref P and P represent the frequency reference value, frequency droop control coefficient, active power reference value, and actual active power value, respectively.

[0119] Furthermore, the frequency secondary control model described in step S2 collects the real-time frequency magnitudes of distributed power sources and uses the central controller to issue control commands to each distributed power source; the secondary control compensates for the frequency of the microgrid, restoring it to its rated value; the frequency secondary control expression is as follows:

[0120] f = f ref +K f (P ref -P)+Δf * (26)

[0121] In the above formula, Δf * This is a secondary frequency control command.

[0122] Furthermore, in step S3, quantum reinforcement learning is used to perform secondary control of the frequency; the quantum reinforcement learning algorithm uses quantum states to represent the data:

[0123] Step S3-1-1: Encode the traditional state set and action set, convert them into quantum states, and limit the frequency to within 50±0.5Hz;

[0124] Step S3-1-2: Retain the frequency deviation to two decimal places and set the state set to: S f = A set of 101 elements of {-0.5,-0.49,-0.48,-0.47,-0.46,-0.45,……,-0.05,-0.04,-0.03,-0.02,-0.01,0,0.01,0.02,0.03,0.04,0.05,……,0.45,0.46,0.47,0.48,0.49,0.50}Hz.

[0125] Step S3-1-3, the frequency action set of the secondary control of the designed integrated photovoltaic and energy storage unit is: A f ={-0.2,-0.15,-0.1,-0.06,-0.01,-0.006,-0.002,0,0.002,0.006,0.01,0.06,0.1,0.15,0.2}Hz.

[0126]

[0127] In the above formula, c f Here, Δf represents the reward function coefficients, and Δf represents the frequency deviation.

[0128] Since quantum gates do not directly handle multiplication operations, and constructing quantum multiplication circuits using quantum gates is quite complex, a quantum-classical combined approach is adopted when calculating the reward function. In quantum reinforcement learning, 4 bits are used to represent the reward function coefficients, and the reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} are represented as 0000, 0001, ..., 1001 respectively, and assigned to the quantum states of the state set. Then, the classical computer assigns the corresponding reward function coefficients to 0000, 0001, ..., 1001, and performs multiplication calculations.

[0129] Given a set of states, actions, and a reward function, we select the optimal action and update the Q-value table to store the expected cumulative reward obtained by the agent in different states by taking different actions. The frequency-based secondary control command, the Q-value update, and the optimal action formula are shown below:

[0130]

[0131] Where, Δf * This is a secondary frequency control command, Q(s) t ,a t ) represents the Q-value table in a classic computer; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9.

[0132] Furthermore, in step S3, the frequency action set of the state set and the secondary control of the integrated optical storage machine are encoded using amplitude encoding, using 7 bits and 4 bits respectively to represent the quantum state set |s> and the quantum action set |a>.

[0133] The quantum state set is represented as:

[0134]

[0135] The reward function coefficients are expressed as follows:

[0136] The quantum action set is represented as:

[0137]

[0138] in,

[0139]

[0140] Where, |ρ n | and |β n | represents the probabilities of the corresponding elements in the quantum state set and the quantum action set, respectively;

[0141] Furthermore, in step S3, the Grover algorithm is used to update the probability amplitude of the quantum action set |a>; during amplitude update, the Grover operation is repeated to reinforce actions that yield high rewards as output; the specific method is as follows:

[0142] Step S3-2-1: Initialize the behavior. The action set initialization formula is shown below:

[0143]

[0144] Among them, |a s (n) > is the initialization action superposition state.

[0145] Step S3-2-2: Construct a U-transformation, defined as a quantum unitary transform Ua = 2|a s (n) > s (n) |-I, keep|a s (n) >, but |a s (n) >It will flip the sign of any orthogonal basis vector; when Ua acts on any vector, it preserves |a s (n) The structure >, but it will flip |a s (n) > orthogonal basis symbols;

[0146] ​Step S3-2-3: Use Grover operation, with the i-th state |a i >Replace |a s (n) >, construct transformation U ai =I-2|a i > i | Define a unitary matrix U Grov =U a U ai Through |a s (n) >Repeated application of U Grov The i-th action |a i The amplitude probability of > is increased, while the amplitude of other behaviors is reduced;

[0147] Step S3-2-4, the initial action f(s) can be expressed as:

[0148]

[0149] In Formula 16, an angle θ is defined that satisfies Established at this time Similarly, perform the Grover operation on |as(n)> again. This is represented as:

[0150]

[0151] Step S3-2-5: By repeating the Grover operation, the probability amplitude of the corresponding action is enhanced according to the reward value, so that any state can be transformed into another specified ground state with a high probability.

[0152] Perform an action from the quantum action set |a>, observe all new states, update the policy and Q value using formula (10), explore the next action, and repeat the Grover iteration operation L times to update the probability amplitude until all states approach 0 Hz.

[0153] To verify the effectiveness of the present invention, the following methods were used: Figure 3 The control system of the photovoltaic energy storage integrated machine shown was verified.

[0154] Implementation Case 1

[0155] A frequency control model for an integrated photovoltaic and energy storage system was established, including a primary frequency control model based on virtual synchronous machine control and a secondary frequency control model based on quantum reinforcement learning. The proposed secondary frequency control based on quantum reinforcement learning was then verified to better and faster recover the frequency.

[0156] ​The droop control coefficient of the photovoltaic-storage integrated machine is set to 9.3e-3, and the rated active power of the load is 2000kW. A 1000kW active power load is connected in the second second, a 1500kW active power load is disconnected in the second and the 800kW active power load is connected in the third second.

[0157] Figure 4 This is a diagram illustrating the effect of a primary voltage control method based on droop control. From... Figure 4 In the simulation waveform, it can be found that the effective value of the inverter output frequency fluctuates around 50Hz, and there is a relatively obvious fluctuation after the load is connected.

[0158] The frequency is quickly restored using secondary control based on primary control. The parameters are the same as in Scenario 1. Figure 5 This is a diagram illustrating the effect of a secondary frequency control method for a distribution network based on droop control. From... Figure 5 Simulation waveforms show that the effective value of the microgrid output frequency fluctuates around 50Hz and can recover the frequency quickly. Compared to primary control, secondary control has a better effect in terms of rapid frequency recovery.

[0159] As can be seen from Table 1, frequency control based on quantum reinforcement learning can effectively and quickly recover frequency fluctuations caused by load mutations. This also verifies the feasibility of the present invention.

[0160] Table 1 Frequency Comparison

[0161]

[0162] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A microgrid frequency control method based on quantum reinforcement learning, characterized in that, Includes the following steps: Step S1: Based on the output characteristics of photovoltaic and energy storage, define that the photovoltaic works at the maximum power output point and is amplified by the Boost circuit, and that the energy storage is charged and discharged through the Buck-Boost circuit, and construct a photovoltaic-energy storage integrated machine model based on the coordinated output of photovoltaic and energy storage. Step S2: Based on the integrated photovoltaic and energy storage model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a secondary frequency control model for external disturbance analysis. Step S3: Analyze the frequency secondary control model established in step S2, establish the state set, the quantum action set and quantum state set of the integrated photovoltaic storage machine secondary control, design the reward function, and use the quantum Q learning method to update and obtain the optimal action; The frequency action set for the secondary control of the integrated photovoltaic and energy storage unit is: A f ={-0.2,-0.15,-0.1,-0.06,-0.01,-0.006,-0.002,0,0.002,0.006,0.01,0.06,0.1,0.15,0.2}Hz; the reward function is calculated using a quantum-classical combined approach, and the reward function is set as follows: (1) In the above formula, c f Here, Δf represents the reward function coefficients, and Δf represents the frequency deviation. In quantum reinforcement learning, the reward function coefficient c is represented by 4 bits. f The reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} are represented as 0000, 0001, ..., 1001 and assigned to the quantum states of the state set. Then, the quantum states are sent to a classical computer to assign the corresponding reward function coefficients to 0000, 0001, ..., 1001 and perform multiplication calculations. Given a set of states, actions, and a reward function, we select the optimal action and update the Q-value table to store the expected cumulative reward obtained by the agent in different states by taking different actions. The frequency-based secondary control command, the Q-value update, and the optimal action formula are shown below: (2) Where, Δf * This is a secondary frequency control command, Q(s) t ,a t ) represents the Q-value table in classical computers; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9, s t Let s be the state at time t. t+1 Let a be the state at time t+1. t Let t be the action taken at time t, and a be the possible actions to be taken. The state set and the frequency action set of the integrated photovoltaic energy storage machine's secondary control are encoded using amplitude encoding, with 7 and 4 bits respectively used to represent the quantum state set. and quantum action set ; The quantum state set is represented as: (3) The quantum action set is represented as: (4) in, (5) (6) Where, |ρ n | and |β n | represents the probability of the element corresponding to the quantum state and quantum action, respectively.

2. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, The method for defining the photovoltaic system's operation at its maximum power output point and amplifying it via a Boost circuit in step S1 is as follows: Step S1-1-1: Define the output current of the photovoltaic cell as follows: (7) In the formula, I is the output current of the photovoltaic cell, I ph I0 is the photogenerated equivalent current, I0 is the PN junction current, ... sh q is the leakage current of the photovoltaic cell, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, and R... sh This is the equivalent parallel resistance value of the actual photovoltaic cell; Step S1-1-2: Obtain its short-circuit current I SC Open circuit voltage U OC Output P at maximum power m Voltage U at maximum power m and the current I at maximum power m Equation (1) is transformed into an equivalent form, and the maximum current I at this time is... m for: (8) The open-circuit current is: (9) Steps S1-1-3, combined with equations (2) and (3), yield the undetermined coefficients of capacitors C1 and C2: (10) Step S1-1-4: By determining the value of equation (4), the operating conditions of the photovoltaic equivalent circuit are obtained. The maximum power point tracking algorithm is used to obtain the specific values ​​of capacitors C1 and C2 at maximum power. Step S1-1-5: Based on the photovoltaic port characteristic curve, the photovoltaic cell is equivalent to a DC voltage source; the front-end uses a Boost converter circuit to boost the photovoltaic array voltage first. Step S1-1-6: By adjusting the duty cycle to control the switching transistor S, the DC voltage on the photovoltaic side is regulated. The input-output relationship of this circuit is as follows: (11) In the above formula, t on t off T D These represent the on and off times and periods of the switching transistor S, respectively, and E is the output voltage of the photovoltaic cell. The output voltage E of the photovoltaic cell rises to U0 and is then fed into the inverter.

3. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, In step S1, the energy storage circuit is connected in parallel with the photovoltaic Boost circuit via a DC-DC circuit. A Buck-Boost circuit is used for the energy storage. The method for charging and discharging the energy storage through the Buck-Boost circuit is as follows: Step S1-2-1: PWM technology can control the on / off state of switching transistors S1 and S2 to achieve bidirectional energy flow. The DC / DC converter has two operating modes: when the energy storage battery needs to output power externally, i.e., when discharging externally, the converter operates in Boost mode, and energy flows from the lithium battery to the DC bus; when the lithium battery needs to absorb power, i.e., to charge itself, the converter operates in Buck mode, and energy flows from the DC bus to the energy storage battery. Step S1-2-2: When the converter operates in Buck mode, switch S2 is turned on and switch S1 is turned off, and current flows from the bus to the battery, satisfying the following: (12) In the above formula, U bat Where is the energy storage voltage, D is the duty cycle, and U is the energy storage voltage. dc This is the DC bus voltage; Step S1-2-3: When the converter operates in Boost mode, switch S1 is turned on and switch S2 is turned off, and current flows from the battery to the bus, satisfying the following: (13) Among them, the switching transistors S1 and S2 are complementary, with one being on and the other off. Only one control signal is needed to control the switching devices of the bidirectional DC / DC converter, enabling the battery to be charged and discharged reasonably as needed, and adjusting the duty cycle during the charging and discharging process to change the DC bus voltage.

4. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, The primary control model based on the virtual synchronizer described in step S2 achieves active power distribution through droop control; the expression for the primary frequency control is as follows: (14) In the above formula, f ref For frequency reference value, K f For frequency droop control coefficient, P ref P is the reference value for active power, and P is the actual value of active power.

5. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, The frequency secondary control model described in step S2 collects the real-time frequency magnitudes of distributed power sources and uses a central controller to issue control commands to each distributed power source; the secondary control compensates for the frequency of the microgrid, restoring it to its rated value; the frequency secondary control expression is as follows: (15) In the above formula, Δf * This is a secondary frequency control command.

6. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, In step S3, quantum reinforcement learning is used to perform secondary control of the frequency; the quantum reinforcement learning algorithm uses quantum states to represent the data. Step S3-1-1: Encode the traditional state set and action set, convert them into quantum states, and limit the frequency to within 50±0.5Hz; Step S3-1-2: Retain the frequency deviation to two decimal places and set the state set to: S f ={-0.5,-0.49,-0.48,-0.47,-0.46,-0.45,……,-0.05,-0.04,-0.03,-0.02,-0.01,0,0.01,0.02,0.03,0.04,0.05,……,0.45,0.46,0.47,0.48,0.49,0.50}Hz.

7. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, In step S3, the Grover algorithm is used to update the probability amplitude of the quantum action set |a>. During amplitude update, the Grover operation is repeated to reinforce actions that yield high rewards as output. The specific method is as follows: Step S3-2-1: Initialize the behavior. The initialization formula for the quantum action set |a> is shown below: (16) in, The action superposition state is initialized; For possible actions; Step S3-2-2: Construct a U-transformation, defined as a quantum unitary transform. ,Keep ,but It will flip the sign of any vector orthogonal basis; when U a Acting on any vector, maintaining The structure, but it will flip. Orthogonal basis symbols; Step S3-2-3: Use Grover operation, using the i-th state. replace Construction transformation Define a unitary matrix as U Grov =U a U ai Through Repeated application of U Grov The i-th action The amplitude probability of one action is increased, while the amplitude of other actions is reduced. Step S3-2-4, the initial action f(s) can be expressed as: (17) In Formula 16, an angle θ is defined such that sinθ = 1 / Established at this time , and repeat the same for Perform Grover operation ( L times, represented as: (18) Step S3-2-5: By repeating the Grover operation, the probability amplitude of the corresponding action is enhanced according to the reward value, so that any state can be transitioned to another specified ground state with a high probability; execute the quantum action set. For a given action, observe all new states, update the policy and Q value using formula (10), explore the next action, and repeat the Grover iteration operation L times to update the probability amplitude until all states approach 0Hz.

8. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, In step S3, the number of iterations of quantum reinforcement learning is determined by the learning control precision and the corresponding reward value.

Citation Information

Patent Citations

  • Intelligent power grid voltage control method based on quantum multi-agent reinforcement learning

    CN118868113A