Micro-grid frequency control method based on quantum reinforcement learning
Through quantum reinforcement learning theory, the photo-storage all-in-one model is constructed, and the frequency recovery of virtual synchronous machine frequency one-time control and quantum state coding are used to accelerate frequency recovery, which solves the problem of slow frequency response caused by intermittent photovoltaic output in isolated island microgrids, and achieves fast, stable and efficient frequency control.
Patent Information
- Application Number
- CN202510352927.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Intermittent photovoltaic output in the isolated microgrid leads to slow frequency response speed, and the traditional reinforcement learning method trains data at a slow speed, resulting in unreliable power supply.
The photoreservoir all-in-one model is constructed using quantum reinforcement learning theory, and secondary control is performed through the virtual synchronous frequency primary control and quantum reinforcement learning method, and frequency recovery is accelerated using quantum state coding and Grover algorithm.
Significantly accelerate frequency recovery speed, improve system dynamic response capabilities and frequency stability, improve energy utilization efficiency, and improve computing efficiency and system real-time performance.
Smart Images

Figure CN120300828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microgrid frequency control, and particularly relates to a microgrid frequency control method based on quantum reinforcement learning. Background Art
[0002] At present, the distributed power sources of island microgrids mainly include photovoltaic and energy storage. However, the output of photovoltaic power has intermittency and randomness, increasing the frequency uncertainty factors, reducing the anti-interference ability of traditional power grids, and affecting the stability and security of power grids to a certain extent. When a microgrid operates in island mode, since there is no large power grid to provide frequency stability support, photovoltaic power must rely on energy storage to coordinate and support the safe operation of frequency. Therefore, it is necessary to use a photovoltaic and energy storage integrated machine to control the active power of photovoltaic and energy storage to achieve frequency stability under the condition of considering the uncertainty of photovoltaic output, so that the system can operate, regulate and manage more stably.
[0003] At present, the coordinated control model of a photovoltaic and energy storage integrated machine is to model photovoltaic and energy storage separately, which can be equivalent to a DC voltage source and then connected to an inverter. The literature "Research on Photovoltaic Energy Storage Grid-Connected Control Technology Based on Virtual Synchronous Machine" proposes a frequency primary and secondary frequency modulation control strategy based on the charge and discharge characteristics of energy storage batteries. The literature "Research on Multi-Level Voltage Control of Distribution Networks Containing Distributed Photovoltaic and Island Frequency Stability" adopts VSG control in island mode to prevent power frequency fluctuations caused by large disturbances.
[0004] Experts and scholars in the field of power systems have also conducted extensive research on secondary control. The literature "Distributed Secondary Optimization Control Based on Reinforcement Learning" proposes a distributed secondary optimization control based on the in-situ feedback method of reinforcement learning for the problems of system frequency and voltage static deviation caused by the droop primary control of distributed power sources in microgrids. The literature "Frequency Coordination Control Strategy of Multi-Photovoltaic and Energy Storage Virtual Synchronous Machines Based on Reinforcement Learning" (Electric Drive) proposes a frequency coordination control strategy of multi-photovoltaic and energy storage virtual synchronous machines based on reinforcement learning.
[0005] The above-mentioned literature has the following disadvantages: When a microgrid operates in island mode disconnected from the large power grid, its frequency is determined by its own distributed power sources. However, the output of photovoltaic power is intermittent, and energy storage is required to immediately generate or absorb active power to achieve frequency stability. However, the control method based on reinforcement learning is slow in terms of training data, resulting in unreliable power supply.
[0006] Therefore, the present invention patent constructs a correlation model of a photovoltaic-storage integrated machine using the theory of quantum reinforcement learning, constructs a secondary frequency control model of the photovoltaic-storage integrated machine, and proposes a microgrid frequency control method based on quantum reinforcement learning. By paralleling the photovoltaic and energy storage and equivalently converting them into a DC voltage source to access the virtual synchronous machine frequency primary control model of the inverter, and then performing secondary control on the frequency deviation output by the inverter, the quantum reinforcement learning method is used to exponentially accelerate the control process. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology in the background art, the present invention proposes a microgrid frequency control method based on quantum reinforcement learning, which uses quantum reinforcement learning to control the output frequency of the photovoltaic-storage integrated machine, and can effectively solve the problem of slow frequency response speed of distributed power sources in an island microgrid.
[0008] To achieve the technical objectives of the present invention, the following technical solutions are adopted:
[0009] A microgrid frequency control method based on quantum reinforcement learning, comprising the following steps:
[0010] Step S1: Based on the output characteristics of the photovoltaic and energy storage, define that the photovoltaic operates at the maximum power output point and is amplified by a Boost circuit, and the energy storage realizes charge and discharge through a Buck-Boost circuit, and construct a photovoltaic-storage integrated machine model based on the coordinated output of the photovoltaic and energy storage;
[0011] Step S2: Based on the photovoltaic-storage integrated machine model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a secondary frequency control model for analysis of external disturbances;
[0012] Step S3: Analyze the secondary frequency control model established in Step S2, establish a state set, a frequency action set for the secondary control of the photovoltaic-storage integrated machine, a frequency state set, design a reward function, and use the quantum Q-learning method to update to obtain the optimal action.
[0013] Further, the method of defining that the photovoltaic operates at the maximum power output point and is amplified by a Boost circuit in Step S1 is as follows:
[0014] Step S1-1-1: Define the output current of the photovoltaic cell as follows:
[0015]
[0016] In the formula, I is the output current of the photovoltaic cell, I ph is the equivalent photocurrent, I0 is the P-N junction current, I shis the leakage current of the photovoltaic cell, q is the electron charge constant, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance value of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, and R sh is the equivalent parallel resistance value of the actual photovoltaic cell;
[0017] Step S1-1-2: Obtain its short-circuit current I SC , open-circuit voltage U OC , output P at maximum power m , voltage U at maximum power m and current I at maximum power m . Equivalently transform Equation (1). The maximum current I at this time m is:
[0018]
[0019] The open-circuit current is:
[0020]
[0021] Step S1-1-3: Combine Equation 2 and Equation 3 to obtain the undetermined coefficients of capacitors C1 and C2:
[0022]
[0023] Step S1-1-4: Determine the operating conditions of the photovoltaic equivalent circuit through the numerical value of Equation 4, and adopt the maximum power point tracking algorithm to obtain the specific value of the capacitor at maximum power;
[0024] Step S1-1-5: According to the photovoltaic port characteristic curve, equivalent the photovoltaic cell to a DC voltage source; the front stage adopts a Boost boost circuit to boost the voltage of the photovoltaic array first;
[0025] Step S1-1-6: Adjust the on and off of the switch tube S by adjusting the duty cycle to adjust the DC voltage on the photovoltaic side. The input-output relationship of this circuit is:
[0026]
[0027] In the above formula, t on , t off , T d are the conduction time, turn-off time and period of the switch tube S respectively, and E is the output voltage of the photovoltaic; the output voltage E of the photovoltaic is increased to U0 and sent to the subsequent inverter.
[0028] Further, in step S1, the energy storage is connected in parallel with the PV Boost circuit through a DC-DC circuit, and a Buck-Boost circuit is adopted for the energy storage. The method for the energy storage to achieve charge and discharge through the Buck-Boost circuit is as follows:
[0029] Step S1-2-1: By using PWM technology, the on-off of switch S1 and switch S2 can be controlled to achieve bidirectional energy flow. The DC / DC converter has two operating modes: when the energy storage battery needs to output power, that is, when discharging externally, the converter operates in the Boost mode, and the energy flow is from the lithium battery to the DC bus; when the lithium battery needs to absorb power, that is, when charging itself, the converter operates in the Buck mode, and the energy flow is from the DC bus to the energy storage battery;
[0030] Step S1-2-2: When the converter operates in the Buck mode, switch S2 is turned on and switch S1 is turned off, and the current flows from the bus to the battery, satisfying:
[0031] U bat = DU dc (6)
[0032] In the above formula, U bat is the energy storage voltage, D is the duty cycle, and U dc is the DC bus voltage.
[0033] Step S1-2-3: When the converter operates in the Boost mode, switch S1 is turned on and switch S2 is turned off, and the current flows from the battery to the bus, satisfying:
[0034] U bat = (1 - D)U dc (7)
[0035] Among them, switch S1 and switch S2 are complementary types. When one is turned on, the other is turned off. Only one control signal is required to control the on-off of the bidirectional DC / DC converter switching device, which can enable the battery to charge and discharge reasonably according to needs, and adjust the duty cycle during the charge and discharge process to change the DC bus voltage.
[0036] Further, in step S2, the primary control model based on the virtual synchronous machine realizes the distribution of active power through droop control; the expression of the primary frequency control is as follows:
[0037] f = f ref + K f (P ref - P) (8)
[0038] In the above formula, f ref , K f , P ref, P are the frequency reference value, the frequency droop control coefficient, the active power reference value, and the actual active power value respectively.
[0039] Further, the frequency secondary control model in step S2 collects the real-time magnitude of the frequency of the distributed power source and uses the central controller to send control instructions to each distributed power source; the secondary control compensates for the frequency of the microgrid to restore it to the rated value; the frequency secondary control expression is as follows:
[0040] f = f ref +K f (P ref -P)+Δf * (9)
[0041] In the above formula, Δf * is the frequency secondary control instruction.
[0042] Further, in step S3, quantum reinforcement learning is used for secondary frequency control; the quantum reinforcement learning algorithm represents data using quantum states:
[0043] Step S3-1-1: Encode the traditional state set and action set, convert them into quantum states, and limit the frequency to be within 50 ± 0.5 Hz;
[0044] Step S3-1-2: Keep the frequency deviation to two decimal places, and set the state set as: S f = {-0.5, -0.49, -0.48, -0.47, -0.46, -0.45, ……, -0.05, -0.04, -0.03, -0.02, -0.01, 0, 0.01, 0.02, 0.03, 0.04, 0.05, ……, 0.45, 0.46, 0.47, 0.48, 0.49, 0.50} Hz, a set of 101 elements,
[0045] The designed frequency action set for the integrated energy storage and photovoltaic system secondary control is: A f = {-0.2, -0.15, -0.1, -0.06, -0.01, -0.006, -0.002, 0, 0.002, 0.006, 0.01, 0.06, 0.1, 0.15, 0.2} Hz.
[0046]
[0047] In the above formula, c f is the reward function coefficient, and Δf is the frequency deviation;
[0048] Since quantum gates do not directly handle multiplication operations and it is relatively complex to construct a quantum multiplication circuit using quantum gates, a quantum-classical combination method is adopted during the calculation of the reward function. In quantum reinforcement learning, 4 bits are used to represent the reward function coefficient c f , and after representing the reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} as 0000, 0001, ……, 1001 respectively and assigning them to the quantum states of the state set, they are then sent to a classical computer to assign the corresponding reward function coefficients to 0000, 0001, ……, 1001 and perform multiplication calculations.
[0049] Having determined the state set, action set, and reward function, the optimal action is selected and the Q-value table is updated to store the expected cumulative rewards obtained by the agent taking different actions in different states. The frequency secondary control instruction, Q-value update, and optimal action formulas are as follows:
[0050]
[0051] Among them, Δf * is the frequency secondary control instruction, Q(s t ,a t ) is the Q-value table in the classical computer; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9, s t is the state at time t, s t+1 is the state at time t + 1, a t is the action at time t, and a is the possible action to be taken.
[0053] Having determined the state set, action set, and reward function, the optimal action is selected and the Q-value table is updated to store the expected cumulative rewards obtained by the agent taking different actions in different states. The frequency secondary control instruction, Q-value update, and optimal action formulas are as follows:
[0054]
[0055] Among them, Δf * is the frequency secondary control instruction, Q(s t ,a t ) is the Q-value table in the classical computer; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9. Further, the frequency action set designed for the secondary control of the integrated energy storage and photovoltaic system in step S3 is: A f ={-0.2, -0.15, -0.1, -0.06, -0.01, -0.006, -0.002, 0, 0.002, 0.006, 0.01, 0.06, 0.1, 0.15, 0.2}Hz.
[0056] Further, in step S3, for the state set and the frequency action set of the integrated photovoltaic and energy storage unit, amplitude encoding is used for encoding, and 7 bits and 4 bits are respectively used to represent the quantum state set |s> and the quantum action set |a>.
[0057] The quantum state set is expressed as:
[0058]
[0059] The quantum action set is expressed as:
[0060]
[0061] Among them,
[0062]
[0063] Among them, |ρ n | and |β n | are respectively the probabilities of the corresponding elements of the quantum state and the quantum action;
[0064] Further, in step S3, the Grover algorithm is used to update the probability amplitude of the quantum action set |a>; when performing amplitude update, by repeating the Grover operation, the actions that obtain high rewards are strengthened as outputs; the specific method is as follows:
[0065] Step S3-2-1: Initialize the behavior, and the initialization formula is as follows:
[0066]
[0067] Among them, |a s (n) > is the initialized action superposition state.
[0068] Step S3-2-2: Construct a U transformation, defined as Ua = 2|a s (n) ><a s (n) | - I, keeping |a s (n) >, but |a s (n) > will flip the signs of the orthogonal bases of any vector; when Ua acts on any vector, keeping |a s ( n) > structure, but will flip the signs of the orthogonal bases of |a s (n) >;
[0069] Step S3-2-3: Adopt Grover operation, and use the i-th state |a i > to replace |as (n) >, construct the transformation U ai = I - 2|a i >] <a i |, define a unitary matrix as U Grov = U a U ai , by repeatedly applying U on |a s (n)> Grov , the amplitude probability of the i-th row being |a i >] is enhanced, while the amplitudes of other rows decrease;
[0070] Step S3-2-4, the initial action f(s) can be expressed as:
[0071]
[0072] In Formula 16, define an angle θ that satisfies holds, at this time Similarly, repeat the Grover operation on |as(n)> times, expressed as:
[0073]
[0074] Step S3-2-5, by repeating the Grover operation, strengthen the probability amplitude of the corresponding action according to the reward value, so that any state can be converted to another specified ground state with a high probability.
[0075] Execute an action in the quantum action set |a>, observe all new states, update the policy and Q value with formula (10), and explore the next action, and repeat the Grover iteration operation L times to update the probability amplitude until all states approach 0Hz.
[0076] Compared with the prior art, the present invention has the following beneficial effects:
[0077] (1) The present invention performs secondary control on the microgrid frequency through the quantum reinforcement learning algorithm, which can significantly accelerate the frequency recovery speed. And compared with the traditional virtual synchronous machine control method, the present invention significantly improves the dynamic response ability of the system.
[0078] (2) Through quantum state encoding and the Grover algorithm, the present invention can perform high-precision adjustment on the frequency deviation. The design of the state set and the action set enables the system to achieve precise control within the range of ±0.5Hz, effectively avoiding the system instability problem caused by excessive frequency fluctuations.
[0079] (3) The proposed integrated photovoltaic and energy storage system model of the present invention can fully utilize the maximum power point tracking algorithm of photovoltaic and the fast charge and discharge characteristics of energy storage through the coordinated output of photovoltaic and energy storage, automatically realizing the charge and discharge of energy storage, and significantly improving the energy utilization efficiency and frequency stability of the microgrid.
[0080] (4) The present invention uses quantum reinforcement learning to accelerate the secondary frequency control, enabling the frequency of the islanded microgrid to quickly recover.
[0081] (5) Through the quantum reinforcement learning algorithm, the present invention utilizes the parallelism of quantum computing and the search acceleration characteristics of the Grover algorithm to significantly improve the computational efficiency of frequency control. Compared with traditional reinforcement learning methods, quantum reinforcement learning can find the optimal control strategy in a shorter time, improving the real-time performance and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a flowchart of a microgrid frequency control method based on quantum reinforcement learning according to the present invention;
[0083] Figure 2 is a diagram of an integrated photovoltaic and energy storage system of a microgrid frequency control method based on quantum reinforcement learning according to the present invention;
[0084] Figure 3 is a Buck-Boost charge and discharge diagram of the energy storage of a microgrid frequency control method based on quantum reinforcement learning according to the present invention;
[0085] Figure 4 is an effect diagram of the primary voltage control of the present invention;
[0086] Figure 5 is an effect diagram of the secondary frequency control of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0087] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0088] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings of the specification and specific embodiments:
[0089] Embodiment 1
[0090] As Figure 1 shown, a microgrid frequency control method based on quantum reinforcement learning includes the following steps:
[0091] Step S1: Based on the output characteristics of photovoltaic and energy storage, define that the photovoltaic operates at the maximum power output point and is amplified through a Boost circuit, and the energy storage realizes charge and discharge through a Buck - Boost circuit, and construct a photovoltaic - energy storage integrated machine model based on the coordinated output of photovoltaic and energy storage;
[0092] Step S2: Based on the photovoltaic - energy storage integrated machine model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a frequency secondary control model for analysis of external disturbances;
[0093] Step S3: Analyze the frequency secondary control model established in Step S2, establish a state set, a frequency action set for the secondary control of the photovoltaic - energy storage integrated machine, a frequency state set, design a reward function, and use the quantum Q - learning method to update and obtain the optimal action.
[0094] Furthermore, the method of defining that the photovoltaic operates at the maximum power output point and is amplified through a Boost circuit in Step S1 is as follows:
[0095] Step S1 - 1 - 1: Define the output current of the photovoltaic cell as follows:
[0096]
[0097] In the formula, I is the output current of the photovoltaic cell, I ph is the equivalent photocurrent, I0 is the P - N junction current, I sh is the leakage current of the photovoltaic cell, q is the electron charge constant, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance value of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, R sh is the equivalent parallel resistance value of the actual photovoltaic cell;
[0098] Step S1 - 1 - 2: Obtain its short - circuit current I SC , open - circuit voltage U OC , output P at maximum power m , voltage U at maximum power m and current I at maximum power m , perform equivalent conversion on Equation (1), and the maximum current I m at this time is:
[0099]
[0100] The open - circuit current is:
[0101]
[0102] Step S1-1-3: The undetermined coefficients of capacitors C1 and C2 can be obtained by combining Formula 2 and Formula 3:
[0103]
[0104] Step S1-1-4: By determining the value of Formula 4, the operating condition of the photovoltaic equivalent circuit is obtained, and the maximum power point tracking algorithm is adopted to obtain the specific value of the capacitor at the maximum power.
[0105] Step S1-1-5: According to the photovoltaic port characteristic curve, the photovoltaic cell is equivalent to a DC voltage source; a Boost boost circuit is adopted in the front stage to boost the voltage of the photovoltaic array first.
[0106] Step S1-1-6: By adjusting the duty cycle to control the on and off of the switch tube S, the DC voltage on the photovoltaic side is adjusted. The input-output relationship of this circuit is:
[0107]
[0108] In the above formula, t on and t off and T d are the conduction time, off time and period of the switch tube S respectively, and E is the output voltage of the photovoltaic; the output voltage E of the photovoltaic is increased to U0 and sent to the subsequent inverter.
[0109] Furthermore, in Step S1, the energy storage is connected in parallel with the photovoltaic Boost circuit through a DC-DC circuit, and a Buck-Boost circuit is adopted for the energy storage. The method for the energy storage to achieve charge and discharge through the Buck-Boost circuit is as follows:
[0110] Step S1-2-1: By PWM technology, the on and off of the switch tube S1 and the switch tube S2 can be controlled to achieve bidirectional energy flow. The DC / DC converter has two working modes: when the energy storage battery needs to output power, that is, when discharging externally, the converter works in the Boost mode, and the energy flow is from the lithium battery to the DC bus; when the lithium battery needs to absorb power, that is, when charging itself, the converter works in the Buck mode, and the energy flow is from the DC bus to the energy storage battery.
[0111] Step S1-2-2: When the converter works in the Buck working mode, the switch tube S2 is turned on and the switch tube S1 is turned off, and the current flows from the bus to the storage battery, satisfying:
[0112] U bat = DU dc (23)
[0113] In the above formula, U bat is the energy storage voltage, D is the duty cycle, and U dcis the DC bus voltage.
[0114] Step S1-2-3: When the converter operates in the Boost mode, switch S1 is turned on and switch S2 is turned off. The current flows from the battery to the bus, satisfying:
[0115] U bat =(1 - D)U dc (24)
[0116] Among them, switch S1 and switch S2 are complementary. When one is turned on, the other is turned off. Only one control signal is needed to control the on / off of the switching devices of the bidirectional DC / DC converter, enabling the battery to charge and discharge reasonably as needed, and adjusting the duty cycle during the charge and discharge process to change the DC bus voltage.
[0117] Furthermore, the primary control model based on the virtual synchronous machine in step S2 realizes the distribution of active power through droop control; the expression of the primary frequency control is as follows:
[0118] f = f ref + K f (P ref - P) (25)
[0119] In the above formula, f ref , K f , P ref , and P are the frequency reference value, the frequency droop control coefficient, the active power reference value, and the actual active power value respectively.
[0120] Furthermore, the secondary frequency control model in step S2 collects the real-time frequency magnitude of the distributed power sources and uses the central controller to send control instructions to each distributed power source; the secondary control compensates the frequency of the microgrid to restore it to the rated value; the expression of the secondary frequency control is as follows:
[0121] f = f ref + K f (P ref - P)+Δf * (26)
[0122] In the above formula, Δf * is the secondary frequency control instruction.
[0123] Furthermore, in step S3, quantum reinforcement learning is used for secondary frequency control; the quantum reinforcement learning algorithm represents data using quantum states:
[0124] Step S3-1-1: Encode the traditional state set and action set and convert them into quantum states, restricting the frequency to be within 50 ± 0.5 Hz;
[0125] Step S3-1-2: Keep the frequency deviation to two decimal places, and set the state set as: S f = a set of 101 elements of {-0.5, -0.49, -0.48, -0.47, -0.46, -0.45, ……, -0.05, -0.04, -0.03, -0.02, -0.01, 0, 0.01, 0.02, 0.03, 0.04, 0.05, ……, 0.45, 0.46, 0.47, 0.48, 0.49, 0.50} Hz.
[0126] Step S3-1-3: The frequency action set designed for the secondary control of the integrated energy storage and photovoltaic system is: A f = {-0.2, -0.15, -0.1, -0.06, -0.01, -0.006, -0.002, 0, 0.002, 0.006, 0.01, 0.06, 0.1, 0.15, 0.2} Hz.
[0127]
[0128] In the above formula, c f is the reward function coefficient, and Δf is the frequency deviation;
[0129] Since quantum gates do not directly handle multiplication operations and it is relatively complex to construct a quantum multiplication circuit using quantum gates, a quantum-classical combination method is adopted in the calculation of the reward function. In quantum reinforcement learning, 4 bits are used to represent the reward function coefficient, and the reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} are respectively represented as 0000, 0001, ……, 1001 and assigned to the quantum states of the state set. Then, it is sent to a classical computer to assign the corresponding reward function coefficients to 0000, 0001, ……, 1001 and perform multiplication calculations.
[0130] After determining the state set, action set, and reward function, the optimal action is selected and the Q-value table is updated to store the expected cumulative rewards obtained by the agent taking different actions in different states. The frequency secondary control instruction, updated Q-value, and optimal action formulas are as follows:
[0131]
[0132] Among them, Δf * is the frequency secondary control instruction, Q(s t ,a t ) is the Q-value table in the classical computer; α is the Q-learning rate, set to 0.1; γ is the discount factor, set to 0.9.
[0133] Further, in step S3, for the state set and the frequency action set of the integrated photovoltaic and energy storage unit, amplitude encoding is used for encoding, and 7 and 4 bits are respectively used to represent the quantum state set |s> and the quantum action set |a>.
[0134] The quantum state set is expressed as:
[0135]
[0136] The reward function coefficient is expressed as:
[0137] The quantum action set is expressed as:
[0138]
[0139] Wherein,
[0140]
[0141] Wherein, |ρ n | and |β n | are respectively the probabilities of the corresponding elements of the quantum state set and the quantum action set;
[0142] Further, in step S3, the Grover algorithm is adopted to update the probability amplitude of the quantum action set |a>; when performing amplitude update, by repeating the Grover operation, the actions with high rewards are strengthened as the output; the specific method is as follows:
[0143] Step S3-2-1: Initialize the behavior, and the action set initialization formula is as follows:
[0144]
[0145] Wherein, |a s (n) > is the initialized action superposition state.
[0146] Step S3-2-2: Construct a U transformation, defined as the quantum unitary transformation Ua = 2|a s (n) ><a s (n) | - I, keeping |a s (n) >, but |a s (n) > will flip the signs of the orthogonal bases of any vector; when Ua acts on any vector, keeping |a s (n) > structure, but will flip the signs of the orthogonal bases of |a s (n) >;
[0147] Step S3-2-3: Use the Grover operation to replace the \(i\)-th state \(|a\rangle\) with \(|a\rangle\), construct the transformation \(U = I - 2|a\rangle\langle a|\), define a unitary matrix \(U = UU\), and by repeatedly applying \(U\) on \(|a\rangle\), the amplitude probability of the \(i\)-th behavior \(|a\rangle\) is enhanced while the amplitudes of other behaviors are reduced. i > with \(|a\rangle\) s (n) >, construct the transformation \(U\) ai \(= I - 2|a\rangle\) i ><a i |, define a unitary matrix as \(U\) Grov \(= U\) a \(U\) ai By repeatedly applying \(U\) on \(|a\rangle\) s (n) >, the amplitude probability of the \(i\)-th behavior \(|a\rangle\) is enhanced while the amplitudes of other behaviors are reduced; Grov The amplitude probability of the \(i\)-th behavior \(|a\rangle\) is enhanced while the amplitudes of other behaviors are reduced; i > is enhanced while the amplitudes of other behaviors are reduced;
[0148] Step S3-2-4: The initial action \(f(s)\) can be expressed as: In Formula 16, define an angle \(\theta\) such that \(\cos\theta = \frac{1}{\sqrt{N}}\) holds. At this time, repeat the Grover operation on \(|a_{s(n)}\rangle\) \(k\) times, which is expressed as:
[0149]
[0150] In Formula 16, define an angle \(\theta\) such that holds. At this time Repeat the Grover operation on \(|a_{s(n)}\rangle\) \(k\) times, which is expressed as:
[0151]
[0152] Step S3-2-5: By repeating the Grover operation, reinforce the probability amplitudes of corresponding actions according to the reward values, so that any state can be transformed into another specified ground state with a high probability.
[0153] Execute an action in the quantum action set \(|a\rangle\), observe all new states, update the policy and Q value using Formula (10), explore the next action, and repeat the Grover iteration operation \(L\) times to update the probability amplitudes until all states approach 0 Hz.
[0154] To verify the effectiveness of the proposed solution of the present invention, a photovoltaic energy storage integrated machine control system as shown in Figure 3 is used for verification.
[0155] Embodiment 1
[0156] Establish a frequency control model for the photovoltaic energy storage integrated machine, and establish a primary frequency control model based on virtual synchronous machine control and a secondary frequency control model based on quantum reinforcement learning respectively; verify whether the proposed secondary frequency control based on quantum reinforcement learning can better and quickly restore the frequency.
[0157] The droop control coefficients of the integrated energy storage and photovoltaic system are set to 9.3e-3 respectively, the rated active power of the load is 2000 kw respectively, and a load with an active power of 1000 kw is connected at the 2nd second, a load with an active power of 1500 kw is disconnected at 2.5 seconds, and a load with an active power of 800 kw is connected at the 3rd second.
[0158] Figure 4 It is the effect diagram of the primary voltage control method based on droop control. From Figure 4 the simulation waveforms, it can be found that the effective value of the inverter output frequency fluctuates around 50 Hz, and there is a relatively obvious fluctuation after the load is connected.
[0159] On the basis of primary control, secondary control is used to quickly restore the frequency. The parameters are the same as those in Scenario 1. Figure 5 It is the effect diagram of the secondary frequency control method of the distribution network based on droop control. From Figure 5 the simulation waveforms, it can be found that the effective value of the microgrid output frequency fluctuates around 50 Hz, and the frequency can be quickly restored. Compared with primary control, secondary control has a better effect in quickly restoring the frequency.
[0160] It can be found from Table 1 that the frequency control based on quantum reinforcement learning can quickly restore the frequency fluctuation caused by load mutation well. This can also verify the practical feasibility of the present invention.
[0161] Table 1 Frequency comparison
[0162]
[0163] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A microgrid frequency control method based on quantum reinforcement learning, characterized in that, It includes the following steps: Step S1: Based on the output characteristics of photovoltaic and energy storage, define that the photovoltaic operates at the maximum power output point and is amplified through a Boost circuit, and the energy storage realizes charge and discharge through a Buck-Boost circuit, and construct a photovoltaic-storage integrated machine model based on the coordinated output of photovoltaic and energy storage; Step S2: Based on the photovoltaic-storage integrated machine model established in Step S1, construct a primary control model based on a virtual synchronous machine, and establish a frequency secondary control model for analysis of external disturbances; Step S3: Analyze the frequency secondary control model established in Step S2, establish a state set, a quantum action set, and a quantum state set for the secondary control of the photovoltaic-storage integrated machine, design a reward function, and update to obtain the optimal action using the quantum Q-learning method.
2. The microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, The method of defining that the photovoltaic operates at the maximum power output point and is amplified through a Boost circuit in Step S1 is as follows: Step S1-1-1: Define the output current of the photovoltaic cell as follows: Wherein, I is the output current of the photovoltaic cell, I ph is the equivalent photocurrent, I0 is the P-N junction current, I sh is the leakage current of the photovoltaic cell, q is the electron charge constant, U is the output voltage of the photovoltaic cell, Rs is the equivalent series resistance value of the actual photovoltaic cell, N is the diode constant, K is the Boltzmann factor, T is the surface temperature of the photovoltaic cell, R sh is the equivalent parallel resistance value of the actual photovoltaic cell; Step S1-1-2, obtain its short-circuit current I SC , open-circuit voltage U OC , output P at maximum power m , voltage U at maximum power m and current I at maximum power m , perform equivalent conversion on Equation (1), and the maximum current I m at this time is as follows: The open-circuit current is: Step S1-1-3: Combining Equation (2) and Equation (3) can obtain the undetermined coefficients of capacitors C1 and C2: Step S1-1-4: Through the numerical determination of Equation (4), obtain the operating conditions of the photovoltaic equivalent circuit, and adopt the maximum power point tracking algorithm to obtain the specific values of capacitors C1 and C2 at the maximum power; Step S1-1-5: According to the photovoltaic port characteristic curve, equivalent the photovoltaic cell to a DC voltage source; the front stage adopts a Boost boost circuit to boost the voltage of the photovoltaic array first; Step S1-1-6: By adjusting the duty cycle to control the on and off of the switch tube S, adjust the DC voltage on the photovoltaic side. The input-output relationship of this circuit is: In the above formula, t on , t off , T d are respectively the on-time, off-time and period of the switching transistor S, and E is the output voltage of the photovoltaic; the output voltage E of the photovoltaic is increased to U0 and sent to the subsequent inverter.
3. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that In Step S1, the energy storage is connected in parallel with the photovoltaic Boost circuit through a DC-DC circuit, and a Buck-Boost circuit is adopted for the energy storage. The method of realizing charge and discharge of the energy storage through the Buck-Boost circuit is as follows: Step S1-2-1: Through PWM technology, the on and off of switch tube S1 and switch tube S2 can be controlled to realize bidirectional energy flow. The DC / DC converter has two working modes: when the energy storage battery needs to output power, that is, when discharging externally, the converter works in the Boost mode, and the energy flow is from the lithium battery to the DC bus; when the lithium battery needs to absorb power, that is, when charging itself, the converter works in the Buck mode, and the energy flow is from the DC bus to the energy storage battery; Step S1-2-2: When the converter works in the Buck working mode, switch tube S2 is turned on and switch tube S1 is turned off, and the current flows from the bus to the storage battery, satisfying: U bat = DU dc (6) In the above formula, U bat is the energy storage voltage, D is the duty cycle, and U dc is the DC bus voltage; Step S1-2-3: When the converter works in the Boost working mode, switch tube S1 is turned on and switch tube S2 is turned off, and the current flows from the storage battery to the bus, satisfying: U bat = (1 - D)U dc (7) Among them, switch tube S1 and switch tube S2 are complementary types. When one is turned on, the other is turned off. Only one control signal is needed to control the on and off of the bidirectional DC / DC converter switching device, which can enable the storage battery to charge and discharge reasonably according to needs, and adjust the duty cycle during the charge and discharge process to change the DC bus voltage.
4. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that The primary control model based on the virtual synchronous machine described in step S2 realizes the distribution of active power through droop control; the expression of the primary frequency control is as follows: f g = f ref + K f (P ref - P) (8) In the above formula, f ref is the frequency reference value, K f is the frequency droop control coefficient, P ref is the active power reference value, and P is the actual active power value.
5. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that The secondary frequency control model described in step S2 collects the real-time magnitude of the frequency of distributed power sources and issues control instructions to each distributed power source by using a central controller; the secondary control compensates for the frequency of the microgrid to restore it to the rated value; the expression of the secondary frequency control is as follows: f = f ref + K f (P ref - P) + Δf * (9) In the above formula, Δf * is the frequency secondary control command.
6. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that In step S3, quantum reinforcement learning is used to perform secondary control on the frequency; the quantum reinforcement learning algorithm represents data using quantum states: Step S3-1-1: Encode the traditional state set and action set, convert them into quantum states, and limit the frequency to be within 50 ± 0.5 Hz; Step S3-1-2: Keep the frequency deviation to two decimal places and set the state set as: S f = a set of 101 elements of {-0.5, -0.49, -0.48, -0.47, -0.46, -0.45, ……, -0.05, -0.04, -0.03, -0.02, -0.01, 0, 0.01, 0.02, 0.03, 0.04, 0.05, ……, 0.45, 0.46, 0.47, 0.48, 0.49, 0.50} Hz.
7. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, The frequency action set of the secondary control of the integrated energy storage and photovoltaic system designed in step S3 is: A f = {-0.2, -0.15, -0.1, -0.06, -0.01, -0.006, -0.002, 0, 0.002, 0.006, 0.01, 0.06, 0.1, 0.15, 0.2} Hz; When calculating the reward function, a quantum-classical combination method is adopted, and the reward function is set as: In the above formula, c f is the reward function coefficient, and Δf is the frequency deviation; In quantum reinforcement learning, 4 bits are used to represent the reward function coefficient c f , and after representing the reward function coefficients {-25, -20, -15, -5, 0, 5, 15, 20, 25} as 0000, 0001, ……, 1001 respectively and assigning them to the quantum states of the state set, they are then sent to a classical computer to assign the corresponding reward function coefficients to 0000, 0001, ……, 1001 and perform multiplication calculations; The state set, action set, and reward function are determined, the optimal action is selected, and the Q-value table is updated to store the expected cumulative rewards obtained by the agent taking different actions in different states. The secondary frequency control instruction, updated Q-value, and optimal action formulas are as follows: Among them, Δf * is the frequency secondary control command, Q(s t , a t ) is the Q-value table in the classical computer; α is the Q learning rate, set to 0.1; γ is the discount factor, set to 0.9, s t is the state at time t, s t+1 is the state at time t+1, a t is the action taken at time t, and a are the possible actions.
8. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that In the described step S3, for the state set and the frequency action set of the integrated energy storage and photovoltaic system for secondary control, amplitude encoding is used for encoding. Seven and four bits are respectively used to represent the quantum state set |s> and the quantum action set |a>; The quantum state set is expressed as: The quantum action set is expressed as: Where, Among them, |ρ n | and |β n | are the probabilities of the corresponding elements of the quantum state and the quantum action, respectively.
9. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that, In step S3, the Grover algorithm is used to update the probability amplitudes of the quantum action set |a>; when performing amplitude update, by repeating the Grover operation, the actions that obtain high rewards are strengthened as outputs; the specific method is as follows: Step S3-2-1: Initialize the behavior, and the initialization formula of the quantum action set |a> is as follows: Among them, |a s (n) > is the initialized action superposition state; a is the possible action to be taken; Step S3-2-2: Construct a U transformation, defined as the quantum unitary transformation U a = 2|a s (n) ><a s (n) |-I, maintaining |a s (n) >, but |a s ( n) > will flip the signs of the orthogonal basis of any vector; when U a acts on any vector, maintaining the structure of |a s (n) >, but will flip the signs of the orthogonal basis of |a s (n) >; Step S3-2-3: Use Grover operation to replace the \(i\)-th state \(\vert a\rangle\) with \(\vert a\rangle\), construct the transformation \(U = I - 2\vert a\rangle\langle a\vert\), define a unitary matrix as \(U = UU\), and by repeatedly applying \(U\) on \(\vert a\rangle\), the amplitude probability of the \(i\)-th behavior with \(\vert a\rangle\) is enhanced while the amplitudes of other behaviors are reduced. i \(\rangle\) with \(\vert a\rangle\) s (n) \(\rangle\), construct the transformation \(U\) ai \(= I - 2\vert a\rangle\) i \(\rangle\langle a\vert\), define a unitary matrix as \(U\) i \(= U\) Grov \(U\) a \(U\) ai \(\rangle\), by repeatedly applying \(U\) on \(\vert a\rangle\) s (n) \(\rangle\), the amplitude probability of the \(i\)-th behavior with \(\vert a\rangle\) is enhanced while the amplitudes of other behaviors are reduced; Grov the \(i\)-th behavior with \(\vert a\rangle\) i \(\rangle\) is enhanced while the amplitudes of other behaviors are reduced; Step S3-2-4: The initial action f(s) can be expressed as: In Equation 16, an angle θ is defined such that holds. At this time Similarly, perform the Grover operation on |as(n)> again times, expressed as: Step S3-2-5: By repeating the Grover operation, according to the reward value, strengthen the probability amplitudes of the corresponding actions, so that any state can be converted to another specified ground state with a high probability; execute an action in the quantum action set |a>, observe all new states, update the policy and Q-value with formula (10), and explore the next action, and repeat the Grover iteration operation L times to update the probability amplitudes until all states approach 0 Hz.
10. A microgrid frequency control method based on quantum reinforcement learning according to claim 1, characterized in that In step S3, the number of iterations of the quantum reinforcement learning is determined by the learning control accuracy and the corresponding reward value.
Citation Information
Patent Citations
Control method and device based on multi-microgrid collaborative optimization, and storage medium
CN113890057A
New energy power supply control method for cigarette factory
CN116316731A
Intelligent power grid voltage control method based on quantum multi-agent reinforcement learning
CN118868113A
Novel low-voltage multi-port single-phase power supply equipment control method
CN119448274A
Optical storage all-in-one machine device
CN208955673U