Wind and light storage micro-grid control optimization method and system based on adaptive fuzzy reinforcement learning
By combining adaptive fuzzy reinforcement learning with the Actor-Critic neural network, the control strategy of the wind, solar and storage microgrid is dynamically adjusted, which solves the randomness and intermittent problems of wind and solar power generation, improves the stability and robustness of the system, and realizes the safe and stable operation of the power grid and the efficient use of energy.
Patent Information
- Application Number
- CN202510716205.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional wind, solar and storage microgrid control strategies are unable to cope with the randomness and intermittency of wind and solar power generation, resulting in power quality problems such as grid frequency offset and voltage fluctuation. In addition, existing intelligent control algorithms such as reinforcement learning lack interpretability and are difficult to be widely used in power grid systems.
An adaptive fuzzy reinforcement learning method based on the Actor-Critic neural network is adopted, combined with real-time information of the wind, solar and storage microgrid for fuzzy processing, dynamically adjusting the reinforcement learning parameters to generate an optimized control strategy. The random fluctuations of wind and light intensity are processed through fuzzy logic to enhance the robustness of the system. The energy storage system SoC is used as a multi-objective optimization indicator to prevent overcharging and discharging.
It improves the stability and ability to cope with fluctuations of wind, solar and storage microgrids, enhances the robustness and interpretability of the system, ensures the safe and stable operation of the power grid, adapts to complex environmental changes, and optimizes energy utilization and the life of energy storage equipment.
Smart Images

Figure CN120675157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of microgrids, and in particular to a wind-solar-storage microgrid control optimization method and system based on adaptive fuzzy reinforcement learning. Background Art
[0002] As the global energy transition accelerates, traditional fossil fuels face resource shortages and environmental pollution. Renewable energy sources, such as wind and solar, are increasingly contributing to the growth of wind, solar, and energy storage microgrids in the power system. Leveraging abundant wind and solar resources, wind, solar, and energy storage microgrids can achieve energy self-sufficiency and green supply. Furthermore, continuous breakthroughs in wind and photovoltaic technologies are leading to increasing power generation efficiency and decreasing costs. Energy storage technology has also made significant progress, with significant improvements in the energy density, charge and discharge efficiency, and lifespan of storage devices such as lithium batteries. Simultaneously, advances in power electronics ensure efficient conversion and coordinated control of multiple energy sources in microgrids, making the construction and operation of wind, solar, and energy storage microgrids more reliable and economical. However, wind and solar energy are characterized by significant randomness and intermittency. Their output fluctuations can easily lead to power quality issues such as frequency offset and voltage fluctuations in the grid-connected system, threatening the safe and stable operation of the power grid. In this context, the wind, solar and storage combined power generation system has become a key technology for improving the new energy absorption capacity by introducing energy storage devices, which can smooth power fluctuations, adjust the DC bus voltage, and promptly compensate for grid safety issues caused by large fluctuations in wind and light intensity.
[0003] Traditional control strategies face numerous challenges in controlling wind, solar, and energy storage microgrids. For one thing, the complex and variable characteristics of wind and solar power generation, influenced by various factors such as weather and time of day, make accurate prediction and control of generated power difficult. For example, sudden changes in wind speed or cloud cover blocking sunlight can cause significant fluctuations in generated power, impacting microgrid stability. Furthermore, the charge and discharge management of energy storage systems requires precise control to balance power generation and consumption, while also preventing overcharging or over-discharging of energy storage devices and extending their service life. Our proposed method integrates the control functions of the energy storage system's SoC to ensure grid stability and extend the life of energy storage units. Currently, many control methods are being applied to wind, solar, and energy storage microgrids. While some rule-based control strategies are simple and easy to implement, they lack the ability to adapt to complex environmental changes, making it difficult to achieve optimal system operation. Advanced intelligent control algorithms, such as reinforcement learning, can learn optimal policies through interaction with the environment, but they suffer from poor interpretability. Reinforcement learning models are often like a "black box" whose decision-making process is difficult to understand, which makes operators lack trust in the control results and is limited in practical applications, especially in systems such as power grids that require safety and stability.
[0004] Therefore, it is necessary to provide a wind-solar-storage microgrid control optimization method and system based on adaptive fuzzy reinforcement learning to improve the system stability and robustness in coping with fluctuations. Summary of the Invention
[0005] The present invention provides a wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning, including: establishing a wind-solar-storage microgrid control optimization model based on an Actor-Critic neural network; obtaining real-time change information of wind and light intensity, and performing fuzzy processing on the real-time change information of wind and light intensity; determining current reinforcement learning parameters based on the real-time change information of wind and light intensity after fuzzy processing; and generating a wind-solar-storage microgrid control optimization strategy based on the current reinforcement learning parameters through the wind-solar-storage microgrid control optimization model.
[0006] Furthermore, the system state of the wind-solar-storage microgrid control optimization model is the set of microgrid frequency and DC bus voltage, and the control quantity is the expected value of wind power generation power, the expected value of photovoltaic power generation power and the expected value of energy storage power.
[0007] Furthermore, the objective function of the wind-solar-storage microgrid control optimization model is:
[0008]
[0009] Among them, F is the objective function of the wind-solar-storage microgrid control optimization model, J(t) is the cost function, and e -λ(τ-t) is the damping factor, λ is the exponential decay rate, τ is the time integral variable, Q(x) is the system state cost, W(u) is the control output cost, Γ(x) is the disturbance impact cost, and Θ(SoC)) is the SoC cost of the energy storage system.
[0010] Furthermore, the energy storage system SoC cost is calculated based on the following function:
[0011]
[0012] Where SoC is the state of the energy storage system, μ is the mean, and σ is the standard deviation.
[0013] Furthermore, the wind-solar-storage microgrid control optimization model uses a radial basis function neural network as an activation function, and the center and width of the radial basis function neural network are determined based on historical operating data of the wind-solar-storage power station.
[0014] Furthermore, the real-time change information of wind force and light intensity is fuzzy processed, including: determining an input domain and an input fuzzy subset; determining an output domain and an output fuzzy subset; based on the input domain, normalizing the real-time change information of wind force and light intensity through a first quantization factor; based on the input fuzzy subset, fuzzy processing is performed on the normalized real-time change information of wind force and light intensity to generate fuzzy processed real-time change information of wind force and light intensity.
[0015] Furthermore, based on the real-time change information of wind force and light intensity after fuzzy processing, the current reinforcement learning parameters are determined, including: determining the input of the fuzzy controller based on the real-time change information of wind force and light intensity through a first quantization factor; determining the membership function of the fuzzy logic output through a preset fuzzy logic control rule; and converting the membership function of the fuzzy logic output into the current reinforcement learning parameters through a second quantization factor.
[0016] Furthermore, the current reinforcement learning parameters include at least a perturbation coefficient and a learning rate.
[0017] Furthermore, a wind-solar-storage microgrid control optimization model is used to generate a wind-solar-storage microgrid control optimization strategy based on the current reinforcement learning parameters, including: inputting the current reinforcement learning parameters into the wind-solar-storage microgrid control optimization model; executing the forward propagation of the Actor-Critic neural network; calculating the policy gradient and updating the Actor network parameters; updating the Critic network parameters; and generating a wind-solar-storage microgrid control optimization strategy based on the updated Actor network.
[0018] The present invention provides a wind, solar and storage microgrid control optimization system based on adaptive fuzzy reinforcement learning, including: applying the above-mentioned wind, solar and storage microgrid control optimization method based on adaptive fuzzy reinforcement learning, including: a model establishment module, used to establish a wind, solar and storage microgrid control optimization model based on an Actor-Critic neural network; an information processing module, used to obtain real-time change information of wind and light intensity, and fuzzy process the real-time change information of wind and light intensity; a parameter adjustment module, used to determine the current reinforcement learning parameters based on the real-time change information of wind and light intensity after fuzzy processing; an optimization control module, used to generate a wind, solar and storage microgrid control optimization strategy based on the current reinforcement learning parameters through the wind, solar and storage microgrid control optimization model.
[0019] Compared with the existing technology, the wind-solar-storage microgrid control optimization method and system based on adaptive fuzzy reinforcement learning provided by the present invention has at least the following beneficial effects:
[0020] By organically combining fuzzy logic and reinforcement learning, an adaptive fuzzy reinforcement learning control strategy for a wind, solar, and energy storage microgrid system was proposed. Fuzzy logic addresses random fluctuations in wind and light intensity, enhancing system robustness and improving the interpretability of the control system. Reinforcement learning dynamically adjusts the control model to obtain the optimal control strategy. This integration overcomes the limitations of traditional reinforcement learning, which suffers from poor interpretability, and fuzzy control, which lacks adaptability, representing an innovative breakthrough in control technology.
[0021] Adaptive dynamic adjustment of the learning rate and perturbation cost coefficient in reinforcement learning is achieved in conjunction with fuzzy control. Based on changes in wind speed and light intensity, the fuzzy controller adjusts reinforcement learning parameters in real time, improving the system's adaptability to complex environmental changes. In the event of sudden changes in wind speed or light intensity, the learning rate and perturbation cost coefficient are promptly increased to accelerate system response. When changes are stable, the parameters are reduced to improve the system's economic efficiency and long-term operational performance.
[0022] In addition, the SoC of the energy storage system is used as one of the multi-objective optimization indicators of reinforcement learning, taking into account the prevention of overcharging and discharging, the response to excess output in emergency situations, and the recovery of SoC during fluctuating hours. Compared with traditional methods that only focus on a single goal, this method not only ensures the safe and stable operation of the energy storage system, but also improves the overall performance and reliability of the system, which is more in line with actual application needs. Compared with traditional reinforcement learning methods, when fluctuations are large, this method can quickly adapt to the fluctuation amplitude and improve the robustness of the control system under sudden disturbances. The experimental results strongly demonstrate the robustness of the proposed control algorithm in improving system stability and the ability to cope with fluctuations, providing reliable technical support for the practical application of wind, solar and storage combined power generation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0024] Figure 1 is a schematic diagram of a single-phase grid-connected wind, solar, and energy storage combined power generation system according to some embodiments of this specification;
[0025] Figure 2 This is a flow chart of a wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to some embodiments of this specification;
[0026] Figure 3 is a schematic diagram of an iterative process of reinforcement learning based on an actor-critic network according to some embodiments of this specification;
[0027] Figure 4 is a schematic diagram of the membership of the input wind speed change according to some embodiments of this specification;
[0028] Figure 5 is a schematic diagram of the membership of an input illumination change according to some embodiments of this specification;
[0029] Figure 6 is a schematic diagram of the membership of the output disturbance coefficient according to some embodiments of this specification;
[0030] Figure 7 is a schematic diagram of the membership of the output learning rate according to some embodiments of this specification;
[0031] Figure 8 is a schematic diagram of adaptive fuzzy reinforcement learning control according to some embodiments of this specification;
[0032] Figure 9 This is a module diagram of a wind-solar-storage microgrid control system based on adaptive fuzzy reinforcement learning according to some embodiments of this specification. DETAILED DESCRIPTION
[0033] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0034] Figure 1 This is a schematic diagram of a single-phase grid-connected wind, solar, and storage combined power generation system according to some embodiments of this specification, such as Figure 1As shown, the wind, solar and storage station has two working modes: grid-connected operation and off-grid operation. In the grid-connected operation mode, its frequency and voltage are supported by the main grid; when it switches to off-grid operation mode, in order to ensure the stability of the internal system of the station, the energy storage system is used to maintain the frequency and voltage of the wind, solar and storage power station. In order to respond to the power command issued by the distribution network control center, the wind, solar and storage microgrid controller adopts a hierarchical control strategy to control the wind, solar and storage microgrid, which includes primary control, secondary control and tertiary control. In the primary control, the grid-connected controller tracks the output power of the distributed generators in the microgrid to the desired power through droop control, and integrates the output power into the DC bus, thereby achieving load power balance and voltage stability. The second-level control realizes the overall stability of the microgrid, balances the power between energy storage and load through automatic power generation control, and distributes power to the power generation equipment through the local network. In the third-level control, the grid control center adjusts the capacity according to the load power demand and the commitment uploaded by each microgrid, and distributes the desired power to the microgrid after measuring the benefits of the power market. Among them, P pre is the preset power value, P * is the actual output active power, Q pre is the preset reactive power, Q* is the actual output reactive power, u wind is the wind turbine control variable, u pv is the photovoltaic control quantity, u bat is the control quantity of the energy storage system, Q wind is the reactive power output by wind power generation, P wind is the active power output of wind power generation, Q pv is the reactive power output by photovoltaic power generation, P pv is the active power output of photovoltaic power generation, Q bat is the battery charging and discharging reactive power, P bat is the battery charging and discharging active power, Q bus is the grid-connected reactive power, P bus is the grid-connected active power.
[0035] In a wind, solar, and energy storage system, all components work together. The pitch angle controller in the wind turbine system uses the optimal tip speed ratio to adjust the blade angle in real time based on varying wind speeds, ensuring that the wind turbine consistently captures wind energy optimally and achieves maximum power point tracking (MPPT). The permanent magnet synchronous generator utilizes a vector control strategy based on rotor field orientation to precisely control the generator's current and torque, efficiently converting wind energy into electricity. The photovoltaic power generation system utilizes the conductance increment method for maximum power point tracking, monitoring the conductance and voltage changes of the photovoltaic cells in real time to quickly and accurately track the maximum power point and improve photovoltaic power generation efficiency. The energy storage system utilizes a bidirectional back-boost circuit for flexible charge and discharge control, storing energy when the system has excess power and releasing it when power is insufficient. Furthermore, the energy storage system controller integrates a system-on-chip (SoC) control function to prevent the energy storage system from being in a state of overcharge and discharge, extending the life of the energy storage units. AC bus power coordination prioritizes photovoltaic and wind power generation, prioritizing them to meet load demand. When their generation is insufficient or excessive, the energy storage system promptly compensates or absorbs power, maintaining reliable operation of the entire system. The inverter grid-connected link uses vector control based on grid voltage orientation to ensure that the output current and grid voltage are in the same frequency and phase, thereby improving the power quality.
[0036] Figure 2 This is a flow chart of a wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to some embodiments of this specification, such as Figure 2 As shown, the wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning can include the following steps.
[0037] Step 210: Establish a wind-solar-storage microgrid control optimization model based on the Actor-Critic neural network.
[0038] Figure 3 is a schematic diagram of the iterative process of reinforcement learning based on the Actor-Critic network according to some embodiments of this specification, such as Figure 3 As shown, the wind-solar-storage microgrid control optimization model is based on an actor-critic network, acquiring real-time environmental information such as wind speed, sunlight intensity, energy storage status, and load demand. This information serves as input data and provides decision-making basis for the actor and critic networks. The system state of the wind-solar-storage microgrid control optimization model is the combination of the microgrid frequency and DC bus voltage, and the control variables are the expected wind power generation, expected photovoltaic power generation, and energy storage power.
[0039] Actor decision: The Actor network uses its internal parameters and neural network structure to generate a specific control action decision based on the current environmental state. For example, when the light intensity is high, the Actor network may decide to increase the power output of the photovoltaic power generation system.
[0040] Action Execution and Environmental Feedback: The control action decisions generated by the Actor network are applied to the wind, solar, and energy storage microgrid system. After the system executes the action, the environmental state changes accordingly, generating an immediate reward J(t). At the same time, the new environmental state information is also fed back to the control system.
[0041] Critic evaluation: The Critic network receives the current state S(t), the action A(t) generated by the Actor network, the immediate reward J(t), and the next state S(t+1) as input. It uses the state value functions V(t) and V(t+1) and the discount factor γ to calculate the temporal difference (TD) error δ(t). The calculation formula is:
[0042] δ(t)=J(t+1)+γV(t+1)-V(t)
[0043] Among them, J(t+1) represents the immediate reward obtained in the next state S(t+1), V(t) and V(t+1) represent the state value functions of the current state and the next state respectively, and γ is the discount factor used to balance the importance of immediate rewards and future rewards.
[0044] Parameter Update: The actor and critic networks each update their parameters based on the calculated TD error δ(t). The actor network uses the TD error to adjust its network parameters using a gradient descent algorithm to optimize generated action decisions and achieve higher long-term rewards. The critic network then updates the parameters of its state-value function based on the TD error to improve its accuracy in assessing the value of different states.
[0045] Iterative Optimization: As the wind, solar, and energy storage microgrid system continues to operate, the above process repeats itself. The Actor-Critic neural network continuously receives environmental feedback and calculates the time-delay error. It then optimizes network parameters through gradient descent, gradually improving the power generation and energy storage control strategies. Over time and with increasing iterations, the Actor network gradually learns the optimal strategy, enabling the wind, solar, and energy storage system to achieve stable operation, efficient energy distribution, and maximized economic benefits under diverse environmental conditions.
[0046] The wind, solar, and energy storage microgrid control optimization model leverages an actor-critic neural network to make precise decisions based on real-time environmental information such as wind speed, sunlight intensity, and energy storage status. For example, when sunlight intensity suddenly changes, the actor network can quickly adjust the photovoltaic system's power output strategy. The critic network uses TD error to evaluate the effectiveness of this decision, providing optimization guidance for subsequent decisions made by the actor network. This allows the system to quickly adapt to environmental changes and ensures stable operation of the power generation system.
[0047] As the wind, solar, and energy storage microgrid system operates, the Actor-Critic neural network continuously receives environmental feedback and calculates the time-delay error. It then continuously optimizes network parameters through gradient descent, gradually improving the power generation and energy storage control strategies. This continuous iteration mechanism enables the system to continuously adapt to new environmental conditions and operational requirements, improving control performance.
[0048] By calculating the TD error, the Critic network assesses the long-term value of the Actor network's decisions, preventing them from making irrational decisions due to short-term rewards and ensuring the stability and reliability of the system's decisions. This allows the system to maintain stable operation in complex and changing energy environments.
[0049] For example, at night, when sunlight is insufficient, wind speeds are low, and loads are light, actors gradually learn to reduce photovoltaic and wind power generation and increase energy storage discharge to avoid energy waste and equipment loss while meeting load demand. Specifically, the actor network generates a control decision based on environmental information such as current light intensity, wind speed, energy storage status, and load demand. For example, it reduces the power output of the photovoltaic system to the minimum level, stops wind power generation, and initiates charging of the energy storage system. The critic network evaluates this decision, calculating the time-delay error (TD error) to determine whether it will generate high long-term rewards. If the TD error is large, indicating a potential problem with the decision, the critic network provides feedback to the actor network, guiding it to adjust parameters and optimize its decision-making strategy. After multiple iterations of learning, the actor network gradually learns the optimal strategy for rational energy allocation and efficient equipment operation at night.
[0050] In a control strategy based on adaptive fuzzy reinforcement learning, the reinforcement learning controller learns the optimal control output by minimizing a cost function. A quantitative metric is used to measure multiple key factors of the power plant system: system state, control output, disturbance impact, and the energy storage system's SoC.
[0051] As a preferred embodiment, the objective function of the wind-solar-storage microgrid control optimization model is:
[0052]
[0053] Among them, F is the objective function of the wind-solar-storage microgrid control optimization model, J(t) is the cost function, and e -λ(τ-t) is the damping factor, λ is the exponential decay rate, which prevents the disturbance ω from making J infinite, τ is the time integral variable, Q(x) is the system state cost, W(u) is the control output cost, Γ(x) is the disturbance impact cost, Θ(SoC)) is the energy storage system SoC cost, and is processed using a Gaussian function. DC ] T is the combination of microgrid frequency and DC bus voltage, u=[u wind ,u PV ,u bat ] T is the expected power of wind, solar and energy storage. T Qx / 2, W(u)=u T Wu / 2, and Q and W are set constant diagonal matrices, which are used to measure the weights of different state elements and control elements of x to achieve priority wind and photovoltaic power generation. Γ(x) is a given positive definite function, which can be set as Γ(x)=ξ||x|| 2 , used to express the upper bound of the perturbation effect. For example only, Q = [1e-4, 0; 0, 8e-5], W = [2e-5, 0, 0; 0, 2e-5, 0; 0, 5e-5].
[0054] The energy storage system SoC cost is calculated based on the following function:
[0055]
[0056] Where SoC represents the state of the energy storage system, μ represents the mean, and σ represents the standard deviation. Calculation is first performed using the state variables after closed-loop control by other controllers and historical data from the SoC. The calculation method uses the K-means clustering algorithm, which divides this historical data into n regions for classification. Each region has a center point, minimizing the sum of the Euclidean distances from all classified points to the cluster center. The cluster center of this region is the mean of the Gaussian function (the radial basis activation function) corresponding to the neuron, and the distance between different cluster centers is the variance of the Gaussian function.
[0057] When the SoC of an energy storage system is too low or too high, its cost increases rapidly, preventing the system from overcharging and discharging for extended periods, thereby protecting the energy storage units. Specifically, when an emergency anomaly occurs (such as a sudden frequency offset), Q(x) in the system performance cost function immediately increases, while the SoC term remains unchanged. Therefore, when minimizing the cost function, greater emphasis is placed on Q(x), allowing the SoC to overcharge and discharge. Subsequently, as the frequency offset decreases and all values return to normal, the SoC gradually recovers, resulting in what is known as a brief overcharge and discharge.
[0058] However, when the system issues an emergency exception, the control system will briefly overcharge and discharge to maintain system safety. By properly adjusting the parameter weights Q, W, μ, and σ in the cost function, it is possible to flexibly balance multiple objectives such as system stability, energy utilization, and energy storage life in different operating scenarios, allowing the system to better adapt to complex and changing operating environments.
[0059] The definition of the actor-critic network in reinforcement learning is as follows:
[0060]
[0061] in, is the state value estimate of the current state by the Critic network, W is the action strategy estimate generated by the Actor network based on the current state. c 、W a and Φ c (·)、Ψ a (·) are the weights and activation functions of the critic and actor networks, respectively.
[0062] Preferably, the wind-solar-storage microgrid control optimization model uses a radial basis function neural network as an activation function, and the center and width of the radial basis function neural network are determined based on the historical operating data of the wind-solar-storage power station.
[0063] Specifically, the initial weights and radial basis of the neurons are initialized. The mean and variance of the radial basis are obtained by performing k-means clustering on the historical state data to obtain its distribution characteristics, thereby initializing the center and width of the radial basis neural network activation function. In addition, the weights are initialized by using the historical data of the control signal, also using k-means clustering to obtain its distribution characteristics.
[0064] Step 220: Acquire the real-time change information of wind force and light intensity, and perform fuzzy processing on the real-time change information of wind force and light intensity.
[0065] Preferably, fuzzy processing is performed on the real-time change information of wind speed and light intensity, including:
[0066] Determine the input universe and input fuzzy subsets;
[0067] Determine the output domain and output fuzzy subset;
[0068] Based on the input domain, the real-time change information of wind speed and light intensity is normalized by the first quantization factor;
[0069] Based on the input fuzzy subset, the normalized real-time change information of wind force and light intensity is fuzzy processed to generate the fuzzy processed real-time change information of wind force and light intensity.
[0070] Specifically, the environmental factors such as wind power and light intensity faced by the wind, solar and storage combined power generation system are highly random and uncertain, and are difficult to describe with an accurate mathematical model. Fuzzy control does not require an accurate mathematical model. It fuzzifies information such as wind speed and light intensity changes and converts them into fuzzy concepts such as "large", "medium" and "small". When dealing with sudden changes in wind speed, fuzzy control can respond quickly based on fuzzy rules, adjust the system control strategy, effectively handle these uncertainties, and ensure stable operation of the system. Therefore, in order to strengthen the role of reinforcement learning in handling uncertainty in wind, solar and storage microgrids, increase its robustness in disturbed environments, and improve the interpretability of intelligent algorithms, fuzzy control methods are used for algorithm optimization. In the online learning process, the wind speed change and light intensity change z=[ΔV wind ,ΔG PV ] is used as the input of fuzzy control, and after fuzzy processing, it is used to adaptively adjust the disturbance coefficient ξ and learning rate α of reinforcement learning. The membership function of the fuzzy controller is designed to be a trigonometric function, and the input is normalized by quantization factors K1 and K2 (i.e., the first quantization factor), such as Figure 4-Figure 7 , set the input universe to [-1, 1] and divide it into five fuzzy subsets, denoted as NB (negative large), NS (negative small), Z (zero), PS (positive small), and PB (positive large). Set the output universe to [0, 1] and divide it into three fuzzy subsets, denoted as Z (zero), PS (positive small), and PB (positive large). The membership functions for the input and output are triangular membership functions.
[0071] Step 230 : Determine current reinforcement learning parameters based on the real-time change information of wind force and light intensity after fuzzy processing.
[0072] Preferably, the current reinforcement learning parameters are determined based on the real-time change information of wind force and light intensity after fuzzy processing, including:
[0073] By presetting fuzzy logic control rules and based on the real-time change information of wind speed and light intensity after fuzzy processing, the membership function of fuzzy logic output is determined;
[0074] The membership function output by the fuzzy logic is converted into current reinforcement learning parameters through the second quantization factor, wherein the current reinforcement learning parameters include at least a perturbation coefficient and a learning rate.
[0075] The preset fuzzy logic control rules can be shown in Table 1:
[0076] Table 1
[0077]
[0078] When the absolute value of the fuzzy input z is too large, ξ and α increase after fuzzy control, thereby enhancing the control system's ability to adjust the frequency during fluctuations and improving the system's robustness; when the absolute value of z is small, ξ and α take smaller values to improve the system's economic benefits and long-term operating performance.
[0079] Step 240: Generate a wind-solar-storage microgrid control optimization strategy based on the current reinforcement learning parameters using the wind-solar-storage microgrid control optimization model.
[0080] Figure 8 is a schematic diagram of adaptive fuzzy reinforcement learning control according to some embodiments of this specification, such as Figure 8 As shown, step 240 specifically includes:
[0081] Input the current reinforcement learning parameters into the wind-solar-storage microgrid control optimization model to guide the model to generate control strategies;
[0082] Execute the forward propagation of the Actor-Critic neural network. Specifically, the Actor network receives the current state information of the microgrid (such as wind power generation, photovoltaic power generation, and energy storage system status) and generates a control action (such as charging and discharging instructions for the energy storage system and power adjustment instructions for the wind turbine) through forward propagation. The Critic network receives the current state information and the action generated by the Actor network, evaluates the value of the action (i.e., the expected long-term return), and outputs a value estimate.
[0083] Calculate the policy gradient and update the Actor network parameters. Specifically, based on the value estimate output by the Critic network, calculate the gradient of the Actor network parameters. The policy gradient represents the direction and magnitude of the impact of the Actor network parameters on long-term returns. Use an optimization algorithm (such as gradient descent) to update the Actor network parameters based on the policy gradient, so that the Actor network can generate more optimal control actions, thereby improving long-term returns and optimizing the Actor network's policy so that it can select more optimal actions in a given state.
[0084] Update the parameters of the critic network. Specifically, based on the actual environmental feedback (such as the actual operation results of the microgrid) and the objective function in the reinforcement learning algorithm, calculate the target value (i.e., the ideal value estimate), and use an optimization algorithm (such as the mean square error loss function) to update the parameters of the critic network so that the value estimate output by the critic network is closer to the target value.
[0085] Based on the updated Actor network, the wind, solar and storage microgrid control optimization strategy is generated.
[0086] By continuously iterating the above steps, the Actor network and the Critic network will gradually converge to the optimal strategy, enabling the microgrid to achieve efficient and stable operation under various working conditions.
[0087] Simulation results using the MATLAB / Simulink platform demonstrate that this method effectively addresses fluctuations in sunlight and wind power. When light intensity fluctuates significantly, the energy storage system responds quickly, stabilizing the overall grid-connected output power, reducing the impact of power generation fluctuations on the grid and ensuring power supply stability. Fast Fourier Transform (FFT) analysis of the grid-connected current demonstrates that the total harmonic distortion (THD) of the grid-connected current for this combined wind, solar, and energy storage system remains low under both normal and fluctuating conditions. This demonstrates that this method effectively optimizes power quality and meets the grid's stringent power quality requirements.
[0088] Figure 9 This is a module diagram of a wind-solar-storage microgrid control system based on adaptive fuzzy reinforcement learning according to some embodiments of this specification, such as Figure 9 As shown, the wind-solar-storage microgrid control system based on adaptive fuzzy reinforcement learning can include a model building module, an information processing module, a parameter adjustment module and an optimization control module.
[0089] Model building module, used to build a wind-solar-storage microgrid control optimization model based on the Actor-Critic neural network;
[0090] An information processing module is used to obtain real-time change information of wind force and light intensity, and perform fuzzy processing on the real-time change information of wind force and light intensity;
[0091] The parameter adjustment module is used to determine the current reinforcement learning parameters based on the real-time change information of wind speed and light intensity after fuzzy processing;
[0092] The optimization control module is used to generate a wind-solar-storage microgrid control optimization strategy based on the current reinforcement learning parameters through the wind-solar-storage microgrid control optimization model.
[0093] The wind-solar-storage microgrid control system based on adaptive fuzzy reinforcement learning can be used to execute the wind-solar-storage microgrid control method based on adaptive fuzzy reinforcement learning, which will not be repeated here.
[0094] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning, characterized by: include: Based on the Actor-Critic neural network, a wind-solar-storage microgrid control optimization model is established; Obtain real-time change information of wind speed and light intensity, and perform fuzzy processing on the real-time change information of wind speed and light intensity; Determine the current reinforcement learning parameters based on the real-time change information of wind speed and light intensity after fuzzy processing; The wind-solar-storage-microgrid control optimization model is used to generate a wind-solar-storage-microgrid control optimization strategy based on the current reinforcement learning parameters.
2. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to claim 1 is characterized in that: The system state of the wind-solar-storage microgrid control optimization model is the set of microgrid frequency and DC bus voltage, and the control quantity is the expected value of wind power generation power, photovoltaic power generation power and energy storage power.
3. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to claim 2 is characterized in that: The objective function of the wind-solar-storage microgrid control optimization model is: F=min(J(t))=∫ t +∞ e -λ(τ-t) (Q(x)+W(u)+Γ(x)+Θ(SoC))dτ Among them, F is the objective function of the wind-solar-storage microgrid control optimization model, J(t) is the cost function, and e -λ(τ-t) is the damping factor, λ is the exponential decay rate, τ is the time integral variable, Q(x) is the system state cost, W(u) is the control output cost, Γ(x) is the disturbance impact cost, and Θ(SoC)) is the SoC cost of the energy storage system.
4. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to claim 3 is characterized in that: The energy storage system SoC cost is calculated based on the following function: Where SoC is the state of the energy storage system, μ is the mean, and σ is the standard deviation.
5. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to any one of claims 1 to 4, characterized in that: The wind-solar-storage microgrid control optimization model uses a radial basis function neural network as an activation function, and the center and width of the radial basis function neural network are determined based on the historical operation data of the wind-solar-storage power station.
6. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to any one of claims 1 to 4, characterized in that: Fuzzy processing of real-time changes in wind speed and light intensity, including: Determine the input universe and input fuzzy subsets; Determine the output domain and output fuzzy subset; Based on the input domain, the real-time change information of wind speed and light intensity is normalized by the first quantization factor; Based on the input fuzzy subset, the normalized real-time change information of wind force and light intensity is fuzzy processed to generate the fuzzy processed real-time change information of wind force and light intensity.
7. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to claim 6 is characterized in that: Based on the real-time change information of wind speed and light intensity after fuzzy processing, the current reinforcement learning parameters are determined, including: Determining the input of the fuzzy controller based on the real-time change information of wind force and light intensity through the first quantization factor; By presetting the fuzzy logic control rules, the membership function of the fuzzy logic output is determined; The membership function output by the fuzzy logic is converted into the current reinforcement learning parameters through the second quantization factor.
8. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to any one of claims 1 to 4, characterized in that: The current reinforcement learning parameters include at least a perturbation coefficient and a learning rate.
9. The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to any one of claims 1 to 4, characterized in that: The wind-solar-storage-microgrid control optimization model is used to generate a wind-solar-storage-microgrid control optimization strategy based on the current reinforcement learning parameters, including: Input the current reinforcement learning parameters into the wind-solar-storage microgrid control optimization model; Perform forward propagation of the Actor-Critic neural network; Calculate policy gradients and update Actor network parameters; Update the critic network parameters; Based on the updated Actor network, the wind, solar and storage microgrid control optimization strategy is generated.
10. A wind-solar-storage microgrid control optimization system based on adaptive fuzzy reinforcement learning, characterized by: The wind-solar-storage microgrid control optimization method based on adaptive fuzzy reinforcement learning according to any one of claims 1 to 9 comprises: Model building module, used to build a wind-solar-storage microgrid control optimization model based on the Actor-Critic neural network; An information processing module is used to obtain real-time change information of wind force and light intensity, and perform fuzzy processing on the real-time change information of wind force and light intensity; The parameter adjustment module is used to determine the current reinforcement learning parameters based on the real-time change information of wind speed and light intensity after fuzzy processing; The optimization control module is used to generate a wind-solar-storage microgrid control optimization strategy based on the current reinforcement learning parameters through the wind-solar-storage microgrid control optimization model.
Citation Information
Patent Citations
Frequency integrated coordination multi-rotation vector superposition control method for wind and light storage system
CN118487291A
MPPT (Maximum Power Point Tracking) control method based on dynamic optimization of lightweight convolutional neural network
CN119668355A
Energy storage grid-connected scheduling decision-making method and system based on deep reinforcement learning
CN119726663A