A power distribution control method and a power storage charging integrated system based on reinforcement learning

By adopting a power allocation control method for photovoltaic, energy storage and charging integrated systems based on reinforcement learning, the power allocation problem between photovoltaic, energy storage and grid under complex operating conditions is solved. This method improves the dynamic response speed, operational stability and energy management efficiency of the system, and ensures grid frequency stability and battery health.

CN121689304BActive Publication Date: 2026-05-12UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UESTC (SHENZHEN) ADVANCED RES INST
Filing Date
2026-02-11
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing photovoltaic-storage-charging systems struggle to balance dynamic response speed, operational stability, and energy management efficiency under complex operating conditions, making it difficult to achieve flexible and precise power allocation between photovoltaics, energy storage, and the grid.

Method used

A power allocation control method for a photovoltaic-storage-charging integrated system based on reinforcement learning is adopted. The working parameters of the photovoltaic power generation unit and the energy storage unit are collected and analyzed in real time by a reinforcement learning power compensation agent. A deep neural network for power compensation is trained using a multi-objective weighted comprehensive reward function to dynamically adjust the power output of the energy storage unit and achieve precise power allocation between photovoltaic, energy storage and grid.

Benefits of technology

It improves the system's dynamic response speed, operational stability, and energy management efficiency under complex operating conditions, enabling flexible and precise power allocation between photovoltaics, energy storage, and the grid, protecting battery health, and maintaining grid frequency stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121689304B_ABST
    Figure CN121689304B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's light storage fills integrated system and power distribution control method, it is related to light storage fills technical field, it is difficult to take into account dynamic response speed, operating stability and energy management efficiency under complex working condition technical problem.The method comprises: based on maximum power point tracking MPPT algorithm determines the maximum power point of photovoltaic power generation unit for electric energy output;Real-time acquisition is carried out to multiple operating parameters by reinforcement learning power compensation agent, the power deviation between total output power and reference power is calculated, and proportioning treatment is carried out, and state vector is inputted jointly;Based on multiple input state vectors, multiple target weighted comprehensive reward function, obtain power compensation deep neural network;Based on real-time operating parameter, power compensation deep neural network carries out inference processing, obtains the power regulation amount of energy storage unit, enters preset working mode.The application realizes the flexible, accurate and safe distribution of power between photovoltaic, energy storage and power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical energy storage and charging technology, and in particular to an integrated optical energy storage and charging system and a power distribution control method based on reinforcement learning. Background Technology

[0002] With the rapid increase in the penetration rate of new energy sources, photovoltaic (PV) power generation is gradually becoming an important energy source for distribution networks and DC microgrids. However, PV power generation is characterized by strong volatility and high uncontrollability. Especially under the influence of factors such as shading, cloud movement, and day-night cycles, its power change rate is much faster than the response speed of traditional regulation devices. When PV is connected to the system as the main energy supply unit, its rapid fluctuations can easily lead to bus voltage deviation, inverter output instability, and load power fluctuations, thereby affecting the safe operation of the entire system. Energy storage batteries, as the system's regulation unit, play a crucial role in peak shaving, valley filling, power smoothing, and buffering load fluctuations. However, the charge and discharge state of batteries is strictly limited by the SOC (State of Charge). Especially when the SOC is close to the boundary, the adjustable range of the battery is significantly reduced. Without an effective coordinated scheduling mechanism, regulation failure can easily occur when the battery energy is insufficient or excessive, thus affecting the system's continuous power supply capability.

[0003] Existing control technologies for photovoltaic-storage-charging systems typically rely on PI (Proportional-Integral) control, fuzzy control, or rule-based decision-making. While these methods are simple in structure and easy to implement, their fixed parameters often prevent them from maintaining stable performance in the face of dynamic factors such as drastic changes in illumination, sudden load increases, and changes in battery state. For example, when illumination suddenly increases, traditional controllers lack predictive capabilities and exhibit delayed responses, leading to increased bus power deviations during transient periods. When the battery's state of charge (SOC) is in the critical region, traditional strategies, due to their limited adjustment range, are prone to causing system power balance failures. Furthermore, the system frequently switches between different operating conditions, such as transitioning from photovoltaic main supply mode to battery compensation mode, or from rapid energy storage discharge mode due to sudden load increases. Fixed-parameter control methods often struggle to maintain the continuity of control signals during mode switching, thus affecting system stability. Although MPC (Model Predictive Control) or optimization algorithms have emerged to achieve multi-objective coordinated scheduling, these methods often rely on accurate models or require extensive real-time computation, making them unsuitable for the high-dimensional, nonlinear, and multi-constraint characteristics of new energy systems. In practical engineering, due to frequent parameter changes, continuous degradation of battery health, and highly random load, control methods based on fixed mathematical models are difficult to maintain stable and effective performance over a long period of time.

[0004] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art:

[0005] Under complex operating conditions, it is difficult to balance dynamic response speed, operational stability and energy management efficiency, and it is difficult to achieve flexible and precise power allocation between photovoltaics, energy storage and the grid. Summary of the Invention

[0006] The purpose of this invention is to provide a power allocation control method and system for a photovoltaic-storage-charging integrated system based on reinforcement learning, in order to solve the technical problems existing in the prior art that make it difficult to balance dynamic response speed, operational stability, and energy management efficiency under complex operating conditions, and difficult to achieve flexible and precise power allocation between photovoltaic, energy storage, and grid. The various technical effects of the preferred solutions among the many technical solutions provided by this invention are detailed below.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] This invention provides a power allocation control method for a photovoltaic-storage-charging integrated system based on reinforcement learning, comprising the following steps: S100: Determining the maximum power point of the photovoltaic power generation unit based on the Maximum Power Point Tracking (MPPT) algorithm, and the photovoltaic power generation unit outputs electrical energy; S200: The reinforcement learning power compensation agent collects multiple operating parameters of the photovoltaic-storage-charging integrated system in real time, calculates the power deviation between the total output power and the reference power, and performs proportional processing, which together serve as the input state vector of the reinforcement learning power compensation agent; S300: Based on multiple input state vectors and a multi-objective weighted comprehensive reward function, the reinforcement learning power compensation agent trains an initial power compensation deep neural network using a reinforcement learning algorithm, and obtains a power compensation deep neural network through repeated training; S400: Based on the real-time operating parameters of the photovoltaic-storage-charging integrated system, the power compensation deep neural network performs reasoning processing on the current state of the photovoltaic-storage-charging integrated system, outputs the target power adjustment amount of the energy storage unit, and combines it with the power deviation adjustment amount to obtain the power adjustment amount of the energy storage unit, and enters a preset working mode.

[0009] Preferably, in step S300, the multi-objective weighted comprehensive reward function is: Where r0 represents the continuous error penalty term, r1 represents the discrete hierarchical reward and penalty term, r2 represents the frequency deviation penalty term, r3 represents the battery SOC protection reward and penalty term, and a, b, c, and d represent the corresponding weights of each term. .

[0010] Preferably, the continuous error penalty term r0 is a continuous error penalty term that is linearly related to the absolute value of the power deviation;

[0011] The discrete hierarchical reward and punishment item r1 is:

[0012]

[0013] This is used to strengthen the control guidance in the critical range. When the absolute value of the power deviation is small, a positive reward is given, and when the absolute value of the power deviation is large, a negative reward is given to prevent large fluctuations in the power deviation.

[0014] The frequency deviation penalty item

[0015] in, The frequency deviation under normal operating conditions is limited to ±0.2Hz, and the threshold is set to 0.05Hz. This indicates that the nominal frequency of China's power system is 50Hz. This indicates the actual grid frequency of the photovoltaic-storage-charging integrated system;

[0016] The battery SOC protection reward / penalty item r3 is as follows:

[0017]

[0018] This indicates that no penalty is imposed when the battery's state of charge (SOC) is between 20% and 80%, but a strong penalty of -10 is imposed when the battery's SOC is not between 20% and 80% to prevent overcharging or over-discharging.

[0019] Preferably, in step S100, the maximum power point tracking (MPPT) algorithm determines the maximum power point based on the perturbation observation method and drives the boost converter through a PWM signal.

[0020] Preferably, in step S200, the input state vector of the reinforcement learning power compensation agent is: ,in, This indicates the output voltage of the photovoltaic power generation unit. This represents the reference d-axis current of the inner current loop. This indicates the power deviation between the total output power and the reference power. This indicates the power percentage of the photovoltaic power generation unit. This indicates the power percentage of the energy storage unit.

[0021] Preferably, in step S400, the target power adjustment amount of the energy storage unit output by the reinforcement learning power compensation agent is: ,in It is the power regulation amount of the energy storage unit, used to compensate for the power fluctuations of the photovoltaic power generation unit, so that the photovoltaic-storage-charging integrated system can maintain accurate tracking of the total output power to the reference power.

[0022] Preferably, in step S400, the preset operating modes include a first operating mode, a second operating mode, a third operating mode, and a fourth operating mode; the first operating mode is that the output of the photovoltaic power generation unit equals the power demand of the grid, and the energy storage unit does not operate; the second operating mode is that the output of the photovoltaic power generation unit is less than the power demand of the grid, and the energy storage unit discharges to compensate for the power deficit; the third operating mode is that the output of the photovoltaic power generation unit is greater than the power demand of the grid, and the energy storage unit charges to absorb the remaining power; the fourth operating mode is that when there is no sunlight, the output of the photovoltaic power generation unit is zero, and the energy storage unit independently supplies power to meet the grid demand.

[0023] A photovoltaic-storage-charging integrated system based on reinforcement learning, which performs power allocation control through any of the above-described power allocation control methods for photovoltaic-storage-charging integrated systems based on reinforcement learning, includes a photovoltaic power generation unit, an energy storage unit, a grid interface, and a control unit. The photovoltaic power generation unit generates photovoltaic power through a photovoltaic array. The energy storage unit is used to store the electrical energy generated by the photovoltaic power generation unit or to output the stored electrical energy. The grid interface is used to transmit the electrical energy from the power generation unit and the energy storage unit to the public power grid. The control unit is used to execute the maximum power point tracking (MPPT) algorithm and to run a neural network algorithm through a reinforcement learning power compensation agent.

[0024] Preferably, the inverter control of the photovoltaic power generation unit adopts a cascaded control structure based on a synchronous reference system, with the outer loop used to stabilize the DC bus voltage and the inner loop used to track the current reference value.

[0025] Preferably, the inverter of the energy storage unit also adopts cascaded control based on a synchronous reference system, wherein the reference current of the d-axis in the synchronous reference system is generated by the sum of the power deviation and the compensation output controlled by the reinforcement learning power compensation agent, and then generated by the PI controller.

[0026] Implementing one of the above-described technical solutions of the present invention has the following advantages or beneficial effects:

[0027] This invention unifies the multi-variable, multi-objective, and multi-mode coordination process of a system into a single control framework through a reinforcement learning control strategy. It can dynamically and intelligently coordinate the power flow between photovoltaic, energy storage, and charging loads, ensuring high-precision power point tracking while effectively protecting battery health and maintaining grid frequency stability. This comprehensively improves the system's dynamic response speed, operational stability, and energy management efficiency under complex operating conditions. Simultaneously, through real-time perception and learning of the system's operating status, it dynamically outputs optimized power adjustments for the energy storage units, thereby achieving flexible, precise, and safe power allocation between photovoltaic, energy storage, and the grid. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0029] Figure 1 This is a flowchart of a power allocation control method for an integrated optical storage and charging system based on reinforcement learning, according to an embodiment of the present invention.

[0030] Figure 2 This is a schematic diagram of the reinforcement learning power compensation agent in Embodiment 1 of the present invention;

[0031] Figure 3 This is a power curve of a power allocation control method for an integrated optical storage and charging system based on reinforcement learning in different operating modes according to an embodiment of the present invention.

[0032] Figure 4 This is an embodiment of the present invention, which describes a power allocation control method for an integrated optical storage and charging system based on reinforcement learning, and the grid-connected a-phase current in different operating modes.

[0033] Figure 5 This is a schematic diagram of an integrated optical storage and charging system based on reinforcement learning, according to Embodiment 2 of the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.

[0035] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0036] To illustrate the technical solution described in this invention, specific embodiments are described below, showing only the parts related to the embodiments of this invention.

[0037] Example 1:

[0038] like Figure 1 , Figure 2As shown, this invention provides a power allocation control method for a photovoltaic-storage-charging integrated system based on reinforcement learning, comprising the following steps: S100: The maximum power point of the photovoltaic power generation unit is determined based on the Maximum Power Point Tracking (MPPT) algorithm, and the photovoltaic power generation unit outputs electrical energy. The output characteristics of the photovoltaic power generation unit are nonlinear, and its maximum power point changes with environmental factors such as light intensity and temperature. The MPPT algorithm continuously adjusts the duty cycle of the DC-DC converter to keep the system operating near the maximum power point, thereby improving the power generation efficiency of the entire photovoltaic-storage-charging integrated system. However, the MPPT algorithm cannot eliminate the impact of power fluctuations on the overall system, and power compensation by the energy storage unit is still required. S200: The reinforcement learning power compensation agent collects multiple operating parameters of the photovoltaic-storage-charging integrated system in real time, calculates the power deviation between the total output power (including the output power of both photovoltaic power generation units and energy storage units) and the reference power (i.e., the target power value that the photovoltaic-storage-charging integrated system needs to track, used to coordinate the power distribution of photovoltaic power generation units, energy storage units, and electrical equipment), and proportionalizes the actual output power and power deviation of photovoltaic power generation units and energy storage units. This transforms the complex power distribution problem into a proportional coordination problem, effectively eliminating dimensional differences, accelerating the convergence speed and stability of neural network training, and enabling the reinforcement learning power compensation agent to establish a relative relationship with the photovoltaic power generation units and energy storage units, which together serve as the input state vector of the reinforcement learning power compensation agent. This vector can fully describe the system's operating mode and power balance. S300: Based on multiple input state vectors (ensuring the diversity of training samples and improving the performance of power-compensated deep neural network models), and a multi-objective weighted comprehensive reward function (which, compared to a single-objective reward, can better consider multiple system operation indicators, such as system performance, stability, and equipment lifespan, and can achieve personalized training for different scenarios and needs through weight adjustment), the reinforcement learning power-compensated agent adopts a reinforcement learning algorithm (in this embodiment, the TD3 algorithm, Twin Delayed Deep Deterministic Policy Gradient, is preferred. In each update, the TD3 algorithm first samples a batch of data from the experience replay pool, then updates the Q-network by minimizing the mean squared error loss of the double Q-network, and finally updates the policy network at a lower frequency, optimizing the policy towards maximizing the Q-value, exhibiting good stability and sample efficiency). Figure 2The blue box represents the TD3 network structure, and the red box represents the actual logic implementation process, specifically: The reinforcement learning power compensation agent outputs actions, which are combined with a PI controller to generate the battery's d-axis reference power. This power is then used to generate PWM to control the inverter via voltage and current dual-loop control. A policy network and a value network are constructed (used to learn the agent's behavioral policy and evaluate the value of a state, respectively; i.e., the policy network and value network work together, the policy network generates actions, the value network evaluates the value of those actions, and then the policy network is updated using the policy gradient method to optimize it towards maximizing value). An initial power compensation deep neural network is then trained. During training, the following methods are used: Numerous typical operating conditions, such as varying light intensity, step load, and extreme battery conditions, enable the strategy to maintain robustness in complex environments. Through repeated training, a power compensation deep neural network is developed. As the output power of the photovoltaic (PV) power generation unit increases, the reinforcement learning power compensation agent gradually increases the battery charging power to absorb excess energy. When the PV power generation unit's output is insufficient, it autonomously selects an appropriate depth of discharge based on the battery's state of charge (SOC) to stabilize the total power output. If the PV power generation unit generates no electricity and the system is in battery-only mode, the reinforcement learning power compensation agent automatically reduces the output power to prevent the battery from entering an unsafe range. S400: Based on the real-time operating parameters of the integrated PV-storage-charging system, the power compensation deep neural network infers and processes the current state of the system, outputting the target power adjustment amount for the energy storage unit. This is combined with the power deviation adjustment amount to obtain the power adjustment amount for the energy storage unit. Power adjustment can be executed through a DC / DC converter while maintaining coordination with other control components of the system. It enters a preset operating mode, allowing for flexible switching between various operating modes according to environmental changes. This embodiment unifies the multi-variable, multi-objective, and multi-mode coordination process of the system into a single control framework through a reinforcement learning control strategy. It can dynamically and intelligently coordinate the power flow between photovoltaic, energy storage, and charging loads, ensuring high-precision power tracking while effectively protecting battery health and maintaining grid frequency stability. This comprehensively improves the system's dynamic response speed, operational stability, and energy management efficiency under complex operating conditions. Simultaneously, through real-time perception and learning of the system's operating status, it dynamically outputs optimized power adjustments for the energy storage units, thereby achieving flexible, precise, and safe power allocation between photovoltaic, energy storage, and the grid.

[0039] As an optional implementation, in step S300, the multi-objective weighted comprehensive reward function is: This is used to achieve system stability, power smoothing, and battery SOC protection. r0 represents a continuous error penalty term, linearly related to the absolute value of the power deviation, ensuring basic power point tracking performance and providing gentle guidance when errors are small. r1 represents a discrete tiered reward / penalty term, dividing power ripple into intervals and setting different levels of rewards or penalties. For example, positive rewards are given for excellent intervals, and negative rewards are applied for excessive fluctuations, thereby strengthening control guidance in critical areas. r2 represents a frequency deviation penalty term, applying a penalty when the system frequency deviation exceeds a set threshold (e.g., ±0.05Hz) to ensure power quality. r3 represents a battery SOC protection reward / penalty term, establishing a safe operating window for battery SOC (e.g., 20%~80%). Once this window is exceeded, a high-weight, strong penalty is applied to ensure battery operation safety and extend its lifespan. a, b, c, and d represent the corresponding weights of each term. For example, a=1.2, b=1.2, c=1, d=1.

[0040] As an optional implementation, a continuous error penalty term r0 is a continuous error penalty term that is linearly related to the absolute value of the power deviation. It is used to ensure power tracking in actual optical storage and charging integrated systems. It provides a gentle penalty for small errors to avoid abrupt policy changes. At the same time, it can assist the agent in exploring strategies and guide the agent to explore in the direction of reducing e1 in the early stage of reinforcement learning.

[0041] The discrete hierarchical reward and punishment term r1 is:

[0042]

[0043] To strengthen the control guidance in the critical range, positive rewards are given when the absolute value of the power deviation is small, and negative rewards are given when the absolute value of the power deviation is large, so as to prevent large fluctuations in power deviation. During training, the reference total power is randomly varied between 10KW and 20KW to encourage the power ripple to be close to 1.5% and to punish to prevent falling into the wrong direction. 3 positive rewards are given to the excellent range (≤50) to form an attraction guidance. When it is greater than 300W, negative rewards are activated to prevent large fluctuations in power. The strategy is accelerated to converge to the low error region by the obvious reward drop.

[0044] Frequency deviation penalty

[0045] in, The frequency deviation under normal operating conditions is limited to ±0.2Hz, and the threshold is set to 0.05Hz. This indicates that the nominal frequency of China's power system is 50Hz. This indicates the actual grid frequency of the photovoltaic-storage-charging integrated system.

[0046] Battery SOC protection reward / penalty item r3 is:

[0047]

[0048] This indicates that no penalty is imposed when the battery's state of charge (SOC) is between 20% and 80%, but a strong penalty of -10 is imposed when the battery's SOC is not between 20% and 80% to prevent overcharging or over-discharging.

[0049] As an optional implementation, in step S100, the maximum power point tracking (MPPT) algorithm determines the maximum power point based on the perturbation-observation method. The perturbation-observation method finds and tracks the maximum power point by periodically perturbing the operating voltage of the photovoltaic array and observing the power changes. It has low computational requirements, low hardware requirements, and is easy to implement in engineering. The boost converter is driven by a PWM signal, that is, the power switching transistor is controlled by a fixed-frequency pulse to make the inductor periodically store and release energy, and then the energy is transferred to the output terminal to obtain a higher stable voltage than the input. The PWM signal is generated by the control unit.

[0050] As an optional implementation, in step S200, the input state vector of the reinforcement learning power compensation agent is: Where S is the state space, This represents the output voltage of the photovoltaic power generation unit, from which the output power of the photovoltaic power generation unit can be calculated. The reference d-axis current, representing the inner current loop, can be used to calculate the output power of the energy storage unit, thereby enabling the power compensation agent to perceive the instantaneous dynamics of the system more deeply. This indicates the power deviation between the total output power and the reference power. This indicates the power percentage of the photovoltaic power generation unit. The power proportions of the energy storage unit, photovoltaic power generation unit, and energy storage unit are all obtained through normalization. Normalization helps accelerate neural network convergence and improve training stability. These operating parameters enhance the reinforcement learning power compensation agent's perception of multi-source coupling relationships and system dynamics.

[0051] As an optional implementation, in step S400, the target power adjustment amount of the energy storage unit output by the reinforcement learning power compensation agent is: ,in It is the power regulation amount of the energy storage unit, used to compensate for the power fluctuations of the photovoltaic power generation unit, so that the photovoltaic-storage-charging integrated system can maintain accurate tracking of the total output power to the reference power.

[0052] As an optional implementation, in step S400, the preset operating modes include a first operating mode, a second operating mode, a third operating mode, and a fourth operating mode; the first operating mode is when the output of the photovoltaic power generation unit is equal to the power demand of the grid (or a situation where it is approximately equal, i.e., the output of the photovoltaic power generation unit is slightly higher or slightly lower than the power demand of the grid), and the energy storage unit does not operate; the second operating mode is when the output of the photovoltaic power generation unit is less than the power demand of the grid, and the energy storage unit discharges to compensate for the power deficit; the third operating mode is when the output of the photovoltaic power generation unit is greater than the power demand of the grid, and the energy storage unit charges to absorb the remaining power; the fourth operating mode is when there is no sunlight, the output of the photovoltaic power generation unit is zero, and the energy storage unit supplies power independently to meet the grid demand.

[0053] This embodiment verifies the dynamic response performance of the photovoltaic-storage-charging integrated system under different photovoltaic power generation operating modes. The experimental setup is a constant reference total power of 10000W and a temperature of 25°C. By simulating changes in solar irradiance, the switching between the third, first, second, and fourth operating modes is achieved, with the specific correspondences as follows: In the third operating mode, with a solar irradiance of 1000W / m², the maximum power point of the photovoltaic power generation unit (14451.4W) exceeds the grid connection requirement, and the energy storage unit absorbs the remaining energy. In the first operating mode, with a solar irradiance of 720W / m², the maximum power point of the photovoltaic power generation unit (10538.7W) is approximately equal to the grid connection requirement. In the second operating mode, with a solar irradiance of 500W / m², the maximum power point of the photovoltaic power generation unit (7352.49W) is less than the grid connection requirement, and the energy storage unit needs to compensate for the energy. In the fourth operating mode, with a solar irradiance of 0W / m², the photovoltaic power generation output is zero, and the energy storage unit alone meets the grid connection requirement. Figure 3 , Figure 4 As shown, the power curves (green represents the photovoltaic power generation unit power curve Ppv, blue represents the actual output power curve Ptotal of the photovoltaic-storage-charging integrated system, and brown represents the energy storage unit power curve PESS) and grid-connected phase a current of the photovoltaic-storage-charging integrated system operating in different working modes are shown. The grid-connected power is 10000W, and the calculated reference phase current is 21.43A. This indicates that by using the power compensation method of this embodiment, the output of the hybrid system can track the grid-connected power demand, realizing flexible switching between different modes.

[0054] The embodiment is merely a specific example and does not indicate that this is the only way to implement the present invention.

[0055] Example 2:

[0056] A reinforcement learning-based integrated optical storage and charging system uses a reinforcement learning-based power allocation control method as described in one embodiment to perform power allocation control. Figure 5As shown, the system includes a photovoltaic (PV) power generation unit, an energy storage unit, a grid interface, and a control unit. The PV power generation unit generates electricity through a PV array, comprising the array itself, a Boost DC-DC converter (a boost-type switching power supply topology that converts a lower DC input voltage to a higher DC output voltage) for MPPT (Multi-Level Transmission Theory), and a two-level DC / AC inverter (currently the most widely used voltage source inverter topology) connected to the grid via an LCL filter (a third-order passive filter network consisting of an inverter-side inductor, a filter capacitor C, and a grid-side inductor, used to suppress high-frequency harmonics generated by PWM switching while considering size and cost). The energy storage unit stores the electrical energy generated by the PV power generation unit or outputs the stored energy. Its core components include a bidirectional DC / DC converter (switching between Boost and Buck modes for discharging and charging, respectively) and a two-level DC / AC inverter connected to the grid via an LCL filter. The grid interface is used to transmit the electrical energy from the power generation unit and energy storage unit to the public power grid. The PV power generation unit and energy storage unit share AC power. The control unit executes the Maximum Power Point Tracking (MPPT) algorithm and runs a neural network algorithm through a reinforcement learning power compensation agent. It also performs dual closed-loop control (outer voltage loop, inner current loop) and PWM signal generation. The reinforcement learning power compensation agent receives "state" and "reward" data and outputs "actions" to form closed-loop control. This embodiment can dynamically and intelligently coordinate the power flow between photovoltaic, energy storage, and charging loads, ensuring high-precision power tracking while effectively protecting battery health and maintaining grid frequency stability. This comprehensively improves the system's dynamic response speed, operational stability, and energy management efficiency under complex operating conditions.

[0057] As an optional implementation, the inverter control of the photovoltaic power generation unit adopts a cascaded control structure based on a synchronous reference system (dq coordinate system). In this cascaded control structure, the outer loop is used to suppress sudden changes in sunlight and grid disturbances, while the inner loop is used for steady-state accuracy, thereby simultaneously achieving high MPPT efficiency, low bus fluctuation, and high power quality. Specifically, the outer loop is used to stabilize the DC bus voltage, and the inner loop is used to track the current reference value.

[0058] As an optional implementation, the inverter of the energy storage unit also adopts cascaded control based on a synchronous reference frame. In this system, the reference current along the d-axis of the synchronous reference frame is generated by the sum of the power deviation and the compensation output controlled by the reinforcement learning power compensation agent, then processed by a PI controller. This allows the decisions of the reinforcement learning power compensation agent to be directly and smoothly embedded into the inner current loop, achieving precise and dynamic adjustment of the energy storage unit's power.

[0059] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A power allocation control method for an integrated optical storage and charging system based on reinforcement learning, characterized in that, Includes the following steps: S100: The maximum power point of the photovoltaic power generation unit is determined based on the maximum power point tracking (MPPT) algorithm, and the photovoltaic power generation unit outputs electrical energy. S200: The reinforcement learning power compensation agent collects multiple operating parameters of the photovoltaic storage and charging integrated system in real time, calculates the power deviation between the total output power and the reference power, and performs proportional processing, which together serve as the input state vector of the reinforcement learning power compensation agent. S300: Based on multiple input state vectors and a multi-objective weighted comprehensive reward function, the reinforcement learning power compensation agent uses a reinforcement learning algorithm to train an initial power compensation deep neural network, and obtains a power compensation deep neural network through repeated training. S400: Based on the real-time operating parameters of the photovoltaic-storage-charging integrated system, the power compensation deep neural network performs reasoning processing on the current state of the photovoltaic-storage-charging integrated system, outputs the target power adjustment amount of the energy storage unit, and combines it with the power deviation adjustment amount to obtain the power adjustment amount of the energy storage unit, and enters the preset working mode. In step S300, the multi-objective weighted comprehensive reward function is: Where r0 represents the continuous error penalty term, r1 represents the discrete hierarchical reward and penalty term, r2 represents the frequency deviation penalty term, r3 represents the battery SOC protection reward and penalty term, and a, b, c, and d represent the corresponding weights of each term. ; The continuous error penalty term r0 is a continuous error penalty term that is linearly related to the absolute value of the power deviation; The discrete hierarchical reward and punishment item r1 is: This is used to strengthen the control guidance in the critical range. When the absolute value of the power deviation is small, a positive reward is given, and when the absolute value of the power deviation is large, a negative reward is given to prevent large fluctuations in the power deviation. The frequency deviation penalty item in, The frequency deviation under normal operating conditions is limited to ±0.2Hz, and the threshold is set to 0.05Hz. This indicates that the nominal frequency of China's power system is 50Hz. This indicates the actual grid frequency of the photovoltaic-storage-charging integrated system; The battery SOC protection reward / penalty item r3 is as follows: This means that there is no penalty when the battery's state of charge (SOC) is between 20% and 80%, but a strong penalty of -10 is applied when the battery's SOC is not between 20% and 80% to prevent overcharging or over-discharging. In step S200, the input state vector of the reinforcement learning power compensation agent is: ,in, This indicates the output voltage of the photovoltaic power generation unit. This represents the reference d-axis current of the inner current loop. This indicates the power deviation between the total output power and the reference power. This indicates the power percentage of the photovoltaic power generation unit. This indicates the power percentage of the energy storage unit.

2. The power allocation control method for an integrated optical storage and charging system based on reinforcement learning according to claim 1, characterized in that, In step S100, the maximum power point tracking (MPPT) algorithm determines the maximum power point based on the perturbation and observation method and drives the boost converter through the PWM signal.

3. The power allocation control method for an integrated optical storage and charging system based on reinforcement learning according to claim 1, characterized in that, In step S400, the target power adjustment amount of the energy storage unit output by the reinforcement learning power compensation agent is: ,in It is the power regulation amount of the energy storage unit, used to compensate for the power fluctuations of the photovoltaic power generation unit, so that the photovoltaic-storage-charging integrated system can maintain accurate tracking of the total output power to the reference power.

4. The power allocation control method for an integrated optical storage and charging system based on reinforcement learning according to claim 1, characterized in that, In step S400, the preset operating modes include a first operating mode, a second operating mode, a third operating mode, and a fourth operating mode. In the first operating mode, the output of the photovoltaic power generation unit equals the power demand of the grid, and the energy storage unit does not operate. In the second operating mode, the output of the photovoltaic power generation unit is less than the power demand of the grid, and the energy storage unit discharges to compensate for the power deficit. In the third operating mode, the output of the photovoltaic power generation unit is greater than the power demand of the grid, and the energy storage unit charges to absorb the remaining power. In the fourth operating mode, when there is no sunlight, the output of the photovoltaic power generation unit is zero, and the energy storage unit supplies power independently to meet the grid demand.

5. A reinforcement learning-based integrated optical storage and charging system, characterized in that, The power allocation control method for a photovoltaic-storage-charging integrated system based on reinforcement learning, as described in any one of claims 1-4, includes a photovoltaic power generation unit, an energy storage unit, a grid interface, and a control unit. The photovoltaic power generation unit generates electricity through a photovoltaic array. The energy storage unit stores the electrical energy generated by the photovoltaic power generation unit or outputs the stored electrical energy. The grid interface transmits the electrical energy from the power generation unit and the energy storage unit to the public power grid. The control unit executes the maximum power point tracking (MPPT) algorithm and runs a neural network algorithm through a reinforcement learning power compensation agent.

6. The optical storage and charging integrated system based on reinforcement learning according to claim 5, characterized in that, The inverter control of the photovoltaic power generation unit adopts a cascaded control structure based on a synchronous reference system. The outer loop is used to stabilize the DC bus voltage, and the inner loop is used to track the current reference value.

7. The optical storage and charging integrated system based on reinforcement learning according to claim 6, characterized in that, The inverter of the energy storage unit also adopts cascaded control based on a synchronous reference system. In this system, the reference current of the d-axis is generated by the sum of the power deviation and the compensation output controlled by the reinforcement learning power compensation agent, and then passed through a PI controller.