Energy storage charging and discharging multi-objective optimization method and system based on reinforcement learning
By using a reinforcement learning-based energy storage charging and discharging optimization method, combined with LSTM and deep deterministic policy gradient algorithm, and dynamically adjusting weights, the flexibility and reliability issues of microgrid energy storage systems in irregular environments are solved, achieving more efficient power management and system security.
Patent Information
- Application Number
- CN202511267803.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-12
AI Technical Summary
Existing microgrid energy storage systems have low flexibility and reliability when facing irregular weather changes and islanding mode switching, and there is a risk of delay during the switching process, which may lead to the risk of load power loss.
A multi-objective optimization method for energy storage charging and discharging based on reinforcement learning is adopted. By acquiring historical and current operating data, the weights are dynamically adjusted using LSTM and fuzzy controller, combined with a deep deterministic policy gradient algorithm, to optimize the charging and discharging trajectory in real time. In the event of islanding, an islanding control strategy is implemented to improve the flexibility and reliability of the system.
It enables flexible and reliable charge and discharge regulation of energy storage systems in complex environments, reduces switching delay, extends battery life, and improves system safety and reliability.
Smart Images

Figure CN121124152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of microgrid control. More specifically, this invention relates to a multi-objective optimization method and system for energy storage charging and discharging based on reinforcement learning. Background Technology
[0002] A microgrid is a small, independently operating distributed energy system that achieves localized energy self-sufficiency by integrating multiple power generation modes and intelligent control technologies. Because microgrids primarily rely on solar and wind power generation, supplemented by grid connection, their renewable energy supply is intermittent, causing significant fluctuations in grid load during morning and evening hours. If the microgrid's output power is insufficient to maintain normal load operation, it must switch from islanded mode to grid-connected mode to obtain energy compensation from the high-voltage grid. To ensure a stable energy supply for the microgrid, energy storage systems are often installed. These systems absorb or release electrical energy through rapid charging and discharging during sudden changes in photovoltaic or wind power output or grid connection switching, providing a smooth power curve and maintaining stable load operation.
[0003] However, current microgrid energy storage systems largely rely on static weighting strategies to adjust charging and discharging power during sudden changes and switching scenarios. While this approach offers some reliability for regular morning and evening load fluctuations in the grid, it often lags behind actual demand and lacks flexibility in the face of unpredictable weather changes. Furthermore, during islanded mode switching, the time from fault detection to mode switching can be several to tens of seconds, potentially leading to load power loss. The lack of scenario-based prediction mechanisms prevents advance parameter adjustments and state preparation, further exacerbating the uncontrollability of the switching process. These technical shortcomings severely restrict the reliable operation of microgrids under complex conditions.
[0004] Therefore, the existing microgrid energy storage charging and discharging regulation has low flexibility and reliability. Summary of the Invention To address the technical issues of low flexibility and reliability in the regulation of energy storage charging and discharging in microgrids, this invention discloses a multi-objective optimization method and system for energy storage charging and discharging based on reinforcement learning.
[0005] In a first aspect, this invention discloses a multi-objective optimization method for energy storage charging and discharging based on reinforcement learning, comprising: Acquire historical operating data, current operating data, and driving factors affecting changes in historical operating data of the energy storage system; Historical operating data and driving factors are input into an LSTM-based fuzzy controller for training to obtain an objective function with multi-objective dynamic weights. The current running data is input into the predictive controller, and the predictive controller is solved by supervising the objective function to obtain the reference charging and discharging trajectory for a future period of time. The energy storage system is controlled to charge and discharge according to a reference charge and discharge trajectory.
[0006] Beneficial effects: The method of this invention combines a fuzzy controller and LSTM to dynamically adjust the multi-objective weights of the objective function in real time, thereby better adapting to the fluctuating environment of microgrids. Based on this, a supervisory predictive controller is used to solve the objective function, obtaining a more accurate reference charging and discharging trajectory. This reference trajectory is then used to control the charging and discharging of the energy storage system, thus achieving flexible and reliable regulation of the energy storage system's charging and discharging.
[0007] Preferably, historical operating data and driving factors are input into an LSTM-based fuzzy controller for training to obtain an objective function with multi-objective dynamic weights, including: Extract input features from historical operating data and driving factors; Based on the inference rules built into the fuzzy controller, the first weight of the input features is dynamically adjusted; Input historical running data into the LSTM module, and output the weight correction coefficients; After multiplying the first weight with the corresponding weight correction coefficient and normalizing, the multi-objective dynamic weight is obtained. Update the multi-objective dynamic weights to the objective function.
[0008] Preferably, controlling the energy storage system to charge and discharge according to a reference charge and discharge trajectory includes: The reference charging and discharging trajectory is input into the deep deterministic policy gradient algorithm for training, and the actual regulated power is output. The inverter is used to verify whether the actual regulated power meets the preset safety constraints. If so, the energy storage system is charged and discharged using the actual adjustable power.
[0009] Preferably, after controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method of the present invention further includes: Write the current running data to the time-series database via a message queue; Using a time-series database as the data source for online learning, an online learning and feedback mechanism is periodically triggered to update the loss function of the deep deterministic policy gradient algorithm.
[0010] Preferably, during the charging and discharging process of the energy storage system controlled by actual power regulation, the method of the present invention further includes: Monitor the battery temperature in the energy storage system; If the battery temperature exceeds the battery temperature threshold, reduce the charging and discharging power of the energy storage system.
[0011] Preferably, the safety constraints include at least SOC boundary constraints, power limiting constraints, and ramp rate constraints.
[0012] Preferably, after controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method of the present invention further includes: A hardware comparator is used to monitor the real-time voltage of the energy storage system; If the real-time voltage is lower than the voltage threshold, the charging and discharging of the energy storage system will be interrupted.
[0013] Preferably, after controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method of the present invention further includes: A hardware comparator is used to monitor whether there is a communication interruption signal in the energy storage system; If so, interrupt the charging and discharging of the energy storage system.
[0014] Preferably, after controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method of the present invention further includes: Determine if an island event has been triggered; If so, extract the feature vector from the currently running data; The feature vectors are matched with a pre-built fault strategy library using a pre-defined supervised learning algorithm to obtain the island control strategy. An islanded control strategy is adopted to control the charging and discharging of the energy storage system.
[0015] Secondly, the present invention discloses a multi-objective optimization system for energy storage charging and discharging based on reinforcement learning, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the multi-objective optimization method for energy storage charging and discharging based on reinforcement learning described in the first aspect is implemented.
[0016] The beneficial effects of this invention are as follows: (1) Compared with the prior art, the method of the present invention solves the problem of low flexibility and reliability of the existing microgrid energy storage charging and discharging regulation.
[0017] (2) Compared with the prior art, the method of the present invention solves the problem of high response delay when switching the microgrid island mode.
[0018] (3) Compared with the prior art, the method of the present invention can extend the life of the battery in the energy storage system. Attached Figure Description
[0019] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 This is a flowchart of the multi-objective optimization method for energy storage charging and discharging based on reinforcement learning in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the structure of the multi-objective optimization system for energy storage charging and discharging based on reinforcement learning in Embodiment 2 of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] This embodiment discloses a multi-objective optimization method and system for energy storage charging and discharging based on reinforcement learning, which is used to solve the problems of low flexibility and reliability of existing microgrid energy storage charging and discharging regulation.
[0022] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0023] Example 1 like Figure 1 As shown, this embodiment discloses a multi-objective optimization method for energy storage charging and discharging based on reinforcement learning, including: S10: Obtain historical operating data, current operating data, and driving factors affecting changes in historical operating data of the energy storage system.
[0024] In this embodiment, the energy storage system is an important component of the microgrid, possessing battery packs supporting charging and discharging, and voltage regulator modules supporting stable current output. The aforementioned historical operating data refers to core historical operating parameters of the battery pack over a past period, including historical State of Charge (SOC), historical State of Health (SOH), historical charging and discharging current, historical battery pack temperature, historical grid frequency, historical node voltage, and historical load power. Current operating data includes real-time monitored parameters such as SOC, SOH, charging and discharging current and temperature, real-time grid frequency (50±0.2Hz), node voltage (10kV±5%), and load power. Driving factors include environmental data such as typhoon warning levels, temperature, and light intensity.
[0025] In other embodiments, the types of data to be acquired can be increased or decreased according to actual circumstances.
[0026] Specifically, in actual data acquisition scenarios, the aforementioned historical operating data can be obtained through a database; the SOC, SOH, charging and discharging current, and temperature of the energy storage system can be collected in real time through a high-precision sensor network; at the same time, key indicators such as real-time frequency, node voltage, and load power on the grid side can be obtained from the SCADA system; and environmental data such as typhoon warning level, temperature, and light intensity can be obtained by connecting to the meteorological department's API interface.
[0027] Preferably, the aforementioned historical operating data, current operating data, and driving factors, as raw data, need to undergo a preprocessing process to eliminate noise and outliers, specifically: The historical operating data, current operating data, and driving factors are filtered using a moving average to eliminate high-frequency noise; then, the data is normalized to the [0,1] interval using a minimization-maximization method, the expression of which is:
[0028] In the formula, This represents the moving average filtered value of any one of the original datasets from historical data, current data, and driving factors. This represents any one of the following raw datasets: historical running data, current running data, and driving factors. express The minimum value in, express The maximum value in.
[0029] Finally, through The principle is to remove outliers, that is, to eliminate... The data points are preprocessed and converted into standardized feature vectors, providing structured input for subsequent algorithm modules. S20: Input historical running data and driving factors into an LSTM-based fuzzy controller for training to obtain an objective function with multi-objective dynamic weights.
[0030] Specifically, step S20 above includes: S21: Extract input features from historical operating data and driving factors.
[0031] In this embodiment, the input features mentioned above are the standardized feature vectors obtained after preprocessing in step S10.
[0032] In a preferred embodiment, the input features involving historical operating data include historical SOC deviation and historical grid frequency deviation, and the input features involving driving factors include electricity price volatility and weather warning level.
[0033] Furthermore, the definition of the above-mentioned electricity price volatility is:
[0034] In the formula, Indicates electricity price volatility. Indicates electricity price over time The range of change, Indicates time Electricity prices.
[0035] Furthermore, the definition of the aforementioned SOC historical deviation is as follows:
[0036] In the formula, Indicates historical deviation of SOC. In time The SOC value.
[0037] Furthermore, the definition of the aforementioned historical frequency deviation of the power grid is as follows:
[0038] In the formula, Indicates time Historical frequency deviation of the power grid Indicates time The power grid frequency.
[0039] It should be noted that the above-mentioned weather warning levels Normalization can be performed using an artificial intelligence model, taking a weight range of 0-1.
[0040] S22: Based on the inference rules built into the fuzzy controller, dynamically adjust the first weight of the input features.
[0041] It should be noted that the aforementioned fuzzy controller has several built-in inference rules, which can be added or removed according to actual needs. Taking one inference rule as an example, its logical expression can be "IF Δf_t>0.2 AND W_alert>0.7 THEN w_2=0.8". After the above input features are input into the fuzzy controller, they will be logically judged according to the inference rules. If the logical judgment meets the conditions, the corresponding first weight will be output.
[0042] S23: Input historical running data into the LSTM module and output the weight correction coefficients.
[0043] In this embodiment, the LSTM module can be used to predict the trend of power grid state changes over the next 30 minutes and output the weight correction coefficient.
[0044] In other embodiments, the above-mentioned prediction time can be adjusted according to the actual situation.
[0045] S24: Multiply the first weight with the corresponding weight correction coefficient and then normalize to obtain the multi-objective dynamic weight.
[0046] In this embodiment, the normalization algorithm described above uses the Softmax normalization algorithm.
[0047] S25: Update the multi-objective dynamic weights to the objective function.
[0048] After outputting the dynamic weights of multiple objectives, they are updated in real time to the objective / reward function used for reinforcement learning, serving as the training objective for the subsequent prediction model.
[0049] Through steps S21-S25 above, the method of this embodiment adjusts the objective function weights in real time using a fuzzy logic controller or an LSTM prediction module, which can better adapt to power grid fluctuations.
[0050] S30: Input the current running data into the predictive controller and use the objective function to supervise the predictive controller to solve for the reference charging and discharging trajectory for a future period of time.
[0051] In this embodiment, the predictive controller performs rolling optimization on a 1-hour cycle, using mixed integer quadratic programming (MIQP) to solve for the optimal charge-discharge plan within a 24-hour time window. The objective function comprises three terms: an economic cost objective function, a SOC balancing objective function, and a power smoothing objective function.
[0052] Preferably, the expression for the above-mentioned economic cost objective function is as follows: The expression for the SOC equilibrium objective function is: The expression for the power smoothing objective function is: .in, and All are adjustment coefficients. Indicates time period Electricity costs within the area Indicates time period Charge and discharge power commands for the internal energy storage system.
[0053] Correspondingly, the constraints of the above objective function include SOC boundary constraints, power limiting constraints, and ramp rate constraints. The expressions for these safety constraints are as follows:
[0054] Among the above safety constraints, Indicates the rated power. This indicates the slope rate limit.
[0055] Using the above technical solution, the solver outputs reference charge / discharge trajectories A_ref and SOC_ref, which are then sent to the short-term control layer via the OPC-UA protocol.
[0056] S40: Control the energy storage system to charge and discharge according to the reference charge and discharge trajectory.
[0057] Specifically, step S40 includes: S41: Input the reference charging and discharging trajectory into the deep deterministic policy gradient algorithm for training, and output the actual regulated power.
[0058] In this embodiment, the deep deterministic policy gradient algorithm performs real-time control with a period of 1 second. The deep deterministic policy gradient algorithm includes an Actor network and a Critic network.
[0059] Among them, Actor network Using a 3-layer fully connected network (each layer containing 256, 128, and 64 nodes respectively), the definition of the input state of the Actor network is as follows:
[0060] In the formula, All of these are dynamic weights output from step S24.
[0061] Actor network input state output action Mapped to actual regulating power :
[0062] Among them, Critic network A dual-network structure is used to suppress overestimation, and its objective function is:
[0063] In this embodiment, the Critic network experience replay pool has a capacity of 10,000 and a batch size of 128, and a priority replay mechanism (TD error) is adopted. Larger samples have higher sampling probabilities. Policy updates use Polyak averaging to ensure training stability.
[0064] S42: Use the inverter to verify whether the actual regulated power meets the preset safety constraints.
[0065] In step S42, the aforementioned safety constraints include SOC boundary constraints, power limiting constraints, and ramp rate constraints.
[0066] Specifically, the inverter receives the power command output by DDPG (the short-term execution layer in step S41). Next, safety constraint verification is performed. For the SOC boundary constraint verification, if the current SOC value is less than 20%, forced charging is performed using the actual regulated power; if the current SOC value is greater than 80%, forced discharging is performed using the actual regulated power. The expression for the power limit constraint is:
[0067] In the formula, This indicates that the current commanded power corresponds to the actual adjusted power in step S41 above. Indicates the maximum permissible power. This represents the available power, and min() represents the minimum value function.
[0068] The expression for the constraint on the gradeability limit is as follows:
[0069] In the formula, This represents the power at the previous moment. This indicates the maximum gradeability.
[0070] Specifically, the current command power, verified by validation, is modulated by PWM to drive the IGBT switching transistor, and the actual executed power... The target value is tracked by closed-loop PID control (error <1%).
[0071] S43: If so, control the energy storage system to charge and discharge using the actual regulating power.
[0072] Through steps S10-S40 described above, the method of this embodiment successfully combines a fuzzy controller and an LSTM, and utilizes this combination for dynamic adjustment of multi-objective weights. This makes the method more suitable for complex and variable microgrid environments, providing greater flexibility. Furthermore, the method of this embodiment utilizes a deep deterministic policy gradient algorithm to calculate the actual execution power and action constraints, fully considering the complexity and safety of real-world scenarios, and enabling more reliable microgrid energy storage charging and discharging control.
[0073] Furthermore, to improve the safety of the energy storage system, in step S43 above, the method of this embodiment further includes: Monitor the battery temperature in the energy storage system.
[0074] If the battery temperature exceeds the battery temperature threshold, reduce the charging and discharging power of the energy storage system.
[0075] Preferably, the battery temperature threshold can be set to 45°C. If the temperature condition is met, the charging and discharging power will be reduced to 80% of the actual power at the previous time.
[0076] Furthermore, following step S40 above, the method of this embodiment also includes an interrupt control method for handling real-time voltage or signal communication anomalies, including: S500: Employs a hardware comparator to monitor the real-time voltage of the energy storage system.
[0077] S501: If the real-time voltage is less than the voltage threshold, interrupt the charging and discharging of the energy storage system.
[0078] S502: Uses a hardware comparator to monitor whether there is a communication interruption signal in the energy storage system.
[0079] S503: If so, interrupt the charging and discharging of the energy storage system.
[0080] Specifically, when the real-time voltage is detected to be lower than the per-unit value (0.9 pu) or a communication interruption signal occurs, an interrupt request is generated within 500 ns. The corresponding FPGA logic unit performs digital filtering on the interrupt signal to eliminate signal glitches with wavelengths less than 1 μs. After confirming its validity, the event feature vector is transmitted to the policy matching module via the DMA channel, with a total delay of less than 1 ms. Simultaneously, a watchdog timer is triggered to ensure that the default safety policy is executed when the algorithm times out. Through the above technical solution, the method of this embodiment can promptly interrupt control based on abnormal signals, thereby further improving the reliability of power control in the energy storage system.
[0081] Furthermore, regarding whether a fault occurs during the handover of an isolated event, after step S40 above, the method in this embodiment further includes: S50: Determine if an island event has been triggered.
[0082] In this embodiment, whether an islanding event is triggered depends on whether the high-voltage main grid connected to the microgrid fails. If the main grid fails, an islanding event will be automatically triggered for the purpose of microgrid isolation and protection.
[0083] S60: If so, extract the feature vector extracted from the currently running data.
[0084] In this embodiment, the above feature vector can be represented as:
[0085] This represents the i-th eigenvector, used to comprehensively describe a certain state of the energy storage system; This represents the average change in voltage; The standard deviation of the frequency; This indicates the initial state of charge of the energy storage system.
[0086] S70: The feature vectors are matched with a pre-built fault strategy library using a pre-set supervised learning algorithm to obtain the island control strategy.
[0087] In this embodiment, the supervised learning algorithm used is the KNN algorithm.
[0088] Specifically, real-time matching using the improved KNN algorithm includes: First, feature weighting is performed, with voltage weighted at 0.6, frequency at 0.3, and SOC at 0.1. Then, the distance metric is calculated, and the three scenes with the smallest distance metrics are voted on. The scene with the highest vote score is selected as the actual scene, and the corresponding control strategy is invoked to complete the matching.
[0089] Upon successful matching, the corresponding control policies are loaded, including: (1) Load the critical load list; (2) Load the diesel generator starting sequence; (3) Loading energy storage power allocation scheme.
[0090] S80: Employs an islanded control strategy to control the charging and discharging of the energy storage system.
[0091] It should be noted that when using an islanded control strategy, non-critical loads must first be disconnected, and then the energy storage output power is adjusted. At this time, the diesel generator is started synchronously, using V / f control mode, and reaches rated speed within 200ms. During this period, the response lag of the diesel engine is compensated by dynamically adjusting the energy storage power to maintain a frequency deviation of less than 0.5Hz.
[0092] It should be further noted that all execution statuses are displayed in real time through the HMI interface and recorded in the event log for post-event analysis.
[0093] Through the technical improvements in steps S50-S80, the strategy execution cycle is shortened to the 10ms level, and the effectiveness of the strategy can be confirmed through a heartbeat mechanism.
[0094] Furthermore, after step S80 above, the method of this embodiment further includes: S900: Writes the current running data to the time series database via a message queue.
[0095] S901: Uses a time-series database as the data source for online learning, and periodically triggers online learning and feedback mechanisms to update the loss function of the deep deterministic policy gradient algorithm.
[0096] Specifically, the current running data is written to the time-series database InfluxDB in real time via the Kafka message queue, and an incremental learning process is triggered every 6 hours, which includes: (1) Remove invalid data.
[0097] (2) Perform feature engineering to construct 100+ dimension feature vectors.
[0098] (3) Fine-tune the deep deterministic policy gradient algorithm by adopting a transfer learning strategy, freezing the first two layers of the Actor network, and updating only the last two layers. During training, the Adam optimizer is used; if the lr value is 0.0001, L2 regularization is added to its loss function. A value of 0.01 was used to prevent overfitting. The updated algorithm was then released in a phased manner after A / B testing, and the online learning and feedback mechanisms further improved the reliability of system control. Simultaneously, the flexibility of the method in this embodiment can be continuously improved by updating the digital twin strategy library and adding strategies to address fault scenarios.
[0099] Example 2 like Figure 2 As shown, this embodiment discloses a multi-objective optimization system for energy storage charging and discharging based on reinforcement learning, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the multi-objective optimization method for energy storage charging and discharging based on reinforcement learning described in Embodiment 1 is implemented.
[0100] The system also includes other components well known to those skilled in the art, such as communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0101] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.
[0102] In the description of this specification, "multiple" means at least two, such as two, three or more, etc., unless otherwise expressly and specifically defined.
[0103] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.
Claims
1. A multi-objective optimization method for energy storage charging and discharging based on reinforcement learning, characterized in that, include: Acquire historical operating data, current operating data, and driving factors affecting changes in the historical operating data of the energy storage system; The historical operating data and the driving factors are input into an LSTM-based fuzzy controller for training to obtain an objective function with multi-objective dynamic weights. The current operating data is input into the predictive controller, and the predictive controller is supervised to solve the objective function to obtain the reference charging and discharging trajectory for a future period of time. The energy storage system is controlled to charge and discharge according to the reference charge and discharge trajectory.
2. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 1, characterized in that, The historical operating data and the driving factors are input into an LSTM-based fuzzy controller for training to obtain an objective function with multi-objective dynamic weights, including: Extract the input features of the historical operating data and the driving factors; Based on the inference rules built into the fuzzy controller, the first weight of the input feature is dynamically adjusted; The historical running data is input into the LSTM module, which outputs the weight correction coefficients. The first weight is multiplied by the corresponding weight correction coefficient and then normalized to obtain the multi-objective dynamic weight. The multi-objective dynamic weights are updated in the objective function.
3. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 1, characterized in that, Controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory includes: The reference charging and discharging trajectory is input into a deep deterministic policy gradient algorithm for training, and the actual adjustment power is output. The inverter is used to verify whether the actual regulated power meets the preset safety constraints. If so, the energy storage system is charged and discharged using the actual regulating power.
4. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 3, characterized in that, After controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method further includes: The current running data is written to the time-series database via a message queue; Using the time-series database as a data source for online learning, an online learning and feedback mechanism is periodically triggered to update the loss function of the deep deterministic policy gradient algorithm.
5. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 3, characterized in that, During the process of controlling the energy storage system to charge and discharge using the actual regulated power, the method further includes: Monitor the battery temperature in the energy storage system; If the battery temperature is greater than the battery temperature threshold, the charging and discharging power of the energy storage system shall be reduced.
6. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 3, characterized in that, The safety constraints include at least SOC boundary constraints, power limiting constraints, and ramp rate constraints.
7. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 1, characterized in that, After controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method further includes: A hardware comparator is used to monitor the real-time voltage of the energy storage system; If the real-time voltage is less than the voltage threshold, the charging and discharging of the energy storage system will be interrupted.
8. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 1, characterized in that, After controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method further includes: A hardware comparator is used to monitor whether there is a communication interruption signal in the energy storage system; If so, the charging and discharging of the energy storage system shall be interrupted.
9. The multi-objective optimization method for energy storage charging and discharging based on reinforcement learning according to claim 1, characterized in that, After controlling the energy storage system to charge and discharge according to the reference charge and discharge trajectory, the method further includes: Determine if an island event has been triggered; If so, extract the feature vector extracted from the currently running data; The feature vectors are matched with a pre-built fault strategy library using a preset supervised learning algorithm to obtain an island control strategy. The islanded control strategy is used to control the energy storage system to charge and discharge.
10. A multi-objective optimization system for energy storage charging and discharging based on reinforcement learning, characterized in that, It includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the multi-objective optimization method for energy storage charging and discharging based on reinforcement learning as described in any one of claims 1-9 is implemented.
Citation Information
Cited By
Optical storage integrated intelligent operation method and system based on deep learning and multi-target model prediction control
CN121395328A