A method and system for optimizing and dispatching power energy storage systems based on reinforcement learning

Through the reinforcement learning-based power storage system optimization scheduling method, combined with battery health status and grid data, effective management of battery aging risks and real-time scheduling in low computing power environments are achieved, improving the system's responsiveness and energy utilization efficiency.

CN120357521BActive Publication Date: 2025-09-16BEIJING LUOHE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510827467.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-16
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing power storage systems fail to effectively assess and control the risk of battery aging during frequent charging and discharging, and find it difficult to achieve real-time scheduling requirements in low computing power environments.

Method used

A reinforcement learning-based method is used to combine battery charging status, health status, grid load and electricity price data to generate a state input data set. Pre-trained intelligent agents are used to optimize charging and discharging scheduling instructions, and power distribution and coordinated control are achieved through multi-port converters. Feedback iterative optimization is performed based on historical data.

Benefits of technology

It improves the real-time response capability and battery life management of the power energy storage system, reduces scheduling delays, improves the system's robustness and energy flow coordination, and enhances energy conversion efficiency and scheduling flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357521B_ABST
    Figure CN120357521B_ABST
Patent Text Reader

Abstract

The present application relates to the field of electrical engineering technology, and provides a method and system for optimizing and scheduling an electric energy storage system based on reinforcement learning, in order to solve the problem that it is difficult to effectively evaluate and control the risk of battery aging during frequent charging and discharging, and it is difficult to meet the real-time scheduling requirements in the low computing power environment on the edge side. The method of the present application includes: obtaining the battery charging status and health status data of the electric energy storage system and the grid load and grid electricity price data of the industrial and commercial park to generate a state input data set; using a reinforcement learning agent to obtain charging and discharging scheduling instructions; allocating power to the electric energy storage system, the multi-port converter power interface and the park load to generate multi-source collaborative power flow data; combining historical operating data, using a reinforcement learning agent to optimize the charging and discharging scheduling instructions; and optimizing and scheduling the electric energy storage system according to the optimized instructions. The present invention realizes efficient, intelligent and adaptive optimization scheduling of the electric energy storage system in the industrial and commercial park.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electrical engineering technology, and in particular to a method and system for optimizing and scheduling an electric energy storage system based on reinforcement learning. Background Art

[0002] In industrial and commercial parks and large complexes, optimizing the scheduling of power storage systems has become a key technology for improving energy efficiency, reducing electricity costs, and enhancing grid stability. With the widespread integration of diverse loads such as distributed photovoltaics and electric vehicle charging stations, park grid loads are becoming more volatile and uncertain. This necessitates an intelligent scheduling method that can dynamically respond to electricity price signals, adapt to load changes in real time, and take battery health into consideration.

[0003] Current research has proposed a scheduling method for power storage systems based on a combination of deep reinforcement learning and model predictive control. This method collects charge state data, campus load forecast information, and time-of-use electricity price curves as input states, uses deep reinforcement learning to build a scheduling decision model, and combines it with a model predictive control framework for multi-step rolling optimization to generate a forward-looking charging and discharging plan. At the same time, the system relies on a centralized computing platform for model inference, and dispatch instructions are sent to multiple independent inverters to perform power control to achieve the goal of coordinated power supply from multiple power sources. However, existing methods have some key flaws. For example, the scheduling model does not fully incorporate battery health status information, making it difficult to effectively assess and control battery aging risks during frequent charging and discharging, affecting the long-term economic operation of the energy storage system. Deep neural networks have complex structures and high inference latency, making it difficult to meet real-time scheduling requirements in low-computing environments on the edge. Summary of the Invention

[0004] The present application provides a method and system for optimizing and scheduling an electric energy storage system based on reinforcement learning, which is used to solve the problems in the prior art of difficulty in effectively evaluating and controlling the risk of battery aging during frequent charging and discharging, and difficulty in meeting the real-time scheduling requirements in a low computing power environment on the edge side.

[0005] In a first aspect, the present application provides a method for optimizing and scheduling an electric energy storage system based on reinforcement learning, comprising:

[0006] Generate a state input data set based on the battery state of charge data and battery health status data of the power energy storage system, as well as the grid load data and grid electricity price data of the industrial and commercial park;

[0007] Using a pre-trained reinforcement learning agent, the state input data set is processed to obtain a charge and discharge scheduling instruction;

[0008] According to the charge and discharge scheduling instructions, power is distributed to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park to generate multi-source collaborative power flow data;

[0009] Combining the multi-source collaborative power flow data and historical operation data, optimizing the charge and discharge scheduling instructions using the charge and discharge scheduling rules preset in the reinforcement learning agent to generate optimized charge and discharge scheduling instructions;

[0010] The power storage system is optimized and scheduled according to the optimized charge and discharge scheduling instructions.

[0011] Optionally, using a pre-trained reinforcement learning agent to process the state input data set to obtain charge and discharge scheduling instructions, including:

[0012] Using the pre-trained reinforcement learning agent, feature extraction is performed on the state input data set to obtain an initial charging power value and an initial discharging power value;

[0013] Adjusting the initial charging power value and the initial discharging power value according to a preset maximum charging power limit value of the battery, a maximum discharging power limit value of the battery, and a safe range of the battery charging state to generate a charging power value and a discharging power value;

[0014] The charging power value and the discharging power value are converted into a charging and discharging scheduling instruction, wherein the charging and discharging scheduling instruction includes a charging power allocation amount and a discharging power allocation amount.

[0015] Optionally, using the pre-trained reinforcement learning agent, feature extraction is performed on the state input data set to obtain an initial charging power value and an initial discharging power value, including:

[0016] coupling the battery charging state data and the grid load data in the state input data set to generate a battery load coupling vector;

[0017] Allocating influencing factors to the battery health status data and the grid load data according to the peak period mark and the off-peak period mark in the grid electricity price data to obtain a health load matrix;

[0018] Using the fusion feature layer of the pre-trained reinforcement learning agent, jointly encode the health load matrix and the battery load coupling vector to generate a joint feature vector;

[0019] The power derivation layer in the pre-trained reinforcement learning agent is used to perform power value regression processing on the joint feature vector to obtain an initial charging power value and an initial discharging power value.

[0020] Optionally, according to the charge and discharge scheduling instruction, power is distributed to the power storage system, the power interface of the multi-port converter, and the industrial and commercial park load to generate multi-source collaborative power flow data, including:

[0021] Determining, according to the charge and discharge scheduling instructions, the charge and discharge demand parameters of the electric energy storage system, the power transmission parameters of the power interface of the multi-port converter, and the real-time power demand parameters of the industrial and commercial park load;

[0022] According to a preset dynamic matching rule, the charging and discharging demand parameters, the power transmission parameters, and the real-time power demand parameters are collaboratively optimized to generate a power allocation plan;

[0023] According to the power allocation scheme and the preset priority rules, channel allocation is performed on the transmission amount of the power interface of the multi-port converter to generate an interface power transmission data set;

[0024] The optimized charge and discharge amount in the power allocation scheme, the actual transmission amount in the interface power transmission data set, and the actual supply amount of the industrial and commercial park load are integrated to generate multi-source collaborative power flow data.

[0025] Optionally, performing optimized scheduling of the electric energy storage system according to the optimized charge and discharge scheduling instruction includes:

[0026] parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power interface of the multi-port converter;

[0027] According to a preset time segment division rule, the target charge amount and target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount;

[0028] Sending the charge and discharge control signal to the electric energy storage system, and simultaneously sending the power adjustment instruction to the control unit of the multi-port converter to execute the scheduling of the electric energy storage system;

[0029] During the dispatching of the power storage system, the battery operating status data, the transmission data of each power interface, and the load consumption data are collected to generate an operation monitoring data set;

[0030] Deviation detection is performed on the operation monitoring data set and the expected value of the optimized charge and discharge scheduling instruction. When it is detected that the charge amount deviation exceeds a first preset threshold or the discharge amount deviation exceeds a second preset threshold, the reinforcement learning agent is triggered to regenerate the optimized charge and discharge scheduling instruction until the detection passes, so as to perform optimized scheduling closed-loop control of the power energy storage system.

[0031] Optionally, parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power supply interface of the multi-port converter includes:

[0032] Identifying an instruction header identification field and an instruction body data block in the optimized charge and discharge scheduling instruction;

[0033] According to a preset instruction version matching rule, the instruction body data block is segmented using a data structure definition template corresponding to the instruction header identification field to obtain a target charge value segment, a target discharge value segment, and a time segment identification code segment;

[0034] extracting a first value from the target charge amount value segment as a target charge amount, and extracting a second value from the target discharge amount value segment as a target discharge amount;

[0035] Determining, according to the time identification information in the time segment identification code segment, a record entry corresponding to the time segment in the interface power transmission data set;

[0036] The power transmission amount distribution value of the power interface of the multi-port converter is extracted from the record entry as the optimized distribution amount.

[0037] Optionally, according to a preset time segment division rule, the target charge amount and the target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount, including:

[0038] Splitting the target charge amount and the target discharge amount according to the preset time segment division rule to obtain charging period sub-amounts and discharging period sub-amounts corresponding to the multiple time segments;

[0039] Dynamically bind the charging period sub-quantity and the discharging period sub-quantity corresponding to each time segment with the corresponding start timestamp and end timestamp to obtain a charging control unit and a discharging control unit;

[0040] Arranging the charging control unit and the discharging control unit in chronological order to generate a charging and discharging control signal with a time stamp;

[0041] According to the optimized allocation amount, the power transmission value, identifier and time segment information of the power interface of the multi-port converter in the corresponding time segment are bound to generate a power adjustment instruction for each power interface.

[0042] In a second aspect, the present application provides a power storage system optimization and scheduling system based on reinforcement learning, comprising:

[0043] An acquisition module is used to generate a state input data set based on the battery state of charge data and battery health status data of the power energy storage system and the grid load data and grid electricity price data of the industrial and commercial park;

[0044] a processing module, configured to process the state input data set using a pre-trained reinforcement learning agent to obtain a charge and discharge scheduling instruction;

[0045] a distribution module for distributing power to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park according to the charge and discharge scheduling instructions, and generating multi-source collaborative power flow data;

[0046] an optimization module, configured to optimize the charge and discharge scheduling instructions by combining the multi-source collaborative power flow data and historical operation data and utilizing the charge and discharge scheduling rules preset in the reinforcement learning agent to generate optimized charge and discharge scheduling instructions;

[0047] The scheduling module is used to optimize the scheduling of the electric energy storage system according to the optimized charging and discharging scheduling instructions.

[0048] In a third aspect, an embodiment of the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a reinforcement learning-based power storage system optimization scheduling method as described in the first aspect above.

[0049] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a method for optimizing and scheduling an electric energy storage system based on reinforcement learning as described in the first aspect.

[0050] In the present application, a state input data set is generated based on obtaining battery charging status data and battery health status data in the electric energy storage system and grid load data and grid electricity price data of the industrial and commercial park; the state input data set is processed using a pre-trained reinforcement learning agent to obtain charge and discharge scheduling instructions; according to the charge and discharge scheduling instructions, power is allocated to the electric energy storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park to generate multi-source collaborative power flow data; in combination with the multi-source collaborative power flow data and historical operation data, the charge and discharge scheduling instructions are optimized using the charge and discharge scheduling rules preset in the reinforcement learning agent to generate optimized charge and discharge scheduling instructions; and the electric energy storage system is optimized and scheduled according to the optimized charge and discharge scheduling instructions.

[0051] Beneficial effects of this application:

[0052] (1) The technical solution provided in this application constructs an input feature space that comprehensively reflects the system operating status by integrating key parameters such as battery charging status and battery health status as well as external operating environment information, providing accurate data support for subsequent scheduling decisions.

[0053] (2) The technical solution provided in this application, with the help of a trained reinforcement learning model, can quickly output the optimal charging and discharging scheduling strategy based on the current system state, and achieve intelligent decision-making on the basis of taking into account economy, energy efficiency and battery life; it realizes the mapping from scheduling strategy to physical execution level, and through unified coordinated control of multiple energy interfaces, improves the coordination and response efficiency of energy flow within the system.

[0054] (3) The technical solution provided in this application introduces real-time power flow and historical experience data for feedback iteration, so that the scheduling strategy has dynamic adaptability, thereby improving the robustness and intelligence level of the system under complex working conditions.

[0055] (4) The technical solution provided in this application introduces the battery health status as a scheduling input factor, which makes up for the long-term operation risks brought about by the existing methods that ignore the risk of battery aging, and improves the sustainable operation capability of the energy storage system; adopts a lightweight reinforcement learning model deployed at the edge to replace the traditional high-computing power centralized inference architecture, reduces the scheduling delay, and enhances the real-time response capability of the system; relies on multi-port converters to achieve integrated power management, effectively integrates multiple energy interfaces, simplifies the system structure, improves energy conversion efficiency and scheduling flexibility, overcomes the bottleneck problems of decentralized control, high loss, and complex maintenance under the traditional multi-inverter architecture, and overall achieves a more efficient, smarter, and more adaptable power storage system optimization scheduling.

[0056] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0058] Figure 1 A flowchart of a method for optimizing and scheduling an electric energy storage system based on reinforcement learning provided by the present application is shown;

[0059] Figure 2 A schematic diagram of the structure of an electric energy storage system optimization and scheduling system based on reinforcement learning provided by the present application is shown;

[0060] Figure 3 A schematic structural diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0062] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0063] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0064] In response to the increasingly complex energy management needs in industrial and commercial parks, the existing power storage scheduling method based on the combination of deep reinforcement learning and model predictive control has exposed multiple key technical bottlenecks in actual deployment. For example, the impact of battery health status on the scheduling strategy is not fully considered, resulting in frequent charging and discharging operations that may accelerate battery aging, affecting the economy and safety of the long-term operation of the system; relying on a centralized high-computing power platform for model inference, it is difficult to meet the real-time response requirements of the edge side, limiting its application feasibility in low-computing power environments, etc. To solve the above problems, this application proposes a power storage system optimization scheduling method based on reinforcement learning. By introducing the joint state data of battery charging status and battery health status, reinforcement learning intelligent agents, and a multi-port converter integrated power management mechanism, a power storage system optimization scheduling method and system based on reinforcement learning that takes into account battery health, adapts to real-time load fluctuations, and supports edge deployment is constructed. This method improves the practicality and deployability of the system while ensuring scheduling performance, effectively filling the application gap of existing technologies in dynamic and complex scenarios of industrial and commercial parks.

[0065] Figure 1The present invention provides a flowchart of a method for optimizing and scheduling an energy storage system based on reinforcement learning, as shown in FIG. Figure 1 As shown, the method includes:

[0066] Step 101: Generate a state input data set based on battery state of charge data and battery health status data in the power energy storage system and grid load data and grid electricity price data of the industrial and commercial park.

[0067] In this step, the power storage system refers to an energy storage device consisting of a lithium-ion battery pack, a battery management system, and a power conversion system, used to store and release electrical energy. Battery charge status data refers to a set of parameters reflecting the battery's real-time charge level, including voltage (V), current (A), and temperature (°C). Battery health status data refers to indicators that characterize battery aging, including capacity decay ((initial capacity - current capacity) / initial capacity × 100%) and internal resistance growth (Ω). Grid load data refers to the total power consumption (kW) of the industrial and commercial park, recorded by equipment category (e.g., air conditioning power consumption + production equipment power consumption). Grid electricity price data refers to time-of-use pricing parameters, including peak, valley, and flat-rate unit prices (yuan / kWh) and tiered billing thresholds. The state input dataset is a two-dimensional table generated by aligning the battery status data, grid load data, and electricity price data by timestamp. Rows represent time points, and columns contain 12 feature fields.

[0068] In an embodiment of the present application, based on obtaining battery charging status data and battery health status data in the power energy storage system, as well as grid load data and grid electricity price data of the industrial and commercial park, the four types of data are integrated into a structured state input data set through timestamp alignment and dimension merging operations. The row records of the data set correspond to the same time point, and the column fields contain 12 dimensions including battery status indicators, grid parameter indicators and electricity price indicators.

[0069] Step 102: Using a pre-trained reinforcement learning agent, the state input data set is processed to obtain a charge and discharge scheduling instruction.

[0070] In this step, the pre-trained reinforcement learning agent refers to a policy network trained offline using historical data. Its neural network weights are fixed, and it takes state data as input and outputs charge and discharge action values. The charge and discharge scheduling instructions are control instructions containing the charging power allocation (positive values) and the discharging power allocation (negative values), expressed in kW.

[0071] In an embodiment of the present application, a pre-trained reinforcement learning agent is used to input the state input data set into the input layer of the agent, and feature extraction is performed through a three-layer fully connected neural network: the first layer calculates the nonlinear relationship between the battery state and the electricity price, the second layer integrates the load fluctuation characteristics, and the third layer outputs the initial charge and discharge power value; the initial value is then truncated according to the battery's maximum charging power limit value (such as ≤100kW), the battery's maximum discharge power limit value (such as ≤150kW) and the battery charging state safety range, and finally a charge and discharge scheduling instruction containing the charging power allocation and the discharge power allocation is generated.

[0072] Step 103: According to the charge and discharge scheduling instruction, power is distributed to the power storage system, the power interface of the multi-port converter, and the industrial and commercial park load to generate multi-source collaborative power flow data.

[0073] In this step, the power interface of a multi-port converter refers to the physical ports that connect to different power sources, such as the photovoltaic input interface and the grid connection interface. The industrial and commercial park load refers to the total power demand of all electrical equipment within the park, with monitoring units divided by production line. Multi-source collaborative power flow data refers to structured data packets containing power source information, transmission path encoding, load consumption, and 15-minute energy accumulation (kWh).

[0074] In an embodiment of the present application, according to the charge and discharge power allocation in the charge and discharge scheduling instruction, the charge and discharge demand parameters of the power storage system, the power transmission parameters of each power interface of the multi-port converter, and the real-time power demand parameters of the industrial and commercial park load are determined; a power balance equation is established through dynamic matching rules: photovoltaic input + grid input + battery discharge = load consumption + battery charge + system loss, and a power allocation scheme is generated after solving the optimal solution; physical channels are then allocated according to the interface priority, and finally the battery charge and discharge amount, the actual interface transmission amount and the load supply amount are integrated to generate multi-source collaborative power flow data with timestamps. For example, the battery needs to be charged 85kW, the photovoltaic interface can output 50kW, the grid interface can input 100kW, and the production line needs to consume 200kW. Through the power balance equation, it is obtained that the photovoltaic is 50kW and the grid is 150kW.

[0075] Step 104: combining the multi-source collaborative power flow data and the historical operation data, optimizing the charge and discharge scheduling instructions using the charge and discharge scheduling rules preset in the reinforcement learning agent, and generating optimized charge and discharge scheduling instructions.

[0076] In this step, historical operating data refers to past operating records for the same time period, including actual power supply output, battery charge and discharge efficiency, and load deviation. The preset charge and discharge scheduling rules refer to the policy network parameters within the reinforcement learning agent, which implement the state-to-action mapping function. The optimized charge and discharge scheduling instructions refer to the final execution instructions after historical data optimization and constraint correction.

[0077] In an embodiment of the present application, the real-time transmission parameters in the multi-source collaborative power flow data are weighted and fused with the records of the same time segment in the historical operation data to generate a reinforcement learning input data set; the preset charge and discharge scheduling rules in the reinforcement learning agent are used to perform state and action mapping on the input data set: the correlation between the load fluctuation characteristics and the electricity price trend is extracted, and the optimized charging power value and the optimized discharge power value are output; finally, the boundary constraints are corrected based on the battery charging state safety range, the converter interface capacity limit and the load demand fluctuation range to generate the optimized charge and discharge scheduling instructions. For example, the actual photovoltaic output of 48kW from 14:00 to 14:15 today and the average photovoltaic output of 52kW from 14:00 to 14:15 yesterday are weighted and fused to generate a reinforcement learning input data set, which is mapped using the policy network parameters to output the optimized charging power value +90kW and the optimized discharge power value -110kW. Finally, based on the battery charging state safety range, the converter interface capacity limit of the grid interface ≤120kW, and the load demand fluctuation range of ±15%, the boundary constraints are corrected to generate the optimized charge and discharge scheduling instructions.

[0078] Step 105: Optimize the scheduling of the power energy storage system according to the optimized charge and discharge scheduling instruction.

[0079] In an embodiment of the present application, the optimized charge and discharge scheduling instructions are parsed to obtain the target charge amount, target discharge amount and optimized distribution amount of the power interface; the target value is converted into a charge and discharge control signal with a timestamp according to a preset time segment, and a power adjustment instruction for each power interface is generated at the same time; the charge and discharge control signal is sent to the battery management system to activate the charge and discharge circuit, and the power adjustment instruction is sent to the converter control unit to switch the interface power; during execution, the battery operating status data, interface transmission data and load consumption data are monitored in real time to generate an operation monitoring data set; when it is detected that the charge amount deviation is greater than the first preset threshold or the discharge amount deviation exceeds the second preset threshold, the reinforcement learning intelligent body is triggered to generate the optimized charge and discharge scheduling instructions until the detection passes, forming a closed-loop control.

[0080] The embodiments of the present application generate high-precision inputs through multi-dimensional state data fusion, use reinforcement learning to generate preliminary scheduling instructions in real time, perform multi-source power dynamic collaborative optimization and historical data-driven instruction optimization, and use deviations to trigger self-correction execution, thereby solving the complex defects of traditional methods such as poor multi-source coordination, delayed dynamic response, and insufficient long-term optimization.

[0081] For example: the remaining battery power is 65%, the health decay rate is 12%, the total load of the park is 825kW, and the electricity price is 1.2 yuan / kWh to generate a state input data set; the reinforcement learning agent outputs a scheduling instruction of charging +100kW and discharging -150kW; the actual photovoltaic output is 120kW and the power grid purchase is 80kW to generate collaborative flow data, with photovoltaic charging of 100kW and load of 20kW and grid load of 80kW; compared with historical data, it is found that the photovoltaic output usually decays by 15%, and the optimized instruction is charging +85kW and discharging -130kW; during execution, the actual charging amount is detected to be only 78kW, which triggers the intelligent body to generate a new instruction of charging +80kW, and the final deviation is ≤2%.

[0082] This application provides a specific embodiment, step 102, using a pre-trained reinforcement learning agent to process the state input data set to obtain a charge and discharge scheduling instruction, specifically including the following steps:

[0083] Step 201: Using the pre-trained reinforcement learning agent, perform feature extraction on the state input data set to obtain an initial charging power value and an initial discharging power value.

[0084] In this step, the initial charging power value refers to the charging power recommendation value directly output by the reinforcement learning agent without considering physical constraints. The initial discharging power value refers to the discharging power recommendation value directly output by the reinforcement learning agent without considering physical constraints.

[0085] In the embodiment of the present application, the state input data set is input into a pre-trained reinforcement learning agent. The agent performs nonlinear mapping on the input features through a three-layer fully connected neural network, and finally generates an initial charging power value and an initial discharging power value in the output layer.

[0086] Step 202: According to a preset maximum charging power limit value of the battery, a maximum discharging power limit value of the battery, and a battery charging state safety range, the initial charging power value and the initial discharging power value are adjusted to generate a charging power value and a discharging power value.

[0087] In this step, the maximum battery charging power limit refers to the charging power safety threshold determined by the battery characteristics. The maximum battery discharging power limit refers to the discharge power safety threshold determined by the battery characteristics. The battery state of charge safety range refers to the safe operating range allowed by the battery state of charge. The charging power value refers to the applicable charging power value after physical constraint adjustments. The discharging power value refers to the applicable discharging power value after physical constraint adjustments.

[0088] In an embodiment of the present application, based on the preset battery maximum charging power limit value, the battery maximum discharging power limit value and the battery charging state safety range, the initial charging power value is processed using an upper truncation formula and the initial discharging power value is processed using a lower truncation formula, and finally a charging power value and a discharging power value that meet the system safety boundary are generated.

[0089] Step 203 converts the charging power value and the discharging power value into a charging and discharging scheduling instruction, wherein the charging and discharging scheduling instruction includes a charging power allocation amount and a discharging power allocation amount.

[0090] In this step, the charging power allocation refers to the normalized charging power ratio parameter. The discharging power allocation refers to the normalized discharging power ratio parameter.

[0091] In an embodiment of the present application, the adjusted charging power value is normalized and mapped to generate the charging power allocation, and the absolute value of the discharge power value is taken and normalized and mapped to generate the discharge power allocation. The two are jointly encapsulated into a charge and discharge scheduling instruction containing a timestamp.

[0092] The embodiment of the present application realizes the transformation from artificial intelligence decision-making to executable instructions, and ensures a high degree of unity between economic optimization strategy and safe operation of equipment through dual constraints.

[0093] This application provides a specific embodiment, step 201, using the pre-trained reinforcement learning agent to perform feature extraction on the state input data set to obtain an initial charging power value and an initial discharging power value, specifically including the following steps:

[0094] Step 211: Couple the battery charging state data and the grid load data in the state input data set to generate a battery load coupling vector.

[0095] In this step, the battery load coupling vector refers to a time series vector generated by linearly weighting the battery state of charge value and the grid load value, reflecting the real-time ability of the energy storage system to respond to grid demand.

[0096] In an embodiment of the present application, the battery charge state data and grid load data in the state input data set are weighted linearly superimposed in time series. Specifically, the battery charge state value is multiplied by the load change sensitivity coefficient and then added to the grid load value to generate a battery load coupling vector that reflects the real-time interaction strength between the energy storage system and the grid.

[0097] Step 212: Allocate influencing factors to the battery health status data and the grid load data according to the peak period mark and the off-peak period mark in the grid electricity price data to obtain a health load matrix.

[0098] In this step, the peak hour marker refers to the identifier of the period of high electricity demand defined in the electricity pricing mechanism, which is used to trigger the high-weight calculation of load data. The off-peak hour marker refers to the identifier of the period of low electricity demand defined in the electricity pricing mechanism, which is used to trigger the attenuation compensation of health data. The impact factor refers to the weight coefficient dynamically assigned based on the period marker. The grid load weight is increased to 1.5 times during peak hours, and the battery health weight is reduced to 0.8 times during off-peak hours. The health load matrix is ​​a two-dimensional data matrix that integrates the battery health attenuation characteristics and the grid load period characteristics. The rows represent time slices, and the columns contain health scores and load demand values.

[0099] In an embodiment of the present application, based on the peak period mark (such as 14:00-17:00) and the valley period mark (such as 00:00-06:00) in the grid electricity price data, the accelerated attenuation factor is applied during the peak period to calculate the battery health value, and the protective attenuation factor is applied during the valley period to calculate the battery health value; the grid load data is simultaneously assigned a period priority factor: the peak period load is assigned a high weight factor, and the valley period load is assigned a low weight factor; the processed battery health status data sequence and the grid load data sequence are matrix-axis aligned and spliced ​​to generate a healthy load matrix containing a time slice index column, a health attenuation feature column, and a load priority feature column.

[0100] Step 213: Utilize the fusion feature layer of the pre-trained reinforcement learning agent to jointly encode the health load matrix and the battery load coupling vector to generate a joint feature vector.

[0101] In this step, the fusion feature layer refers to the neural network module in the reinforcement learning agent that implements spatiotemporal feature fusion. It generates a unified feature representation through time series modeling and feature cross-pollination. The joint feature vector is the high-dimensional vector output by the fusion feature layer, which encodes the coupled relationship between battery status, grid load, and health decay.

[0102] In an embodiment of the present application, a fusion feature layer of a pre-trained reinforcement learning agent is used to perform forward and reverse bidirectional feature extraction on the time dimension health decay characteristic sequence in the health load matrix to generate a time context feature sequence. At the same time, a multi-level spatial relationship modeling of the battery load coupling vector is performed through a fully connected deep neural network to generate a high-dimensional spatial feature vector. The two types of features are then input into the cross-attention mechanism module for joint encoding. By calculating the attention distribution of the spatial feature sequence to the time feature sequence, the dynamic weighted fusion of the two types of features is realized, and a joint feature vector representing the multi-dimensional coupling relationship is output.

[0103] Step 214: Utilize the power derivation layer in the pre-trained reinforcement learning agent to perform power value regression processing on the joint feature vector to obtain an initial charging power value and an initial discharging power value.

[0104] In this step, the power inference layer refers to the output layer structure of the reinforcement learning agent, which maps the joint feature vector into a regression network that can execute power instructions.

[0105] In an embodiment of the present application, the joint feature vector is input into the power derivation layer of the reinforcement learning agent, and the power value regression processing of the joint feature vector is performed through a fully connected regression network. Specifically, three layers of nonlinear transformation calculation are performed, including the first layer for feature dimension compression, the second layer for activation to generate nonlinear response, and the third layer for independently outputting two regression values, which serve as the initial charging power value and the initial discharging power value respectively. The two together constitute the core instruction parameters of the unconstrained optimization strategy.

[0106] The embodiment of the present application solves the problem of coordinated optimization of grid load fluctuations and battery health degradation through spatiotemporal feature fusion and dynamic factor allocation, and achieves unified decision-making on economy and safety.

[0107] This application provides a specific embodiment, step 103, which allocates power to the power storage system, the power interface of the multi-port converter, and the industrial and commercial park load according to the charge and discharge scheduling instruction, and generates multi-source coordinated power flow data, specifically including the following steps:

[0108] Step 301: Determine the charge and discharge requirement parameters of the power energy storage system, the power transmission parameters of the power interface of the multi-port converter, and the real-time power requirement parameters of the industrial and commercial park load according to the charge and discharge scheduling instruction.

[0109] In this step, the charge / discharge demand parameter refers to the total energy (kWh) planned to be charged / discharged by the energy storage system during the scheduling period. It is calculated by multiplying the charge / discharge power allocation by the time slice length. The power transfer parameter refers to the real-time maximum allowable transfer power of each power port of the multi-port converter. The real-time power demand parameter refers to the immediate total power demand of all electrical equipment in the industrial and commercial park.

[0110] In an embodiment of the present application, the charge and discharge power allocation and discharge power allocation are extracted from the charge and discharge scheduling instructions, and the charge and discharge demand parameters are calculated in combination with the rated parameters of the energy storage system; the available power values ​​of each power interface of the multi-port converter are synchronously collected to generate power transmission parameters; and at the same time, the real-time power consumption data of each electrical equipment in the industrial and commercial park is monitored to generate real-time power demand parameters.

[0111] Step 302: According to a preset dynamic matching rule, the charging and discharging demand parameters, the power transmission parameters, and the real-time power demand parameters are collaboratively optimized to generate a power allocation plan.

[0112] In this step, dynamic matching rules are implemented using a real-time optimization algorithm with economic, safety, and environmental goals as the goals. The power allocation plan includes a time slice (e.g., 15 minutes), energy storage charge and discharge capacity, and an optimized schedule for power allocation to each interface.

[0113] In an embodiment of the present application, through preset dynamic matching rules, constraints are established so that the charging and discharging demand parameters do not exceed the total transmission capacity upper limit of the power transmission parameters; the actual supply rate of the real-time power demand parameters is ensured to be no lower than the preset threshold; the power transmission balance between the power interfaces of the multi-port converter is used as the optimization target, and a power distribution scheme is obtained that includes time segment identifiers, charging and discharging power distribution values, and interface power distribution values.

[0114] Step 303: performing channel allocation on the transmission volume of the power interface of the multi-port converter according to the power allocation scheme and the preset priority rule, and generating an interface power transmission data set.

[0115] In this step, priority rules establish a hard policy for the order in which interfaces are called, based on energy type. The transmission capacity of a multi-port converter's power interface refers to the actual power value transmitted by a single power interface within a specific time slice. The interface power transmission dataset is the aggregated transmission capacity of all interfaces.

[0116] In an embodiment of the present application, based on the total power demand value of each time slice in the power allocation scheme, through preset priority rules, the maximum available transmission capacity of the photovoltaic interface is preferentially allocated to the basic load demand, the remaining demand is supplemented by the transmission capacity of the grid interface, and the emergency gap calls the transmission capacity of the generator interface, and finally generates an interface power transmission data set containing the power transmission value of each power interface in the corresponding time slice.

[0117] Step 304: Integrate the optimized charge and discharge amount in the power allocation scheme, the actual transmission amount in the interface power transmission data set, and the actual supply amount of the industrial and commercial park load to generate multi-source collaborative power flow data.

[0118] In this step, the optimized charge and discharge capacity refers to the energy storage system charge and discharge energy values ​​planned for execution in the power allocation plan. The actual transmission capacity refers to the real-time transmission value recorded by the power metering device at each power interface. The actual supply capacity refers to the real-time consumption value recorded by the load-side energy meter.

[0119] In an embodiment of the present application, the optimized charging and discharging amount of the energy storage system in the power allocation scheme, the actual transmission amount in the interface power transmission data set, and the actual supply amount of the load metering device in the industrial and commercial park are collected, aligned according to the time axis, and integrated into a structured data table to generate multi-source collaborative power flow data that records the energy flow of the entire link.

[0120] The embodiments of the present application eliminate the energy waste caused by the lag in source-grid-load coordination in traditional methods through millisecond-level dynamic matching of power storage systems, multiple types of power interfaces and industrial and commercial loads, and maximize the utilization rate of clean energy while ensuring load demand.

[0121] This application provides a specific embodiment, step 105, performing optimized scheduling of the electric energy storage system according to the optimized charge and discharge scheduling instruction, specifically including the following steps:

[0122] Step 501: parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power supply interface of the multi-port converter.

[0123] In this step, the target charge capacity refers to the total planned charge energy set in the optimized charge and discharge scheduling instructions. The target discharge capacity refers to the total planned release energy set in the optimized charge and discharge scheduling instructions. The optimized allocation capacity refers to the power allocation ratio parameters of each power interface within the scheduling cycle.

[0124] In an embodiment of the present application, the optimized charge and discharge scheduling instructions are subjected to structured decomposition processing, and the key variables of the charging operation parameter part are extracted as the target charging amount, the key variables of the discharging operation parameter part are extracted as the target discharging amount, and the multi-interface allocation list of the power interface allocation part is extracted as the optimized allocation amount.

[0125] Step 502: According to a preset time segment division rule, the target charge amount and the target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount.

[0126] In this step, the time segment division rule refers to the basic time axis division method that divides 24 hours into 96 15-minute segments. The timestamped charge and discharge control signal refers to a timing control instruction in the form of [timestamp, charging power, discharging power]. The power adjustment instruction for each power interface refers to a control command that includes the interface code and power setting value.

[0127] In an embodiment of the present application, based on a preset time segment division rule, the target charging amount is apportioned according to the time segment length to generate a charging power sequence, and the target discharging amount is apportioned according to the time segment length to generate a discharging power sequence, and then the two sequences are superimposed with a timestamp code to form a charging and discharging control signal; synchronously, the power adjustment instructions for each power interface are generated according to the interface allocation ratio in the optimized allocation amount multiplied by the maximum available power of the interface.

[0128] Step 503: Send the charge and discharge control signal to the electric energy storage system, and simultaneously send the power adjustment instruction to the control unit of the multi-port converter to execute the scheduling of the electric energy storage system.

[0129] In an embodiment of the present application, the charge and discharge control signals are sent to the battery management system of the power storage system through the power communication network to trigger power conversion. At the same time, the power adjustment instructions of each power interface are sent to the control unit of the multi-port converter through the control bus to drive the power semiconductor switch to execute the scheduling of the power storage system.

[0130] Step 504: During the process of executing the scheduling of the electric energy storage system, the battery operation status data, the transmission data of each power interface, and the load consumption data are collected to generate an operation monitoring data set.

[0131] In this step, battery operating status data refers to operational safety parameters including cell voltage, temperature, and total battery state of charge. Transmission data for each power interface refers to the real-time recording of the actual transmission power value and direction of each interface. Load consumption data refers to the total active power and power factor of the industrial and commercial park. Operational monitoring data sets refer to time-stamped monitoring records for batteries, interfaces, and loads.

[0132] In an embodiment of the present application, during the execution of instructions, the battery operating status data of the power storage system, the transmission data recorded by the power metering devices of each power interface, and the load consumption data uploaded by the smart meter on the load side are collected in real time and integrated into a three-dimensional operation monitoring data set with a timestamp.

[0133] Step 505: Perform deviation detection on the operation monitoring data set and the expected value of the optimized charge and discharge scheduling instruction. When it is detected that the charge amount deviation exceeds a first preset threshold or the discharge amount deviation exceeds a second preset threshold, trigger the reinforcement learning agent to regenerate the optimized charge and discharge scheduling instruction until the detection passes, so as to perform optimized scheduling closed-loop control of the power energy storage system.

[0134] In this step, the theoretical charge and discharge energy standard value set in the expected value instruction of the optimized charge and discharge scheduling instruction is used. The charge capacity deviation refers to the absolute deviation rate between the actual total charge amount and the expected value. The first preset threshold value is the maximum allowable charge deviation rate. The discharge capacity deviation refers to the absolute deviation rate between the actual total discharge amount and the expected value. The second preset threshold value is the maximum allowable discharge deviation rate.

[0135] In an embodiment of the present application, the actual cumulative value of the charging amount is extracted from the operation monitoring data set and the expected value of the optimized charging and discharging scheduling instruction is used to calculate the charging amount deviation, and the actual cumulative value of the discharging amount is extracted to calculate the discharge amount deviation; when the charging amount deviation exceeds the first preset threshold or the discharge amount deviation exceeds the second preset threshold, the reinforcement learning agent is restarted to generate a new instruction; until the two deviations are lower than the corresponding preset thresholds, the closed-loop control is completed.

[0136] The embodiment of the present application solves the problem of strategy failure caused by environmental changes in traditional open-loop control through a real-time deviation feedback mechanism, thereby ensuring the reliability of the implementation of the economic scheduling strategy.

[0137] This application provides a specific embodiment, step 501, parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power interface of the multi-port converter, specifically comprising the following steps:

[0138] Step 511: Identify the instruction header identification field and instruction body data block in the optimized charge and discharge scheduling instruction.

[0139] In this step, the instruction header identification field refers to the fixed-length control field at the beginning of the instruction, which contains the version number and cyclic redundancy check code, used to verify the integrity of the instruction. The instruction body data block refers to the core data segment that stores the actual scheduling parameters, and the specific fields are divided according to the version template.

[0140] In the embodiment of the present application, the optimized charge-discharge scheduling instruction is subjected to binary stream parsing to identify the instruction header identification field at the start position of the instruction and the subsequent instruction body data block.

[0141] Step 512: According to the preset instruction version matching rule, the instruction body data block is segmented using the data structure definition template corresponding to the instruction header identification field to obtain the target charge value segment, the target discharge value segment and the time segment identification code segment.

[0142] In this step, the instruction version matching rule refers to the mapping rules between version numbers and field templates. The data structure definition template refers to a pre-set binary field parsing template that defines field length, type, and storage order. The target charge value segment instruction body specifically stores a binary segment of the charge value. The target discharge value segment instruction body specifically stores a binary segment of the discharge value. The time segment identifier code segment refers to the field that stores the standardized time slice number.

[0143] In an embodiment of the present application, based on the preset instruction version matching rules, the version code value in the instruction header identification field is parsed, and the corresponding data structure definition template is matched in the pre-stored data structure template library; according to the predefined field length sequence and field order rules in the data structure definition template, the instruction body data block is sequentially cut off in equal length: the first fixed-length data segment is cut off from the starting position of the instruction body data block as the target charge value segment, followed by the second fixed-length data segment being cut off as the target discharge value segment, and the third fixed-length data segment being cut off as the time segment identification code segment.

[0144] Step 513: extracting a first value from the target charge amount value segment as a target charge amount, and extracting a second value from the target discharge amount value segment as a target discharge amount.

[0145] In this step, the first value is the decimal floating-point value parsed from the target charge capacity value segment. The target charge capacity is the total charging demand (in kWh) directly assigned by the first value. The second value is the decimal floating-point value parsed from the target discharge capacity value segment.

[0146] In an embodiment of the present application, the binary data in the target charge amount value segment is converted into a decimal floating point number as a first value (such as 101.5), and the target charge amount is directly assigned; the binary data in the target discharge amount value segment is converted into a decimal floating point number as a second value (such as 0.0), and the target discharge amount is directly assigned.

[0147] Step 514: Determine a record entry corresponding to the time segment in the interface power transmission data set according to the time identification information in the time segment identification code segment.

[0148] In this step, the time identification information refers to the readable time slot number in the time slot identification code segment, and the record entry refers to the data row in the interface power transmission data set that matches the specific time identification information.

[0149] In an embodiment of the present application, the time slice identification code segment is converted from binary to text format to extract the structured time identification information; then, a hash matching query is performed in the hash index of the interface power transmission data set using the time slice number as the index key; when there is an index key that is completely consistent with the time identification information, the corresponding data record row is locked and returned as the target record entry.

[0150] Step 515: extract the power transmission amount allocation value of the power interface of the multi-port converter from the record entry as the optimized allocation amount.

[0151] In this step, the power transmission amount allocation value refers to the power value corresponding to a power interface in the record entry.

[0152] In an embodiment of the present application, the power transmission amount allocation value corresponding to each power interface is extracted from the matching record entry, such as the photovoltaic interface value PV:50.2, which is used as the optimized allocation amount of each interface.

[0153] The embodiment of the present application solves the problem of separation between instruction parameters and execution data in traditional methods through a time slice coding matching mechanism, ensuring the temporal and spatial consistency of the optimization strategy.

[0154] This application provides a specific embodiment, step 502, converting the target charge amount and target discharge amount into a charge and discharge control signal with a timestamp according to a preset time segment division rule, and generating a power adjustment instruction for each power interface according to the optimized allocation amount, specifically including the following steps:

[0155] Step 521: split the target charge amount and the target discharge amount according to the preset time segment division rule to obtain charging period sub-amounts and discharging period sub-amounts corresponding to multiple time segments.

[0156] In this step, the charging period sub-quantity refers to the energy value (in kWh) planned to be charged in a single time segment. It is calculated as follows: target charging amount ÷ total number of time segments × time segment duration coefficient. The discharging period sub-quantity refers to the energy value (in kWh) planned to be released in a single time segment. It is calculated as follows: target discharge amount ÷ total number of time segments × time segment duration coefficient.

[0157] In an embodiment of the present application, the scheduling period is divided into multiple fixed-length segments based on a preset time segment division rule, and the target charging amount is proportionally allocated according to the length of each time segment to obtain a charging period sub-amount of each segment. At the same time, the target discharging amount is proportionally allocated according to the length of each time segment to obtain a discharging period sub-amount of each segment.

[0158] Step 522: Dynamically bind the charging period sub-quantity and the discharging period sub-quantity corresponding to each time segment with the corresponding start timestamp and end timestamp to obtain a charging control unit and a discharging control unit.

[0159] In this step, the start timestamp refers to the standard time at which the time slice begins. The end timestamp refers to the standard time at which the time slice ends. The charge control unit refers to a triple data structure consisting of the time slice start timestamp, end timestamp, and charge period sub-quantity. The discharge control unit refers to a triple data structure consisting of the time slice start timestamp, end timestamp, and discharge period sub-quantity.

[0160] In an embodiment of the present application, the charging period sub-quantity and the discharging period sub-quantity generated for each time segment are respectively data-bound with the preset start timestamp and end timestamp of the time segment to generate a charging control unit including a charging value and a time interval, and a discharging control unit including a discharging value and a time interval.

[0161] Step 523: Arrange the charging control unit and the discharging control unit in chronological order to generate a charging and discharging control signal with a time stamp.

[0162] In this step, the charge and discharge control signal with a time stamp refers to a sequence of charge control units and discharge control units arranged in chronological order.

[0163] In an embodiment of the present invention, the charging control unit and the discharging control unit corresponding to each time slice are arranged in ascending order according to the chronological order of the start timestamps of the time slices, forming an ordered control unit sequence consisting of a start timestamp, an end timestamp, a charging period sub-quantity, and a discharging period sub-quantity. This sequence is a time-stamped charging and discharging control signal containing a complete time dimension mark and energy control parameters.

[0164] Step 524: Bind the power transmission values, identifiers, and time segment information of the power interfaces of the multi-port converter in corresponding time segments according to the optimized allocation amount, and generate power adjustment instructions for each power interface.

[0165] In this step, the power transfer value refers to the power value (in kW) that the power interface needs to transfer during a specific time slice. It is calculated by multiplying the optimized allocation ratio by the total power demand. The identifier is the unique identifier of the power interface. The time slice information is a time interval encapsulated object with a start and end timestamp.

[0166] In an embodiment of the present application, based on the power transmission ratio value of each power interface of the multi-port converter in the optimized allocation amount and combined with the time segment information, the power transmission value of each interface in the corresponding time slice is calculated, and then the value, interface identifier and time segment information are bound to generate a power adjustment instruction for the independent power interface.

[0167] The embodiment of the present application solves the problem of asynchrony between charge and discharge instructions and interface control signals in traditional methods through a time slice binding mechanism, thereby ensuring millisecond-level coordinated response of the multi-source system.

[0168] Figure 2 The present invention provides a structural diagram of an electric energy storage system optimization and scheduling system based on reinforcement learning, as shown in FIG. Figure 2 As shown, the system includes:

[0169] An acquisition module 21 is configured to generate a state input data set based on acquiring battery state of charge data and battery health status data of the power energy storage system and grid load data and grid electricity price data of the industrial and commercial park;

[0170] A processing module 22 is configured to process the state input data set using a pre-trained reinforcement learning agent to obtain a charge and discharge scheduling instruction;

[0171] a distribution module 23 for distributing power to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park according to the charge and discharge scheduling instructions, and generating multi-source collaborative power flow data;

[0172] An optimization module 24 is configured to combine the multi-source coordinated power flow data and historical operation data, optimize the charge and discharge scheduling instructions using the charge and discharge scheduling rules preset in the reinforcement learning agent, and generate optimized charge and discharge scheduling instructions;

[0173] The scheduling module 25 is used to optimize the scheduling of the electric energy storage system according to the optimized charge and discharge scheduling instructions.

[0174] Figure 2 The power storage system optimization and dispatching system based on reinforcement learning can be executed Figure 1 The implementation principles and technical effects of the reinforcement learning-based power storage system optimization and scheduling method described in the illustrated embodiment are not further elaborated. The specific manner in which each module and unit performs operations in the reinforcement learning-based power storage system optimization and scheduling system in the above embodiment has been described in detail in the embodiments of the method and will not be elaborated on here.

[0175] In one possible design, Figure 2 The embodiment shown is a power storage system optimization and scheduling system based on reinforcement learning, which can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0176] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0177] The processing component 32 is used for the above Figure 1 The embodiment provides a method for optimizing and scheduling an electric energy storage system based on reinforcement learning.

[0178] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0179] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0180] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0181] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.

[0182] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0183] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0184] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment provides a method for optimizing and scheduling an electric energy storage system based on reinforcement learning.

[0185] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0187] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for optimizing and scheduling an electric energy storage system based on reinforcement learning, characterized in that: include: Generate a state input data set based on the battery state of charge data and battery health status data of the power energy storage system, as well as the grid load data and grid electricity price data of the industrial and commercial park; Using a pre-trained reinforcement learning agent, the state input data set is processed to obtain a charge and discharge scheduling instruction; According to the charge and discharge scheduling instructions, power is distributed to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park to generate multi-source collaborative power flow data; Combining the multi-source collaborative power flow data and historical operation data, optimizing the charge and discharge scheduling instructions using the charge and discharge scheduling rules preset in the reinforcement learning agent to generate optimized charge and discharge scheduling instructions; Optimizing the scheduling of the electric energy storage system according to the optimized charge and discharge scheduling instructions; Optimizing the scheduling of the power energy storage system according to the optimized charge and discharge scheduling instructions includes: parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power interface of the multi-port converter; According to a preset time segment division rule, the target charge amount and target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount; Sending the charge and discharge control signal to the electric energy storage system, and simultaneously sending the power adjustment instruction to the control unit of the multi-port converter to execute the scheduling of the electric energy storage system; During the dispatching of the power storage system, the battery operating status data, the transmission data of each power interface, and the load consumption data are collected to generate an operation monitoring data set; Deviation detection is performed on the operation monitoring data set and the expected value of the optimized charge and discharge scheduling instruction. When it is detected that the charge amount deviation exceeds a first preset threshold or the discharge amount deviation exceeds a second preset threshold, the reinforcement learning agent is triggered to regenerate the optimized charge and discharge scheduling instruction until the detection passes, so as to perform optimized scheduling closed-loop control of the power energy storage system.

2. The method according to claim 1, characterized in that The state input data set is processed using a pre-trained reinforcement learning agent to obtain charge and discharge scheduling instructions, including: Using the pre-trained reinforcement learning agent, feature extraction is performed on the state input data set to obtain an initial charging power value and an initial discharging power value; Adjusting the initial charging power value and the initial discharging power value according to a preset maximum charging power limit value of the battery, a maximum discharging power limit value of the battery, and a safe range of the battery charging state to generate a charging power value and a discharging power value; The charging power value and the discharging power value are converted into a charging and discharging scheduling instruction, wherein the charging and discharging scheduling instruction includes a charging power allocation amount and a discharging power allocation amount.

3. The method according to claim 2, characterized in that Using the pre-trained reinforcement learning agent, feature extraction is performed on the state input data set to obtain an initial charging power value and an initial discharging power value, including: coupling the battery charging state data and the grid load data in the state input data set to generate a battery load coupling vector; Allocating influencing factors to the battery health status data and the grid load data according to the peak period mark and the off-peak period mark in the grid electricity price data to obtain a health load matrix; Using the fusion feature layer of the pre-trained reinforcement learning agent, jointly encode the health load matrix and the battery load coupling vector to generate a joint feature vector; The power derivation layer in the pre-trained reinforcement learning agent is used to perform power value regression processing on the joint feature vector to obtain an initial charging power value and an initial discharging power value.

4. The method according to claim 1, wherein According to the charge and discharge scheduling instructions, power is distributed to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park to generate multi-source collaborative power flow data, including: Determining, according to the charge and discharge scheduling instructions, the charge and discharge demand parameters of the electric energy storage system, the power transmission parameters of the power interface of the multi-port converter, and the real-time power demand parameters of the industrial and commercial park load; According to a preset dynamic matching rule, the charging and discharging demand parameters, the power transmission parameters, and the real-time power demand parameters are collaboratively optimized to generate a power allocation plan; According to the power allocation scheme and the preset priority rules, channel allocation is performed on the transmission amount of the power interface of the multi-port converter to generate an interface power transmission data set; The optimized charge and discharge amount in the power allocation scheme, the actual transmission amount in the interface power transmission data set, and the actual supply amount of the industrial and commercial park load are integrated to generate multi-source collaborative power flow data.

5. The method according to claim 1, characterized in that Parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power interface of the multi-port converter includes: Identifying the instruction header identification field and instruction body data block in the optimized charge and discharge scheduling instruction; According to a preset instruction version matching rule, the instruction body data block is segmented using a data structure definition template corresponding to the instruction header identification field to obtain a target charge value segment, a target discharge value segment, and a time segment identification code segment; extracting a first value from the target charge amount value segment as a target charge amount, and extracting a second value from the target discharge amount value segment as a target discharge amount; Determining, according to the time identification information in the time segment identification code segment, a record entry corresponding to the time segment in the interface power transmission data set; The power transmission amount distribution value of the power interface of the multi-port converter is extracted from the record entry as the optimized distribution amount.

6. The method according to claim 1, characterized in that According to a preset time segment division rule, the target charge amount and the target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount, including: Splitting the target charge amount and the target discharge amount according to the preset time segment division rule to obtain charging period sub-amounts and discharging period sub-amounts corresponding to the multiple time segments; Dynamically bind the charging period sub-quantity and the discharging period sub-quantity corresponding to each time segment with the corresponding start timestamp and end timestamp to obtain a charging control unit and a discharging control unit; Arranging the charging control unit and the discharging control unit in chronological order to generate a charging and discharging control signal with a time stamp; According to the optimized allocation amount, the power transmission value, identifier and time segment information of the power interface of the multi-port converter in the corresponding time segment are bound to generate a power adjustment instruction for each power interface.

7. A power energy storage system optimization and scheduling system based on reinforcement learning, characterized in that: include: An acquisition module is configured to generate a state input data set based on the acquired battery state of charge data and battery health status data of the power energy storage system and the grid load data and grid electricity price data of the industrial and commercial park; a processing module, configured to process the state input data set using a pre-trained reinforcement learning agent to obtain a charge and discharge scheduling instruction; a distribution module for distributing power to the power storage system, the power interface of the multi-port converter, and the load of the industrial and commercial park according to the charge and discharge scheduling instructions, and generating multi-source collaborative power flow data; an optimization module, configured to optimize the charge and discharge scheduling instructions by combining the multi-source collaborative power flow data and historical operation data and utilizing the charge and discharge scheduling rules preset in the reinforcement learning agent to generate optimized charge and discharge scheduling instructions; A scheduling module, configured to optimize the scheduling of the electric energy storage system according to the optimized charge and discharge scheduling instructions; Optimizing the scheduling of the power energy storage system according to the optimized charge and discharge scheduling instructions includes: parsing the optimized charge and discharge scheduling instruction to obtain a target charge amount, a target discharge amount, and an optimized allocation amount of the power interface of the multi-port converter; According to a preset time segment division rule, the target charge amount and target discharge amount are converted into a charge and discharge control signal with a timestamp, and at the same time, a power adjustment instruction for each power interface is generated according to the optimized allocation amount; Sending the charge and discharge control signal to the electric energy storage system, and simultaneously sending the power adjustment instruction to the control unit of the multi-port converter to execute the scheduling of the electric energy storage system; During the dispatching of the power storage system, the battery operating status data, the transmission data of each power interface, and the load consumption data are collected to generate an operation monitoring data set; Deviation detection is performed on the operation monitoring data set and the expected value of the optimized charge and discharge scheduling instruction. When it is detected that the charge amount deviation exceeds a first preset threshold or the discharge amount deviation exceeds a second preset threshold, the reinforcement learning agent is triggered to regenerate the optimized charge and discharge scheduling instruction until the detection passes, so as to perform optimized scheduling closed-loop control of the power energy storage system.

8. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a power energy storage system optimization scheduling method based on reinforcement learning as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the method for optimizing and scheduling an electric energy storage system based on reinforcement learning as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Day-ahead and intra-day economic dispatching method for wind and light storage system

    CN120150103A