Energy storage charging and discharging strategy generation method and system based on reinforcement learning

By applying reinforcement learning SAC algorithm in the energy storage system, optimizing the charging and discharging strategy of the energy storage system, the problem that the existing technology is difficult to adapt to complex scenarios is solved, and the stability and economic benefits of the power system are maximized.

CN120033728APending Publication Date: 2025-05-23STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510158641.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing energy storage charging and discharging strategies are difficult to adapt to complex scenarios where there are instability on both the power generation and power consumption sides, making it difficult to maximize the stability and economic benefits of the power system.

Method used

Using reinforcement learning-based methods, especially the Soft Actor-Critic (SAC) algorithm, the state space, action space and reward function of the energy storage system are constructed to optimize the charging and discharging strategy of the energy storage system by analyzing historical load data and multi-source data.

Benefits of technology

It improves the stability and economic benefits of the power system, can adapt to different types of energy storage systems and power system scenarios, realizes the optimal charging and discharging strategy of the energy storage system, smoothes the fluctuations on the power generation side and meets the dynamic needs of the power consumption side.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120033728A_ABST
    Figure CN120033728A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent power grids, and particularly relates to an energy storage charging and discharging strategy generation method and system based on an SAC reinforcement learning method, and the method comprises the steps: data collection and preprocessing: collecting the real-time output data of a power generation side and the load demand data of a power utilization side, and carrying out the preprocessing of the data, comprising data cleaning, missing value filling and normalization processing; building a reinforcement learning model: selecting an SAC algorithm as a core algorithm of reinforcement learning, and carrying out dynamic strategy adjustment to adapt to a real-time operation environment, so that the method aims at optimizing charging and discharging behaviors of the energy storage system, stabilizing power generation side fluctuation, meeting dynamic requirements of a power utilization side and improving stability of a power system through an intelligent algorithm; and maximization of economic benefits and environmental benefits is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart grids, and in particular relates to a method and system for generating energy storage charging and discharging strategies based on a SAC reinforcement learning method. Background Art

[0002] With the widespread use of renewable energy, the power output on the power generation side (especially green electricity, such as solar and wind power) has significant volatility. At the same time, the load demand on the power consumption side also shows the characteristics of dynamic change. This two-sided uncertainty poses a challenge to the stable operation of the power system. As a technical means that can instantly smooth out the impact of the randomness of renewable energy generation on the supply side and respond promptly to the dynamic changes in the demand side load, the design of the charging and discharging strategy of the energy storage system is particularly important.

[0003] Most existing energy storage charging and discharging strategies are based on simple rules or fixed thresholds, which make it difficult for them to adapt to complex scenarios where instability exists on both the power generation and consumption sides.

[0004] Therefore, the present invention proposes a method for generating energy storage charging and discharging strategies based on reinforcement learning, aiming to optimize the charging and discharging behavior of the energy storage system through an intelligent algorithm to cope with the challenges of bilateral uncertainty. Summary of the invention

[0005] Purpose of the invention: With the widespread application of renewable energy and the dynamic changes in user electricity demand, the power system faces the challenge of uncertainty on both the power generation side and the power consumption side. This uncertainty leads to an imbalance in power supply and demand, affecting the stable operation of the power grid. Therefore, the purpose of the present invention is to provide a method for generating energy storage charging and discharging strategies based on reinforcement learning. This method can comprehensively consider the uncertainties on the power generation side and the power consumption side, and use the SAC (Soft Actor-Critic) algorithm to design the optimal energy storage charging and discharging strategy to achieve stable operation of the power system and maximize economic benefits. The present application also provides a system for generating energy storage charging and discharging strategies based on reinforcement learning.

[0006] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is:

[0007] In a first aspect, the present invention provides a method for generating an energy storage charging and discharging strategy based on reinforcement learning, the method comprising:

[0008] Collect historical load data of a period of time in the distribution network, and classify the historical load data to obtain a feature data set containing multiple key information;

[0009] Constructing a state space and a reward function of a battery energy storage system according to the characteristic data set of the key information, wherein the battery energy storage system is used to store and release electric energy in response to changes in the power generation side and the power consumption side of the power system; constructing an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system;

[0010] The flexible action-evaluation algorithm is used to analyze the correlation between the power generation and load power consumption under different categories;

[0011] Continue to fuse multi-source data, and use the time series feature extraction method to update the state space, action space and reward function after feature extraction, perform strategy evaluation after correlation analysis again, and adjust the strategy according to the strategy evaluation results.

[0012] Further, including:

[0013] The historical load data includes the historical load power data of the substations within the distribution network of the study area, the wind power and irradiance data within the substations, and the distributed photovoltaic power generation power within the substations collected from multiple systems over a period of time, and the historical load data are classified, specifically:

[0014] The historical load data is classified according to the special transformer substation and the public transformer substation; the daily load data of all users under the special transformer substation is aggregated and feature-recognized to form substation data classification processing of different types of loads, and different types of load data are obtained; the public transformer substation is classified into rural, township and urban areas to obtain load data of different regions; the historical load data is divided into summer and winter to obtain load data of different seasons; the sample data is classified according to working days and non-working days to obtain load data of different times; the ratio of the photovoltaic, wind power and hydropower installed capacity of the substation to the maximum load of the substation is calculated to obtain the green electricity penetration rate of the substation, and the historical load data is divided into three intervals of [0%-30%], (30%-80%) and (80%-100%) according to the penetration rate to obtain load data of different green penetration rates; the data change rate is calculated, and the change rate is divided into three intervals of [0%-30%], (30%-80%) and (80%-100%), and the load data corresponding to the change rate of the three intervals are extracted respectively.

[0015] Further, including:

[0016] The step of constructing a state space of the battery energy storage system according to the characteristic data set of the key information specifically includes:

[0017] A state space set of a battery energy storage system is defined, wherein the state space set includes real-time data, classified load data and meteorological data, wherein the real-time data includes: current time, current output data of the power generation side of the distribution network, current user power demand of the power consumption side of the distribution network and current state of charge of the battery energy storage system; the classified load data is a type of classified load data;

[0018] The action space is obtained by the difference between the current charging power of the battery energy storage system and the current discharging power of the battery energy storage system;

[0019] The reward function of the battery energy storage system is constructed based on the characteristic data set of the key information, including the following factors: the power factor of purchasing or selling electricity from the power grid at the current moment, the energy change factor of the battery energy storage system corresponding to the next moment and the current moment, the ratio factor of the predicted power generation of renewable energy and the user's electricity demand at the current moment, and the charge factor of the current battery energy storage system.

[0020] Further, including:

[0021] The flexible action-evaluation algorithm is used to analyze the correlation between the power generation and the load power under different categories, specifically including training the flexible action-evaluation algorithm:

[0022] Initialize the Actor network and two Critic networks; the Actor network is responsible for outputting the probability distribution of each action in the current state, while the Critic network is responsible for estimating the value of the state-action pair; for each time step, obtain the current state space from the environment;

[0023] Use the Actor network to output the probability distribution of each action space under the current state space; sample an action space from the probability distribution and execute it; obtain a new state space and reward function from the environment; use two Critic networks to estimate the action value and entropy of the state-action pair respectively; update the Actor network and Critic network according to the action value and entropy; repeat the above until the algorithm converges or reaches the maximum number of iterations, so as to predict the best action probability and state value estimation under the current classification of the load data;

[0024] Repeat the above steps to predict the best action probability and state value estimation under all load data categories, thereby obtaining the strategy under each load data category.

[0025] Further, including:

[0026] The method of continuing to fuse multi-source data and extracting features using a time series feature extraction method to update the state space, action space and reward function includes:

[0027] The classified historical load data is updated according to the time series to ensure that the historical load data is always the data of the most recent period, and real-time data is collected to form a multi-source data set under a classification;

[0028] Extracting features from the multi-source data set using a time series feature extraction method to form a key feature set including multiple features;

[0029] Then, based on the multi-source data sets under all categories, the key feature sets under all categories are obtained.

[0030] Further, including:

[0031] After the correlation analysis is performed again, a strategy evaluation is performed, and the strategy is adjusted according to the strategy evaluation result, including:

[0032] According to the key feature set under each category, the state space, the action space and the corresponding reward function are updated, and the flexible action-evaluation algorithm is trained using the above data to obtain the optimal decision;

[0033] The optimal decision is evaluated using strategy evaluation indicators. If the set strategy effect is achieved, the correlation between the power generation and load power consumption under different time series is obtained according to the current decision. Otherwise, the current decision is adjusted and the strategy is re-evaluated.

[0034] Further, including:

[0035] The reward function corresponding to the update is expressed as: the current reward plus the adjustment item under the current category, and the adjustment item under the current category is obtained according to the actual changes of the control parameters under the current category at different times.

[0036] Further, including:

[0037] The strategy adjustment is expressed as a strategy adjustment amount calculated based on the current strategy plus a strategy evaluation result, and the strategy adjustment amount represents an adjustment amount of a control parameter of the battery energy storage system at the current time step.

[0038] In a second aspect, the present invention further provides a system for generating energy storage charging and discharging strategies based on reinforcement learning, the system comprising:

[0039] A historical load data collection module is used to collect historical load data of a period of time in the distribution network, and classify the historical load data to obtain a feature data set containing multiple key information;

[0040] A parameter construction module, for constructing a state space and a reward function of a battery energy storage system according to a characteristic data set of the key information, wherein the battery energy storage system is used to store and release electric energy in response to changes in the power generation side and the power consumption side in the power system; and constructing an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system;

[0041] Algorithm training module, used to analyze the correlation between power generation and load power consumption under different categories using flexible action-evaluation algorithm;

[0042] The model updating module is used to continue fusing multi-source data, and to update the state space, action space and reward function after feature extraction using a time series feature extraction method, to perform strategy evaluation after performing correlation analysis again, and to adjust the strategy according to the strategy evaluation results.

[0043] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for generating energy storage charging and discharging strategies based on reinforcement learning.

[0044] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0045] 1. Improve the stability of the power system: By optimizing the charging and discharging behavior of the energy storage system, the present invention can smooth out the volatility of the power generation side and meet the dynamic needs of the power consumption side, thereby improving the stability of the power system.

[0046] 2. Improve economic benefits: By rationally scheduling the charging and discharging behavior of the energy storage system, the present invention can maximize economic benefits while ensuring the stable operation of the power system.

[0047] 3. Strong adaptability: The method of the present invention can be applied to different types of energy storage systems and power system scenarios, and has strong versatility and adaptability. Specifically:

[0048] In this application, the SAC algorithm indirectly guides the charging and discharging strategy of the energy storage system by analyzing the correlation between power generation and load power under different classifications (such as load type, region type, season type, green electricity penetration rate, data change rate). This correlation analysis reveals the degree of matching between renewable energy generation and load demand, thereby providing the following guidance for the energy storage system:

[0049] Optimized scheduling: When renewable energy generation is positively correlated with load demand, the energy storage system can be charged during the peak period of renewable energy generation and discharged during the peak period of load, thereby smoothing fluctuations and improving the utilization rate of renewable energy.

[0050] Prediction and prevention: By analyzing the correlation patterns in historical data, we can predict the trend of load demand changes in the future, so as to formulate charging and discharging strategies in advance and prevent imbalance in electricity supply and demand.

[0051] Dynamic adjustment: The SAC algorithm can dynamically adjust the charging and discharging behavior of the energy storage system based on real-time data and correlation analysis results to adapt to real-time changes in the power generation and consumption sides and ensure stable operation of the power system.

[0052] Therefore, by classifying and calculating the data, the present application can achieve the following technical effects:

[0053] More targeted: Data of different categories have different characteristics and rules. The classification calculation strategy can formulate more accurate charging and discharging strategies for different scenarios and improve the effectiveness of the strategy.

[0054] Higher accuracy: The classification calculation strategy can reduce the interference between data and improve the accuracy of correlation analysis, thereby obtaining more reliable strategy guidance.

[0055] Greater adaptability: The classification computing strategy can be flexibly adjusted for different scenarios to improve the adaptability and robustness of the strategy.

[0056] Higher economic benefits: Formulating strategies for different categories can make more effective use of renewable energy, reduce electricity purchase costs and improve economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of a strategy generation process after feature extraction using a time series feature extraction method provided by an embodiment of the present invention;

[0058] Figure 2 The historical load data classification method provided by the embodiment of the present invention is as follows: Figure 1 ;

[0059] Figure 3 The historical load data classification method provided by the embodiment of the present invention is as follows: Figure 2 ;

[0060] Figure 4 It is a simplified flow chart of the energy storage charging and discharging strategy generation method provided by an embodiment of the present invention;

[0061] Figure 5 It is a flow chart of the energy storage charging and discharging strategy generation method described in Example 1 of the present invention. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] Example 1

[0064] like Figure 5 As shown, the present invention provides a method for generating energy storage charging and discharging strategies based on reinforcement learning, the method comprising the following steps:

[0065] S1 collects historical load data of a period of time in the distribution network, and classifies the historical load data to obtain a feature data set containing multiple key information.

[0066] In this embodiment, if Figure 2 As shown, the data collection refers to the collection of historical load data of the area load power, wind power and irradiance data within the area, and distributed photovoltaic power generation power within the area for at least one year from various data sources such as the distribution automation system and the power consumption collection system for all the areas within the distribution network of the study area. The data is expressed as power and irradiance values ​​and their time scales at 288 or 96 points per day. The above-mentioned collected historical load data will be used to construct the state space of the reinforcement learning model, and combined with real-time load data for prediction to improve the adaptability and accuracy of the strategy.

[0067] like Figure 3 As shown, the feature data set, such as user type features, geographic features, time features, green electricity features and dynamic features, is mainly processed for classification and extraction of sample data, mainly including:

[0068] 1) Classify the sample data according to the special transformer area and public transformer area;

[0069] 2) Aggregate and identify the daily load data of all users under the dedicated transformer area to form a classification process for the area data of different types of loads;

[0070] 3) Classify public transformer stations into rural, township and urban areas;

[0071] 4) Divide the sample data into summer and winter;

[0072] 5) Classify the sample data into working days and non-working days;

[0073] 6) Calculate the ratio of the installed capacity of photovoltaic, wind power and hydropower generation in the area to the maximum load of the area to obtain the green electricity penetration rate of the area, and divide the sample data into three intervals according to the penetration rate: [0%-30%], (30%-80%), and (80%-100%);

[0074] 7) Calculate the data change rate, divide the change rate into three intervals: [0%-30%], (30%-80%), and (80%-100%), and extract the corresponding power generation power and load power consumption data in the three intervals respectively.

[0075] Through the above-mentioned characteristic data processing steps, a characteristic data set containing a variety of key information is obtained, such as load type, regional type, seasonal type, green electricity penetration rate, data change rate, etc. Next, these characteristic data will be used to analyze the correlation between the green electricity generation power and the load power consumption in the substation area, and a corresponding analysis model will be constructed to evaluate the charging and discharging strategy of the energy storage system. Therefore, the advantage of the classification calculation correlation used in the present invention is that the analysis is refined and the decision is made accurately: data of different classifications reflect different electricity consumption characteristics and laws. For example, there are obvious differences in the electricity demand of industrial loads and residential loads, and the load demand in rural and urban areas is also different. Through classification, more accurate charging and discharging strategies can be formulated for different scenarios. For example, for industrial loads, charging can be performed during the low-peak period at night and discharging can be performed during the peak period during the day; for residential loads, more refined scheduling can be performed according to electricity consumption habits. Improve the accuracy of correlation analysis: There may be interference between data of different classifications. For example, mixing industrial loads and residential loads for analysis may mask their respective electricity consumption characteristics. Through classification, the interference between data can be reduced, the accuracy of correlation analysis can be improved, and more reliable strategy guidance can be obtained. Enhance the adaptability of the strategy: The power system is a complex system that is affected by many factors, such as weather, seasons, and economic development. Through classification, flexible adjustments can be made for different scenarios to improve the adaptability and robustness of the strategy. For example, the charging and discharging strategy of the energy storage system can be adjusted according to seasonal changes. Improve economic benefits: By classifying the calculation strategy, renewable energy can be used more effectively, the cost of purchasing electricity can be reduced, and economic benefits can be improved. For example, different energy storage system charging and discharging strategies can be formulated according to the green electricity penetration rate in different regions to improve the utilization rate of renewable energy. Easy to manage and maintain: Classifying and storing data and managing it can facilitate data query, analysis, and maintenance, and improve work efficiency.

[0076] S2 constructs a state space and a reward function of a battery energy storage system according to the characteristic data set of the key information, wherein the battery energy storage system is used to store and release electrical energy in response to changes in the power generation side and the power consumption side of the power system; and constructs an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system.

[0077] S3 uses a flexible action-evaluation algorithm to analyze the correlation between power generation and load power consumption under different categories.

[0078] Specifically, this embodiment first describes the flexible action-criteria algorithm, namely the SAC algorithm, which is a reinforcement learning algorithm that allows the agent to obtain efficient learning and execution in the continuous action space by combining the discrete and continuous action space methods. SAC is a further extension of the deterministic policy gradient algorithm and the deep Q network algorithm, and is intended to solve control problems in high-dimensional, continuous and multi-modal action spaces.

[0079] It includes an actor network, 2 Q networks or 4 Critic networks, among which the 4 Critic networks include 2 V Critic networks and 2 Q Critic networks, one actor network, four critic networks, respectively, state value estimation υ and Targetv network; action-state value estimation Q 0 and Q 1 network.

[0080] The input of the actor network is the state, and the output is the action probability π(a t |s t ) or action probability distribution parameters, the input of the critic network is the state, and the output is the value of the state. The output of the V Critic network is u(s), which represents the estimate of the state value; the output of the Q Critic network is q(s,a), which represents the estimate of the action-state pair value. Because the concept of entropy is added to the SAC algorithm to encourage exploration, the training objectives of its actor and critic networks are different from those of conventional algorithms that do not contain entropy. In the SAC algorithm, if the action output by the actor network can make a comprehensive indicator larger, then the better.

[0081] In this embodiment, Figure 4As shown, the establishment of the analysis model mainly refers to analyzing the correlation coefficient between the green power generation power and the load power consumption of the station area by using the reinforcement learning SAC algorithm according to the data after the special data processing steps of the sample data. For example, the data classified according to the characteristics of load type, regional type, seasonal type, green power penetration rate, data change rate, etc. are used to analyze the correlation between the power generation power and the load power consumption under different categories. For example, the correlation coefficients of the power generation power and the load power consumption under different load types (industrial, commercial, and residential) can be calculated respectively to analyze the impact of different types of loads on the consumption of renewable energy. In addition, the correlation between the power generation power and the load power consumption under different regional types (rural, township, and urban), different seasonal types (summer, winter), different green power penetration rates, and different data change rates can also be analyzed to comprehensively evaluate the consumption of renewable energy and provide a basis for the charging and discharging strategy of the energy storage system. If the absolute value of the correlation coefficient is larger, the correlation is stronger, and the closer the correlation coefficient is to 0, the weaker the correlation is.

[0082] The specific implementation steps include:

[0083] The step of constructing a state space of the battery energy storage system according to the characteristic data set of the key information specifically includes:

[0084] A state space set of a battery energy storage system is defined, wherein the state space set includes real-time data, classified load data and meteorological data. The real-time data includes: current time, current output data on the power generation side of the distribution network, current user power demand on the power consumption side of the distribution network and current state of charge of the battery energy storage system; the classified load data is a type of classified load data.

[0085] As an example of this embodiment, the state space is represented as follows:

[0086] s t =(t,P gen,t ,P load,t ,E t ,H load ,W);

[0087] Where t is the current time; P gen,t is the current renewable energy generation on the power generation side of the distribution network; P load,t E is the current power demand of users on the user side of the distribution network; t H is the current state of charge (SOC) of the battery energy storage system; loadis the historical load data; W is the meteorological data, i.e. the weather forecast data, including but not limited to: temperature: maximum temperature, minimum temperature, average temperature, etc.; precipitation: rainfall, snowfall, etc.; wind speed: wind direction, wind speed level, etc.; humidity: relative humidity; cloud cover: cloud thickness, cloud cover percentage, etc.; sunshine: sunshine time, solar radiation intensity, etc.; among them, the user's electricity demand comes from the electricity collection system or smart meter, which records the user's real-time electricity load data. The renewable energy power generation comes from the distribution automation system or distributed energy management system, which records the real-time power generation power of renewable energy power generation equipment. The current state of charge of the battery energy storage system comes from the energy storage system's own monitoring system, which records the battery's charging level. The historical load data in this embodiment is a type of classified data. Therefore, this embodiment performs correlation analysis on each type of data, such as calculating the correlation coefficients between the power generation and load power consumption under different load types (industrial, commercial, and residential) to analyze the impact of different types of loads on the consumption of renewable energy. This requires forming a state space action space and related reward functions corresponding to all classified data. The SAC algorithm is used for cyclic calculations on each type of data to obtain the correlation between the power generation and load power consumption under each classification.

[0088] In this embodiment, the action space is obtained by the difference between the current charging power of the battery energy storage system and the current discharging power of the battery energy storage system.

[0089] Specifically, the implementation method corresponding to the action space of an energy storage system in this embodiment is expressed as follows:

[0090] a t =P charge,t -P discharge,t ;

[0091] Among them, P charge is the current charging power of the battery energy storage system; P discharge is the current discharge power of the battery energy storage system.

[0092] The reward function of the battery energy storage system is constructed based on the characteristic data set of the key information, including the following factors: the power factor of purchasing or selling electricity from the power grid at the current moment, the energy change factor of the battery energy storage system corresponding to the next moment and the current moment, the ratio factor of the predicted power generation of renewable energy and the user's electricity demand at the current moment, and the charge factor of the current battery energy storage system.

[0093] Specifically, the implementation method of a reward function of a battery energy storage system in this embodiment is expressed as follows:

[0094] r t =-c 1 *|P grid,t|-c 2 *ΔE t +c 3 *(P renew,t / P load,t )+c 4 *(1-E t / E max );

[0095] Among them, P grid,t =P load,t -P gen,t -a t is the power purchased or sold from the power grid; ΔE t =E t+1 -E t is the energy change of the energy storage device; c 1 is the cost coefficient of purchasing or selling electricity from the grid; c 2 is the energy change cost coefficient of the energy storage device; c 3 is the incentive coefficient for improving the utilization rate of renewable energy; c 4 It is to reduce the state of charge reward coefficient of the energy storage device; P renew,t is the current renewable energy power generation; P load,t E is the current user electricity demand; t E is the current state of charge of the energy storage device; max is the maximum state of charge of the battery energy storage system.

[0096] The flexible action-evaluation algorithm is used to analyze the correlation between the power generation and the load power under different categories, specifically including training the flexible action-evaluation algorithm:

[0097] Initialize the Actor network and two Critic networks; the Actor network is responsible for outputting the probability distribution of each action in the current state, while the Critic network is responsible for estimating the value of the state-action pair; for each time step, obtain the current state space from the environment;

[0098] Use the Actor network to output the probability distribution of each action space under the current state space; sample an action space from the probability distribution and execute it; obtain a new state space and reward function from the environment; use two Critic networks to estimate the action value and entropy of the state-action pair respectively; update the Actor network and Critic network according to the action value and entropy; repeat the above until the algorithm converges or reaches the maximum number of iterations, so as to predict the best action probability and state value estimation under the current classification of the load data;

[0099] Repeat the above steps to predict the best action probability and state value estimation under all load data categories, thereby obtaining the strategy under each load data category.

[0100] Specifically, a SAC algorithm training process of this embodiment includes:

[0101] Step 1. Initialize the network and parameters:

[0102] Actor network: used to select actions.

[0103] Critic 1 and Critic 2 networks: used to estimate Q-values.

[0104] Target Critic 1 and Target Critic 2: Same as the Critic network architecture, used to generate more stable target Q values.

[0105] Step 2: Calculate the target Q value:

[0106] Use the target network to calculate the Q value at the next state.

[0107] Take the minimum value of the two Q network outputs to prevent overestimation of the Q value,

[0108] Introduce the entropy regularization term, and the calculation formula is:

[0109] y=r+γ·min(Q 1 ,Q 2 )-α·logπ(a|s);

[0110] Step 3: Update the Critic network:

[0111] Minimize the mean square error (MSE) between the target Q value and the current Q value.

[0112] Step 4. Update the Actor network:

[0113] Maximize the target loss:

[0114] L = α·logπ(a|s)-Q 1 (s,π(s));

[0115] That is, high-value actions are selected while ensuring exploration.

[0116] Step 5: Soft update target network:

[0117] Soft update the target Q network parameters so that the target network parameters slowly approach the current network to avoid oscillation.

[0118] Therefore, the SAC algorithm combined with soft action selection can improve the robustness of the strategy.

[0119] In this example, in order to make the strategy more adaptable to uncertainty and variability, we use an action selection method based on probability distribution, namely soft action selection. This method allows the policy network to output a probability distribution of a series of actions instead of a single deterministic action. In this way, the strategy can weigh between multiple possible actions, thereby improving its robustness in the face of environmental changes.

[0120] The design of the policy network is as follows: it receives the current state s as input and outputs a probability distribution of actions. This distribution is determined by the parameters θ of the policy network and can be expressed by the following formula:

[0121] π(a|s)=softmax(π θ (s));

[0122] Among them, π(a|s) represents the probability distribution of executing action a in state s; π θ (s) represents the output of the policy network, which is a vector in which each element represents the probability value of the corresponding action; softmax is an activation function used to convert the vector output by the policy network into a probability distribution so that the sum of the probabilities of all possible actions is 1.

[0123] In this way, the policy network is able to not only recommend the best action, but also take into account the likelihood of other potential actions, thereby keeping the policy flexible and robust in the presence of environmental noise or high model uncertainty.

[0124] Temperature parameter α: controls the weight of entropy and balances exploration and exploitation.

[0125] π(a|s)=softmax(π θ (s) / α);

[0126] Therefore, when α is small, the strategy tends to choose the action with the largest probability value, that is, to use the known best action; when α is large, the strategy tends to choose the action with a more uniform probability value, that is, to explore new actions.

[0127] After the above model training, real-time data or test data is used for correlation calculation or testing, and the final output strategy is not a specific charge and discharge instruction, but an action probability distribution predicted based on the current state and future trends.

[0128] Assume that the current status is: renewable energy generation is high, load demand is low, and the energy storage system charge state is low.

[0129] According to the action probability distribution output by the SAC algorithm, the energy storage system may take the following actions:

[0130] Charging with a higher probability: Because the renewable energy generation is high, the charging cost is low, and the energy storage system has a low state of charge and needs to be replenished.

[0131] Discharging with lower probability: Discharging may result in wasted energy because the load demand is low.

[0132] Strategy adjustment: The SAC algorithm dynamically adjusts the action probability distribution based on real-time data and correlation analysis results. For example, when the power generation of renewable energy decreases, the probability of charging will decrease. When the load demand increases, the probability of discharging will increase. When the state of charge of the energy storage system is high, the probability of charging will decrease and the probability of discharging will increase.

[0133] The ultimate goal of the SAC algorithm is to find an optimal action strategy that enables the energy storage system to maximize economic benefits while smoothing fluctuations on the power generation side and meeting the dynamic demands on the power consumption side.

[0134] In order to improve the accuracy of the evaluation, this embodiment also includes the following solutions:

[0135] S4 continues to fuse multi-source data, and uses a time series feature extraction method to extract features to update the state space, action space and reward function, performs a correlation analysis again and then performs a strategy evaluation, and adjusts the strategy according to the strategy evaluation results.

[0136] In this embodiment, Figure 1 As shown, the classified historical load data is updated according to the time series to ensure that the historical load data is always data within the most recent period of time, and real-time data is collected at the same time to form a multi-source data set under a classification;

[0137] Extracting features from the multi-source data set using a time series feature extraction method to form a key feature set including multiple features;

[0138] Then, based on the multi-source data sets under all categories, the key feature sets under all categories are obtained.

[0139] Furthermore, this embodiment also includes:

[0140] After the correlation analysis is performed again, a strategy evaluation is performed, and the strategy is adjusted according to the strategy evaluation result, including:

[0141] According to the key feature set under each category, the state space, the action space and the corresponding reward function are updated, and the flexible action-evaluation algorithm is trained using the above data to obtain the optimal decision;

[0142] The optimal decision is evaluated using strategy evaluation indicators. If the set strategy effect is achieved, the correlation between the power generation and load power consumption under different time series is obtained according to the current decision. Otherwise, the current decision is adjusted and the strategy is re-evaluated.

[0143] The reward function corresponding to the update is expressed as: the current reward plus the adjustment item under the current category, and the adjustment item under the current category is obtained according to the actual changes of the control parameters under the current category at different times.

[0144] Specifically, an implementation of this embodiment includes:

[0145] 1) Data Collection

[0146] D={P gen,t ,P load,t ,E t ,H load ,W,P renew,t};

[0147] Where D is the collected multi-source dataset; P gen,t It is the real-time output data of the power generation side; load,t E is the load demand data on the electricity consumption side; t is the current state of charge of the battery energy storage system; H load is the historical load data; W is the weather forecast data; P renew,t For renewable energy power generation prediction, H load The historical load data is classified data at different times, that is, this application not only considers the classification of data horizontally, but also considers the relevance of data vertically from the time, so as to obtain a more accurate correlation analysis.

[0148] 2) Data preprocessing: clean the data, fill in missing values, normalize the data, and ensure data quality.

[0149] 3) Feature Engineering: During the training phase of the model, the historical data is first analyzed in depth to extract a set of key features. These features are extracted using time series feature extraction methods such as autocorrelation function and Fourier transform.

[0150] X={f 1 (D),f 2 (D),...,f n (D)};

[0151] Among them, X is the extracted key feature set; f i (D) is the i-th feature extracted from the dataset.

[0152] Since the extracted features are closely related to the state space, action space, and reward function defined above, they together constitute the input and output of the reinforcement learning model. Raw data usually contains a lot of redundant information, and feature extraction can help the model focus on important information and ignore irrelevant information. This can reduce the amount of data that the model needs to process.

[0153] The dynamic adjustment strategy of this embodiment adapts to the real-time operating environment:

[0154] 1) Real-time data collection:

[0155] D t = {P gen,t ,P load,t ,E t};

[0156] Among them, D t The real-time dataset is collected. This dataset contains the status information of the current environment, which is obtained through sensors, log files or other real-time data collection tools. The purpose of the real-time dataset is to capture the current status of the system in real time after the model is deployed, so that the model can make decisions based on the latest information.

[0157] In this embodiment, although the key feature set X and the real-time data set may differ in data source and collection time, they complement each other in the application of the model. The key feature set provides a training basis for the model, while the real-time data set provides input for the model during execution. That is, the key feature set can characterize the entire data, and the collection of real-time data is to correct the data, such as deviation.

[0158] 2) Take the extracted key features as part of the state space. Based on the previously defined state space, we now further refine the state space to adapt to the real-time operating environment. The initially defined state space provides us with a basic framework that contains all possible states that the model needs to consider during the training phase. However, in actual operation, it may be necessary to dynamically adjust the state space based on real-time data, which may include the following:

[0159] s t =(t,P gen,t ,P load,t ,E t ,H load ,W,r road ,η renew ,E gap ,r soc );

[0160] Among them, t is the current time, r road ,η renew ,Egap ,r soc They are the reward component that complies with the change rate, the reward component based on the green electricity penetration rate, the reward component related to the season type, and the reward component related to the data change rate.

[0161] In this embodiment, the state space initially defined may be to construct the basic framework of the reinforcement learning model. This state space defines all possible states that the model can observe and process, and is the starting point for the model to understand and interact with the environment.

[0162] In order to make the model more adaptable to the changes and uncertainties in the real-time operating environment, the state space is further refined or adjusted here. This redefined state space takes into account new states or more specific conditions that may be encountered in actual operation, such as extreme weather events, sudden fluctuations in market prices, or equipment failures. In this way, the dynamic adjustment strategy of the model can more accurately reflect the complexity and variability in actual operations.

[0163] In this embodiment, the reward function is updated as follows: In order to ensure that the reward function can effectively reflect the dynamic behavior of the system, several key features are incorporated into the design.

[0164] Here is a concrete example showing how to incorporate load change rate into the reward function:

[0165] Considering the impact of load change rate on system stability, a reward component related to load change rate is defined. Specifically, the immediate reward r in the reward function t It can be expressed as the current reward plus an adjustment term related to the load change rate as follows:

[0166] r t =r t +a*r load ;

[0167] Where: r t is the instantaneous reward at time step t; α is a weight coefficient used to adjust the impact of load change rate on the total reward; r load is the bonus component of the load change rate, which can be calculated based on the actual change of the load, for example, r load It can be the ratio of the load change to the load at the previous moment.

[0168] In this way, the reward function can dynamically adjust the reward value according to the actual operation of the system, thereby guiding the model to learn how to make optimal decisions under different load changes.

[0169] The strategy adjustment is expressed as a strategy adjustment amount calculated based on the current strategy plus a strategy evaluation result, and the strategy adjustment amount represents an adjustment amount of a control parameter of the battery energy storage system at the current time step.

[0170] In this embodiment, a strategy adjustment method is implemented as follows:

[0171] Define a concise evaluation metric D t , which directly reflects the effectiveness of the strategy. The evaluation indicators of this embodiment can evaluate the strategies under each classification and time series, and are specifically defined as follows:

[0172] D t is a policy evaluation metric that represents the performance of the policy at time step t. t The calculation of depends only on the output of the strategy and the actual response of the system and is independent of other parameters.

[0173] D t =g(strategy output, system response);

[0174] Here, g(·) is a function that calculates the evaluation value based on the decision of the strategy and the actual response of the system. This function should be designed to ensure that D t The quality of a policy can be effectively measured, for example, by comparing the actual load balance resulting from the policy with the expected load balance.

[0175] In this way, the effectiveness of the strategy can be evaluated directly and clearly without being influenced by other irrelevant parameters.

[0176] Strategy adjustment formula:

[0177] a t =a t +Δa t ;

[0178] Among them, Δa t is the strategy adjustment amount calculated based on the evaluation results of energy storage system operation performance and market demand. t Represents a control parameter of the energy storage system at time step t, such as discharge power or charging power.

[0179] Finally, in order to make the policy more adaptable to uncertainty and variability, an action selection method based on probability distribution, namely soft action selection, is adopted. This method allows the policy network to output a probability distribution of a series of actions instead of a single deterministic action. In this way, the policy can weigh between multiple possible actions, thereby improving its robustness in the face of environmental changes.

[0180] Therefore, the present invention proposes a method for generating energy storage charging and discharging strategies based on reinforcement learning SAC. The method first processes the collected data set, and then combines and improves the SAC algorithm on this basis, aiming to optimize the charging and discharging behavior of the energy storage system through an intelligent algorithm, smooth out fluctuations on the power generation side, meet the dynamic needs of the power consumption side, improve the stability of the power system, and maximize economic and environmental benefits.

[0181] Example 2

[0182] In a second aspect, the present invention further provides a system for generating energy storage charging and discharging strategies based on reinforcement learning, the system comprising:

[0183] A historical load data collection module is used to collect historical load data of a period of time in the distribution network, and classify the historical load data to obtain a feature data set containing multiple key information;

[0184] A parameter construction module, for constructing a state space and a reward function of a battery energy storage system according to a characteristic data set of the key information, wherein the battery energy storage system is used to store and release electric energy in response to changes in the power generation side and the power consumption side in the power system; and constructing an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system;

[0185] Algorithm training module, used to analyze the correlation between power generation and load power consumption under different categories using flexible action-evaluation algorithm;

[0186] The model updating module is used to continue fusing multi-source data, and to update the state space, action space and reward function after feature extraction using a time series feature extraction method, to perform strategy evaluation after performing correlation analysis again, and to adjust the strategy according to the strategy evaluation results.

[0187] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for generating energy storage charging and discharging strategies based on reinforcement learning.

[0188] The other technical features of the second and third aspects of the present application are similar to the corresponding methods for generating energy storage charging and discharging strategies based on reinforcement learning, and will not be elaborated here.

[0189] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0190] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0191] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for generating energy storage charging and discharging strategies based on reinforcement learning, characterized in that: The method includes: Collect historical load data of a period of time in the distribution network, and classify the historical load data to obtain a feature data set containing multiple key information; Constructing a state space and a reward function of a battery energy storage system according to the characteristic data set of the key information, wherein the battery energy storage system is used to store and release electric energy in response to changes in the power generation side and the power consumption side of the power system; constructing an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system; The flexible action-evaluation algorithm is used to analyze the correlation between the power generation and load power consumption under different categories; Continue to fuse multi-source data, and use the time series feature extraction method to update the state space, action space and reward function after feature extraction, perform strategy evaluation after correlation analysis again, and adjust the strategy according to the strategy evaluation results.

2. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 1, characterized in that: The historical load data includes the historical load power data of the substations within the distribution network of the study area, the wind power and irradiance data within the substations, and the distributed photovoltaic power generation power within the substations collected from multiple systems over a period of time, and the historical load data are classified, specifically: The historical load data is classified according to the special transformer substation and the public transformer substation; the daily load data of all users under the special transformer substation is aggregated and feature-recognized to form substation data classification processing of different types of loads, and different types of load data are obtained; the public transformer substation is classified into rural, township and urban areas to obtain load data of different regions; the historical load data is divided into summer and winter to obtain load data of different seasons; the sample data is classified according to working days and non-working days to obtain load data of different times; the ratio of the photovoltaic, wind power and hydropower installed capacity of the substation to the maximum load of the substation is calculated to obtain the green electricity penetration rate of the substation, and the historical load data is divided into three intervals of [0%-30%], (30%-80%) and (80%-100%) according to the penetration rate to obtain load data of different green penetration rates; the data change rate is calculated, and the change rate is divided into three intervals of [0%-30%], (30%-80%) and (80%-100%), and the load data corresponding to the change rate of the three intervals are extracted respectively.

3. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 2, characterized in that: The step of constructing a state space of the battery energy storage system according to the characteristic data set of the key information specifically includes: A state space set of a battery energy storage system is defined, wherein the state space set includes real-time data, classified load data and meteorological data, wherein the real-time data includes: current time, current output data of the power generation side of the distribution network, current user power demand of the power consumption side of the distribution network and current state of charge of the battery energy storage system; the classified load data is a type of classified load data; The action space is obtained by the difference between the current charging power of the battery energy storage system and the current discharging power of the battery energy storage system; The reward function of the battery energy storage system is constructed based on the characteristic data set of the key information, including the following factors: the power factor of purchasing or selling electricity from the power grid at the current moment, the energy change factor of the battery energy storage system corresponding to the next moment and the current moment, the ratio factor of the predicted power generation of renewable energy and the user's electricity demand at the current moment, and the charge factor of the current battery energy storage system.

4. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 3 is characterized in that: The flexible action-evaluation algorithm is used to analyze the correlation between the power generation and the load power under different categories, specifically including training the flexible action-evaluation algorithm: Initialize the Actor network and the Critic network; the Actor network is responsible for outputting the probability distribution of each action in the current state, while the Critic network is responsible for estimating the value of the state-action pair; for each time step, obtain the current state space from the environment; Use the Actor network to output the probability distribution of each action space under the current state space; sample an action space from the probability distribution and execute it; obtain a new state space and reward function from the environment; use the Critic network to estimate the action value and entropy of the state-action pair respectively; update the Actor network and the Critic network according to the action value and entropy; repeat the above until the algorithm converges or reaches the maximum number of iterations, so as to predict the best action probability and state value estimation under the current classification of the load data; Repeat the above steps to predict the best action probability and state value estimation under all load data categories, thereby obtaining the strategy under each load data category.

5. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 3 is characterized in that: The method of continuing to fuse multi-source data and extracting features using a time series feature extraction method to update the state space, action space and reward function includes: The classified historical load data is updated according to the time series to ensure that the historical load data is always the data of the most recent period, and real-time data is collected to form a multi-source data set under a classification; Extracting features from the multi-source data set using a time series feature extraction method to form a key feature set including multiple features; Then, based on the multi-source data sets under all categories, the key feature sets under all categories are obtained.

6. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 5, characterized in that: After the correlation analysis is performed again, a strategy evaluation is performed, and the strategy is adjusted according to the strategy evaluation result, including: According to the key feature set under each category, the state space, the action space and the corresponding reward function are updated, and the flexible action-evaluation algorithm is trained using the above data to obtain the optimal decision; The optimal decision is evaluated using strategy evaluation indicators. If the set strategy effect is achieved, the correlation between the power generation and load power consumption under different time series is obtained according to the current decision. Otherwise, the current decision is adjusted and the strategy is re-evaluated.

7. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 6, characterized in that: The reward function corresponding to the update is expressed as: the current reward plus the adjustment item under the current category, and the adjustment item under the current category is obtained according to the actual changes of the control parameters under the current category at different times.

8. The method for generating energy storage charging and discharging strategies based on reinforcement learning according to claim 6, characterized in that: The strategy adjustment is expressed as a strategy adjustment amount calculated based on the current strategy plus a strategy evaluation result, and the strategy adjustment amount represents an adjustment amount of a control parameter of the battery energy storage system at the current time step.

9. A system for generating energy storage charging and discharging strategies based on reinforcement learning, characterized in that: The system includes: A historical load data collection module is used to collect historical load data of a period of time in the distribution network, and classify the historical load data to obtain a feature data set containing multiple key information; A parameter construction module, for constructing a state space and a reward function of a battery energy storage system according to a characteristic data set of the key information, wherein the battery energy storage system is used to store and release electric energy in response to changes in the power generation side and the power consumption side in the power system; and constructing an action space of the battery energy storage system according to the current charging power and discharging power of the battery energy storage system; Algorithm training module, used to analyze the correlation between power generation and load power consumption under different categories using flexible action-evaluation algorithm; The model updating module is used to continue fusing multi-source data, and to update the state space, action space and reward function after feature extraction using a time series feature extraction method, to perform strategy evaluation after performing correlation analysis again, and to adjust the strategy according to the strategy evaluation results.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating energy storage charging and discharging strategies based on reinforcement learning as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Micro-grid energy management method and system based on reinforcement learning

    CN121689158A

  • A method and system for three-phase imbalance treatment of energy storage power distribution network based on DSAC algorithm

    CN122553264A