A flexible resource scheduling decision adaptive adjustment method based on reinforcement learning

By constructing a flexible resource scheduling decision-making method based on deep reinforcement learning, and dynamically adjusting the peak-valley coefficient and distributed renewable energy configuration, the problem of poor adaptability of traditional scheduling methods in the face of complexity and uncertainty is solved, thereby improving the stability and efficiency of the power system.

CN120258425BActive Publication Date: 2026-04-14STATE GRID SHANXI ELECTRIC POWER CO ECONOMIC & TECH RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID SHANXI ELECTRIC POWER CO ECONOMIC & TECH RES INST
Filing Date
2025-03-21
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional resource scheduling methods struggle to adapt to changes in system structure or load characteristics when faced with the complexity and uncertainty of integrated energy systems. This leads to decreased control effectiveness and a lack of robustness, making it difficult to respond quickly to load fluctuations and uncertainties in distributed energy sources, thus affecting the stability and accuracy of the power system.

Method used

We construct a flexible resource scheduling decision-making method based on deep reinforcement learning. By using a two-layer model for load fluctuation adjustment and a user distributed resource allocation model, we can dynamically adjust the peak-valley coefficient and the allocation of distributed renewable energy through Markov decision process and deep deterministic policy gradient reinforcement learning to optimize the resource scheduling strategy of the power system.

Benefits of technology

It enhances the flexibility and stability of the power system, enabling it to automatically adapt to system changes, reduce operation and maintenance complexity, effectively cope with load surges and uncertainties, and improve the automation level of resource scheduling and the operating efficiency of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258425B_ABST
    Figure CN120258425B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of flexibility resource scheduling decision adaptability adjustment method based on reinforcement learning, belong to integrated energy field, solve the problem of low flexibility of resource scheduling.It includes the load fluctuation adjustment double-layer model of resource scheduling based on the net load fluctuation level and carbon emission total amount of power system, and user distributed renewable energy and energy storage configuration, including load fluctuation adjustment optimization model and user distributed resource configuration optimization model;Obtain the net load data of typical day in certain historical period of power system, based on dynamic peak-valley period determination method, the net load power curve of each typical day is divided into peak-valley period, and the corresponding normal period, trough period and peak period are obtained;Based on normal period, trough period and peak period, the load fluctuation adjustment double-layer model is iteratively solved using deep reinforcement learning method, and the generator output result predicted after load fluctuation level optimization and the distributed renewable energy configuration and energy storage configuration capacity of each user are obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated energy technology, and in particular to a flexible resource scheduling decision-making adaptive adjustment method based on reinforcement learning. Background Technology

[0002] With the transformation of the energy structure and the rapid development of new energy sources, integrated energy systems are becoming increasingly complex, and load fluctuation problems have become particularly prominent. Traditional resource scheduling methods mainly rely on fixed mathematical models and parameter settings, such as PI control (Proportional-Integral Control). These methods often struggle to adapt when system structure or load characteristics change, leading to a decline in control effectiveness. For example, when new generating units are added to the power system or load characteristics change, the original control model may no longer be applicable, requiring manual intervention to adjust parameters, increasing the complexity and workload of operation and maintenance.

[0003] Furthermore, traditional resource scheduling methods lack sufficient robustness when facing disturbances such as sudden load changes and line faults, as well as uncertainties such as the randomness of distributed energy generation and changes in weather conditions. These factors may cause the power system frequency to deviate from the set value, making it difficult to respond quickly and accurately, thus making it difficult to maintain the stability and accuracy of the system.

[0004] In recent years, with the development of artificial intelligence technology, reinforcement learning has been increasingly applied to load management and optimization of power grids. Through interaction with the environment, reinforcement learning continuously learns and optimizes control strategies, effectively addressing dynamic changes and uncertainties in power systems. However, reinforcement learning models also face some challenges in application. For example, the model's state space and action space grow exponentially with the number of system nodes, leading to a sharp increase in exploration costs. Furthermore, directly using traditional reinforcement learning models for automatic adjustment of power grid operation modes may result in low exploration flexibility, low efficiency, and the generation of ineffective operating modes.

[0005] Traditional resource scheduling methods have many shortcomings when facing the complexity and uncertainty of integrated energy systems, while reinforcement learning has certain advantages, but it still needs further optimization and improvement in practical applications. Summary of the Invention

[0006] In view of the above analysis, the present invention aims to provide a reinforcement learning-based adaptive adjustment method for flexible resource scheduling decisions, in order to solve the problems of poor adaptability, complex operation and maintenance and low flexibility in existing methods in resource scheduling, and difficulty in effectively dealing with the uncertainties of distributed renewable resources, which leads to the net load power of the power system deviating from the preset value and failing to respond quickly and accurately, thus making it difficult to maintain the stability and accuracy of the system.

[0007] The objective of this invention is mainly achieved through the following technical solutions:

[0008] This invention provides a method for adaptive adjustment of flexible resource scheduling decisions based on reinforcement learning, comprising the following steps:

[0009] Based on the net load fluctuation level and total carbon emissions of the power system, as well as the configuration of distributed renewable energy and energy storage by users, a two-layer model for load fluctuation adjustment of resource scheduling is constructed, including a load fluctuation adjustment optimization model and a user distributed resource configuration optimization model.

[0010] The net load data of a typical day within a certain historical period of the power system is obtained. Based on the dynamic peak-valley time period determination method, the net load power curve of each typical day is divided into peak-valley time periods to obtain the corresponding normal period, low-valley period and peak period.

[0011] Based on the normal, off-peak, and peak periods, a deep reinforcement learning method is used to iteratively solve the two-layer load fluctuation adjustment model. The net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration of the power system are defined as state variables in the state space of a Markov decision process. The Markov agent is optimized using an actor-critic network based on deep deterministic policy gradient reinforcement learning. When the loss function of the critic network converges, the predicted generator output and the distributed renewable energy configuration capacity and energy storage configuration capacity of each user after the load fluctuation level optimization are obtained.

[0012] Furthermore, the load fluctuation adjustment optimization model includes: an objective function aimed at minimizing the system's net load fluctuation level and the system's total carbon emissions, as well as a set of constraints;

[0013] The objective function of the load fluctuation adjustment optimization model is as follows:

[0014]

[0015] Where f1 and f2 represent the total system carbon emissions and the system net load fluctuation level of the power system, respectively; f1' and f2' represent the normalized total system carbon emissions and the system net load fluctuation level, respectively; s represents a typical day; and θ represents the system net load fluctuation level.s Ω represents the probability of a typical day occurring. s For a typical day set, Ω T For the set of time periods, Ω G For a collection of thermal power units, e g Let g be the carbon emission intensity of the thermal power unit. Let L be the power output of thermal power unit g, Δt be the time interval of a period, and L be the power output of thermal power unit g. av This represents the average net load of the system.

[0016] Furthermore, the set of constraints for the load fluctuation adjustment optimization model includes: peak-valley coefficient boundary limit constraints, system power balance power constraints, and minimum limit constraints for peak-valley coefficients.

[0017] The peak-valley coefficient boundary limit constraints are as follows:

[0018]

[0019] in, Ω represents the peak-to-valley coefficient for time period t. v For the set of low periods, Ω m For the set of ordinary time intervals, Ω p For peak hours, This is the coefficient for the trough period. The coefficient difference between the trough and the normal period. This represents the difference in coefficients between normal and peak hours;

[0020] The electrical energy balance constraints are as follows:

[0021]

[0022] in, For user i, the electricity consumption under a typical day s. For user i's internet usage, For the power output of thermal power unit g, Ω D Ω represents the set of users in the system affected by the peak-valley coefficient. L This refers to the set of users in the system who are not affected by the peak-valley coefficient. The power output of thermal power units that are not affected by peak-valley coefficients;

[0023] The minimum limit constraint for the peak-valley coefficient is as follows:

[0024]

[0025] Among them, h R The recovery coefficient, L, is used to recover excess electrical energy from users. i,t,s For the initial load demand of user i, Let g be the unit electrical energy generation cost of thermal power unit g. The cost of transmitting electricity per unit of electrical energy in a power system.

[0026] Furthermore, the user distributed resource configuration optimization model includes: an objective function and a set of constraints aimed at minimizing the distributed renewable energy configuration capacity and energy storage configuration capacity;

[0027] The objective function of the user distributed resource configuration optimization model is as follows:

[0028]

[0029] Among them, w Re The annualized cost per unit capacity for configuring distributed renewable energy. w represents the renewable energy configuration capacity of user i. Es The annualized cost per unit capacity of energy storage. This represents the energy storage configuration capacity of user i.

[0030] Furthermore, the set of constraints for the user distributed resource configuration optimization model includes: power balance constraints, distributed renewable energy configuration and operation constraints, and energy storage configuration and operation constraints.

[0031] The power balance constraint is as follows:

[0032]

[0033] in, It is the actual power generation of distributed renewable energy configured by user i. and These are the energy storage charging and discharging power, respectively.

[0034] The configuration and operational constraints of the distributed renewable energy source are as follows:

[0035]

[0036] in, and α i It refers to the power generation efficiency and capacity factor of distributed renewable energy. Configure the maximum capacity for distributed renewable energy sources;

[0037] The energy storage configuration and operational constraints are as follows:

[0038]

[0039] in, and These are the upper limits of the charging and discharging power of energy storage, respectively. Let be the current energy storage capacity of user i at times t and t-1 on a typical day s, respectively. The charging / discharging efficiency of the energy storage for user i. and These represent the energy storage charging and discharging power at time t-1. This is the maximum capacity configured for energy storage.

[0040] Furthermore, the dynamic peak-valley time period determination method is used to divide each typical day into different time periods, including:

[0041] Input the net load data array x = {x1, x2, ... x3} for each typical day. 24};

[0042] Initialize the load time-sharing fluctuation matrix dp and the minimum fluctuation segment matrix divider, set the elements in the dp matrix to ∞, and set dp[0][0] to 0; set the elements in the divider matrix to 0;

[0043] Set the maximum number of segments, max_intervals, to 8;

[0044] The values ​​in the dp matrix and divider matrix are calculated using three nested loops; where...

[0045] The first loop iterates through the typical daily hour number i from 1 to 24;

[0046] The second loop iterates through the number of segments j from 1 to min(i, max_intervals);

[0047] The third loop variable k iterates from j-1 to i-1 to calculate the variance of the net load time series, as follows:

[0048]

[0049] Where n = ik is the net load data volume within this segment, x t Let t be the net load value for hour t. This represents the average net load data from the (k+1)th hour to the i-th time period;

[0050] Determine whether dp[k][j-1]+variance is less than dp[i][j]; where dp[i][j] represents the minimum sum of variances of the net load time series achieved by dividing the first i hours into j segments;

[0051] If so, update dp[i][j] to dp[k][j-1]+variance, and update divider[i][j] to k; where divider[i][j] represents the corresponding segmentation point;

[0052] Determine the segmentation points sequentially from dp

[24] [max_intervals] to obtain the time interval values; merge consecutive time interval values ​​with the same value.

[0053] Output the normal period, off-peak period, and peak period, along with their corresponding hourly values;

[0054] Among them, the power value of the net load data is less than the valley threshold, which is the valley period; it is greater than the peak threshold, which is the peak period; and the period between the valley threshold and the peak threshold is the normal period.

[0055] Furthermore, the iterative solution of the load fluctuation adjustment optimization model and the user distributed resource allocation optimization model based on the normal period, off-peak period, and peak period using the DRL learning method includes:

[0056] The peak-valley coefficient is obtained by inputting the normal period, off-peak period and peak period into the load fluctuation adjustment and optimization model.

[0057] Based on the peak-valley coefficient, the user distributed resource configuration optimization model obtains the corresponding distributed renewable energy configuration and energy storage configuration;

[0058] Based on the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration, a Markov decision process model is performed, including:

[0059] The state variables in the Markov state space include the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration;

[0060] Within each time step b, the Markov agent observes the state variable s. b ,as follows:

[0061] s b =(h b-1 N b-1 )

[0062] Among them, h b-1 N represents the peak-to-valley coefficient at step b-1; b-1 This represents the net load at step b-1;

[0063] The agent's action a b The settings are as follows:

[0064]

[0065] During the operation of the power system, the smaller the objective function value of the load fluctuation adjustment optimization model, the higher the reward of the Markov agent; the smaller the objective function value of the user distributed resource allocation optimization model, the lower the reward of the Markov agent.

[0066] The Markov agent is optimized using a deep deterministic policy gradient reinforcement learning algorithm with an actor-critic structure.

[0067] The critic network is based on the state variable s b The 48-dimensional vector and the action a b The output is the Q-value estimate of the current state-action pair, which scores the action. The actor network optimizes its parameters based on the Q-value estimate output by the critic network until the loss function of the critic network converges and the iteration stops.

[0068] Furthermore, the reward function r of the Markov agent b The settings are as follows:

[0069]

[0070] in, The sum of the user's distributed renewable energy configuration and energy storage configuration.

[0071] Furthermore, the loss function of the critic network is as follows:

[0072] min MSE[Q θμ (s b ,a b ),r b +γQ θQ′ (s′ b ,a′ b )]

[0073] Where MSE represents the mean square error, θ μ For the critic network parameters, (s b ,a b ) and (s′ b ,a′ b ) represent the state-action pair of the previous time step and the state-action pair of the next time step, respectively; r b Let θ be the reward for the current step; γ be the discount factor for the reward; optimize the network parameters θ by minimizing the loss function. μ This enhances the ability to predict future states.

[0074] Furthermore, the action value function Q(s) of the actor networkb ,a b The parameters of the actor network are optimized according to the following formula:

[0075]

[0076] Where, θ μ The parameters θ are the values ​​of the actor network, which optimizes the parameters θ based on the Q-value estimate from the critic network output. μ Perform a gradient ascent operation to maximize the action value function Q.

[0077] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0078] 1. The flexible resource scheduling decision-making adaptive adjustment method based on reinforcement learning in this invention can enhance the exploratory nature of resource scheduling and avoid premature convergence, thereby identifying better or near-optimal load fluctuation levels. By minimizing net load fluctuation levels and total system carbon emissions, it effectively smooths load fluctuations, reduces system operation risks, and improves the stability of the power system. Dynamic peak-valley time period division ensures the operating efficiency of the power system, reduces peak-valley differences, reduces the peak-valley differences in system electrical energy, and further improves the flexibility, stability, and reliability of resource scheduling.

[0079] 2. This invention utilizes Markov Decision Process (MDP) modeling and Deep Reinforcement Learning (DRL) optimization to dynamically adjust resource scheduling strategies during load fluctuations, adapting to changes in power system structure and load characteristics without requiring manual intervention for parameter adjustments. The agent learns optimal strategies through interaction with the environment, automatically adapting to new generator units or changes in load characteristics, reducing operational complexity and workload, and enhancing the adaptability of power system resource scheduling.

[0080] 3. This method can effectively cope with uncertainties such as sudden changes in power system load, line faults, randomness of distributed renewable energy generation, and changes in meteorological conditions, maintaining the stability and accuracy of the power system; by adjusting load fluctuation adjustment model parameters such as peak-valley coefficient, it encourages users to configure distributed renewable energy and energy storage devices, thereby achieving sustainable development of the power system and effective control of carbon emissions.

[0081] 4. This method utilizes deep reinforcement learning algorithms to enable the agent to autonomously learn optimal strategies, achieve intelligent decision-making, and improve the automation level of power system resource scheduling. Based on dynamic learning using historical and real-time data, the strategy is continuously optimized to adapt to the dynamic changes of the power system, enhancing the system's flexibility, resource scheduling decision-making, and operational efficiency; it also achieves dynamic adaptive adjustment during peak and valley periods, thereby reducing the peak-valley difference in system power energy, improving the safety and reliability of system operation, and promoting the consumption of distributed renewable energy.

[0082] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description

[0083] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0084] Figure 1 This is a flowchart of a flexible resource scheduling decision adaptive adjustment method based on reinforcement learning in an embodiment of the present invention;

[0085] Figure 2 This is a schematic diagram of a two-layer model for load fluctuation adjustment in resource scheduling during flexible resource scheduling decisions, as described in this embodiment of the invention. Detailed Implementation

[0086] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0087] A specific embodiment of the present invention, such as Figure 1 As shown, a flexible resource scheduling decision-making adaptive adjustment method based on reinforcement learning is disclosed, including the following steps:

[0088] Step S1: Based on the net load fluctuation level and total carbon emissions of the power system, as well as the configuration of distributed renewable energy and energy storage for users, construct a two-layer model for load fluctuation adjustment of resource scheduling, including a load fluctuation adjustment optimization model and a user distributed resource configuration optimization model.

[0089] Step S2: Obtain the net load data of typical days within a certain historical period of the power system, and divide the net load power curve of each typical day into peak and valley periods based on the dynamic peak and valley period determination method to obtain the corresponding normal period, valley period and peak period.

[0090] Step S3: Based on the normal period, off-peak period, and peak period, the two-layer model for load fluctuation adjustment of resource scheduling is iteratively solved using a deep reinforcement learning method. The net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration of the power system are defined as state variables in the state space of the Markov decision process. The Markov agent is optimized using an actor-critic network of deep deterministic policy gradient reinforcement learning. When the loss function of the critic network converges, the predicted generator output and the distributed renewable energy configuration capacity and energy storage configuration capacity of each user after the load fluctuation level optimization are obtained.

[0091] Step S1 includes steps S11-S12.

[0092] Step S11: Construct a load fluctuation optimization model.

[0093] like Figure 2 As shown. The load fluctuation adjustment optimization model includes: an objective function aimed at minimizing the system's net load fluctuation level and total system carbon emissions, and a set of constraints;

[0094] The objective function of the load fluctuation adjustment optimization model is as follows:

[0095]

[0096] Where f1 and f2 represent the total system carbon emissions and the system net load fluctuation level of the power system, respectively; f1' and f2' represent the normalized total system carbon emissions and the system net load fluctuation level, respectively; s represents a typical day; and θ represents the system net load fluctuation level. s Ω represents the probability of a typical day occurring. s For a typical day set, Ω T For the set of time periods, Ω G For a collection of thermal power units, e g Let g be the carbon emission intensity of the thermal power unit. Let L be the power output of thermal power unit g, Δt be the time interval of a period, and L be the power output of thermal power unit g. av This represents the average net load of the system.

[0097] Since f1 and f2 have different dimensions, they were normalized and then added together.

[0098] The set of constraints for the load fluctuation adjustment optimization model includes: peak-valley coefficient boundary limit constraints, system power balance power constraints, and minimum limit constraints for peak-valley coefficients.

[0099] The peak-valley coefficient boundary limit constraints are as follows:

[0100]

[0101] in, Ω represents the peak-to-valley coefficient for time period t. v For the set of low periods, Ω m For the set of ordinary time intervals, Ω p For peak hours, This is the coefficient for the trough period. The coefficient difference between the trough and the normal period. This represents the difference in coefficients between normal and peak hours;

[0102] The electrical energy balance constraints are as follows:

[0103]

[0104] in, For user i, the electricity consumption under a typical day s. For user i's internet usage, For the power output of thermal power unit g, Ω D Ω represents the set of users in the system affected by the peak-valley coefficient. L This refers to the set of users in the system who are not affected by the peak-valley coefficient. The power output of thermal power units that are not affected by peak-valley coefficients;

[0105] The minimum limit constraint for the peak-valley coefficient is as follows:

[0106]

[0107] Among them, h R The recovery coefficient, L, is used to recover excess electrical energy from users. i,t,s For the initial load demand of user i, Let g be the unit electrical energy generation cost of thermal power unit g. The cost of transmitting electricity per unit of electrical energy in a power system.

[0108] The purpose of step S11 is to construct a load fluctuation adjustment optimization model for resource scheduling. This model aims to minimize the net load fluctuation level of the power system and the total carbon emissions of the system. By defining the objective function and the set of constraints, it provides a mathematical optimization framework for the resource scheduling of the power system and ensures the stability of the power system operation.

[0109] Step S12: Construct a user distributed resource configuration optimization model.

[0110] The load fluctuation optimization model in step S11 can obtain the optimized peak-valley coefficient. Users adjust their electricity consumption behavior based on the optimized peak-valley coefficient. When the peak-valley coefficient fluctuates greatly, users can utilize the capacity of distributed renewable energy configuration and energy storage configuration to improve the absorption of renewable energy.

[0111] The user distributed resource configuration optimization model includes: an objective function and a set of constraints aimed at minimizing the distributed renewable energy configuration capacity and energy storage configuration capacity;

[0112] The objective function of the user distributed resource configuration optimization model is as follows:

[0113]

[0114] Among them, w Re The annualized cost per unit capacity for configuring distributed renewable energy. w represents the renewable energy configuration capacity of user i. Es The annualized cost per unit capacity of energy storage. This represents the energy storage configuration capacity of user i.

[0115] The configuration costs of distributed renewable energy (such as solar and wind power) and energy storage devices are relatively high. By minimizing these capacities, investment and operation and maintenance costs can be reduced while meeting user electricity demand. For example, in distributed energy storage systems, rationally configuring the distributed renewable energy configuration and energy storage capacity can avoid over-investment and improve the system's economics.

[0116] The output of distributed renewable energy is intermittent and uncertain, while energy storage systems are used to smooth power fluctuations and provide backup capacity. By optimizing the configuration, distributed energy sources and energy storage systems can be better coordinated, enhancing the flexibility and stability of the system; minimizing the energy storage configuration capacity can improve the peak-shaving and frequency regulation capabilities of the power grid and optimize grid operation.

[0117] The set of constraints for the user distributed resource configuration optimization model includes: power balance constraints, distributed renewable energy configuration and operation constraints, and energy storage configuration and operation constraints.

[0118] The power balance constraint is as follows:

[0119]

[0120] in, It is the actual power generation of distributed renewable energy configured by user i. and These are the energy storage charging and discharging power, respectively.

[0121] The configuration and operational constraints of the distributed renewable energy source are as follows:

[0122]

[0123] in, α i These are the power generation efficiency and capacity factor of renewable energy, respectively. Configure the maximum capacity for renewable energy;

[0124] The energy storage configuration and operational constraints are as follows:

[0125]

[0126] in, and These are the upper limits of the charging and discharging power of energy storage, respectively. Let be the current energy storage capacity of user i at times t and t-1 on a typical day s, respectively. The charging / discharging efficiency of the energy storage for user i. and These represent the energy storage charging and discharging power at time t-1. This is the maximum capacity configured for energy storage.

[0127] Step S12 aims to construct a user distributed resource configuration optimization model. With the goal of minimizing the configuration capacity of distributed renewable energy and the configuration capacity of energy storage, it optimizes the configuration and operation of the user's distributed renewable energy and energy storage equipment by defining power balance constraints, distributed renewable energy configuration operation constraints, and energy storage configuration and operation constraints, so as to achieve flexible resource scheduling, efficient resource utilization, and maximize the absorption of distributed renewable energy.

[0128] The purpose of step S1 is to construct a two-layer model for load fluctuation adjustment in resource scheduling, including a load fluctuation adjustment optimization model and a user distributed resource allocation optimization model. By defining the objective function and the set of constraints, a mathematical optimization framework is provided for load fluctuation adjustment of the power system and the optimal allocation of user-side distributed renewable energy and energy storage, so as to achieve the stability of power system operation and the efficient utilization of distributed renewable energy.

[0129] Step S2 includes:

[0130] The algorithm uses dynamic programming to divide the peak, flat, and valley periods of each typical day. The input to the algorithm is the 24-hour net load data x = {x1, x2, ... x} for each typical day within a certain historical period. 24 The algorithm outputs the hourly values ​​corresponding to normal, off-peak, and peak periods, respectively.

[0131] A typical day refers to a representative day selected from historical data, whose load curve and operating characteristics can reflect the typical operating state of the power system under specific seasons or conditions. The selection of a typical day is usually based on factors such as load patterns, weather conditions, and user behavior, with the aim of simplifying model calculations while ensuring that the model can effectively reflect actual operating conditions.

[0132] For example, both large and small fluctuations in net load can be considered typical days.

[0133] A specific historical period refers to a set of data collected over a past period used for analysis and modeling. For example, a specific historical period could be a few days, several weeks, several months, or a year, depending on the research objectives and data availability. By analyzing data within a historical period, patterns in load fluctuations, the output characteristics of distributed energy resources, and the operating modes of energy storage systems can be extracted, thus providing a basis for optimizing a two-tier model of load fluctuation adjustment in resource scheduling.

[0134] The dynamic peak-valley time period determination method is used to divide each typical day into different time periods, including:

[0135] Input the net load data array x = {x1, x2, ... x3} for each typical day. 24};

[0136] Initialize the load time-sharing fluctuation matrix dp and the minimum fluctuation segment matrix divider, set the elements in the dp matrix to ∞, and set dp[0][0] to 0; set the elements in the divider matrix to 0;

[0137] Set the maximum number of segments, max_intervals, to 8;

[0138] The values ​​in the dp matrix and divider matrix are calculated using three nested loops;

[0139] The first loop iterates through the typical daily hour number i from 1 to 24;

[0140] The second loop iterates through the number of segments j from 1 to min(i, max_intervals);

[0141] The third loop variable k iterates from j-1 to i-1 to calculate the variance of the net load time series, as follows:

[0142]

[0143] Where n = ik is the net load data volume within this segment, x t Let t be the net load value for hour t. This represents the average net load data from the (k+1)th hour to the i-th time period;

[0144] Determine whether dp[k][j-1]+variance is less than dp[i][j]; where dp[i][j] represents the minimum sum of variances of the net load time series achieved by dividing the first i hours into j segments;

[0145] If so, update dp[i][j] to dp[k][j-1]+variance, and update divider[i][j] to k; where divider[i][j] represents the corresponding segmentation point;

[0146] Determine the segmentation points sequentially from dp

[24] [max_intervals] to obtain the time interval values; merge consecutive time interval values ​​with the same value.

[0147] Output the normal period, off-peak period, and peak period, along with their corresponding hourly values;

[0148] Among them, the power value of the net load data is less than the valley threshold, which is the valley period; it is greater than the peak threshold, which is the peak period; and the period between the valley threshold and the peak threshold is the normal period.

[0149] For example, the valley threshold is set to 25%, the peak threshold is set to 75%, and max_intervals is set to 8.

[0150] The peak, flat, and valley periods are divided based on a dynamic programming algorithm. The pseudocode for the specific algorithm is shown below:

[0151]

[0152] In the pseudocode above, `max_intervals` represents the maximum number of consecutive time intervals. The algorithm aims to minimize the total load variation across all time intervals and ensure consistent load within each time interval, thereby determining the segmentation points for peak, flat, and valley periods. Afterward, we iterate through the segmented time intervals and merge consecutive periods of the same type. Finally, based on the average load within each period, each period is labeled as peak, mid-peak, or valley. The valley and peak thresholds are set to the 1 / 4 and 3 / 4 quantiles of the net load power curve, respectively, with the intervals in between considered normal.

[0153] In the process of dividing peak, flat, and valley periods using a dynamic programming algorithm, the variance calculated in each iteration refers to the variance of the net load time series within the current segment. Specifically, the variance is calculated as the average of the squares of the differences between the net load value within the current segment and the average net load value within that segment. This variance measures the degree of fluctuation in the net load within the current segment; the smaller the variance, the smaller the fluctuation level of the net load within that segment, and the more stable it is.

[0154] Step S2 is to use a dynamic programming algorithm to divide the net load power curve of the power system into off-peak, peak and flat periods. By minimizing the variance of the net load time series in each period, the hourly values ​​of the flat, off-peak and peak periods are dynamically determined, providing a basis for the flexible load fluctuation adjustment of subsequent resource scheduling.

[0155] Step S3, specifically.

[0156] The step of iteratively solving the load fluctuation adjustment optimization model and the user distributed resource allocation optimization model using the DRL learning method based on the normal period, off-peak period, and peak period includes:

[0157] The peak-valley coefficient is obtained by inputting the normal period, off-peak period and peak period into the load fluctuation adjustment and optimization model.

[0158] Based on the peak-valley coefficient, the user distributed resource configuration optimization model obtains the corresponding distributed renewable energy configuration and energy storage configuration;

[0159] Based on the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration, a Markov decision process model is performed, including:

[0160] The state variables in the Markov state space include the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration;

[0161] Within each time step b, the Markov agent observes the state variable s. b ,as follows:

[0162] s b =(h b-1 N b-1 ) Formula (14)

[0163] Among them, h b-1 N represents the peak-to-valley coefficient at step b-1; b-1 This represents the net load at step b-1;

[0164] The agent's action a b The settings are as follows:

[0165]

[0166] During the operation of the power system, the smaller the objective function value of the load fluctuation adjustment optimization model, the higher the reward of the Markov agent; the smaller the objective function value of the user distributed resource allocation optimization model, the lower the reward of the Markov agent.

[0167] The Markov agent is optimized using a deep deterministic policy gradient reinforcement learning algorithm with an actor-critic structure.

[0168] The critic network is based on the state variable s b The 48-dimensional vector and the action a b The output is the Q-value estimate of the current state-action pair, which scores the action. The actor network optimizes its parameters based on the Q-value estimate output by the critic network until the loss function of the critic network converges and the iteration stops.

[0169] The process of adjusting and optimizing the parameters of the design load fluctuation model is modeled as a Markov decision process. By constructing an agent to set different peak-valley coefficients and gradually interacting with the simulated power system application environment, the optimal peak-valley coefficient result is finally obtained.

[0170] In Markov decision-making, each time step is a scheduling cycle, labeled with index b. At time step b, the agent observes its current state and selects an action. If the action stabilizes within a round, the round ends. To maintain training efficiency, if the action does not stabilize within 500 steps, the next training round begins.

[0171] The agent's actions are encapsulated as a three-dimensional continuous vector.

[0172] The reward function for the reinforcement learning agent is set. The setting of the reward function mainly considers two influencing factors: power system operation (including net load fluctuation level and total carbon emissions) and user capacity configuration (including distributed renewable energy configuration and energy storage configuration). The lower the total carbon emissions of the system operation, the higher the reward received by the agent; conversely, the higher the user capacity configuration, the lower the reward.

[0173] The reward function r of a Markov agent b The settings are as follows:

[0174]

[0175] in, The sum of the user's distributed renewable energy configuration and energy storage configuration.

[0176] The agent's policy is optimized using the DDPG (Deep Deterministic Policy Gradient) reinforcement learning algorithm. Considering that the agent's actions reside in a continuous decision space, this algorithm employs an actor-critic structure. The actor network directly outputs the actions, while the critic network scores the actions based on rewards from environmental feedback.

[0177] The goal of an actor network is to maximize the expected cumulative reward. The action-value function of an actor network is Q(s). b ,a b The parameters of the actor network are optimized according to the following formula:

[0178]

[0179] Where, θ μ The parameters θ are the values ​​of the actor network, which optimizes the parameters θ based on the Q-value estimate from the critic network output. μ Perform a gradient ascent operation to maximize the action value function Q.

[0180] Considering that in the process of determining peak and trough coefficients, each decision in the actor network influences subsequent decisions, forming a temporal dependency chain, a Long Short-Term Memory (LSTM) network is used to capture this dependency. Each LSTM network receives a batch of samples as input, and each sample contains a state variable s. b Considering that the actor network needs to determine three different actions, the architecture adopts a stacked structure consisting of three independent long short-term memory layers, each corresponding to the agent's action 'a'. b One dimension. The output of each Long Short-Term Memory network represents the result of a specific action, which is then processed by a sigmoid function to ensure that it falls within the (0,1) interval, and then scaled proportionally to obtain the peak-to-valley coefficient.

[0181] Each long short-term network outputs a value representing the result of a specific action, with the output value ranging from (0,1). To convert these output values ​​into actual peak-valley coefficients, they are scaled proportionally according to the application scenario.

[0182] For example, the target range of the peak-valley coefficient is determined based on the actual application scenario and requirements. The coefficient during the trough period may be in the range of [0.5, 1.0].

[0183] For each LSTM output value x, scale using the following formula:

[0184] Peak-valley coefficient = x × (maximum value - minimum value) + minimum value (Formula 18)

[0185] Assuming the target range is [0.5, 1.5], then the peak-valley coefficient = x × (1.5 - 0.5) + 0.5 = x + 0.5.

[0186] Since the actor network needs to determine three different actions, each LSTM output corresponds to one action dimension, and each output value is converted into the actual peak-valley coefficient through the scaling formula mentioned above.

[0187] The output of the LSTM network is converted into actual peak-valley coefficients to optimize load fluctuation adjustment strategies for the power system. This approach not only considers time dependence but also ensures the practical availability of the output values ​​through scaling.

[0188] The learning process of the critic network is influenced by the currently estimated Q-value Q(s). b ,a b ) and the target Q value Q(s) b ′,a b The impact of ′) and the loss function of the critic network are as follows:

[0189]

[0190] Where MSE represents the mean square error, θ μ For the critic network parameters, (s b ,a b ) and (s b ′,a b ′) represent the state-action pair of the previous time step and the state-action pair of the next time step, respectively; r b Let θ be the reward for the current step; γ be the discount factor for the reward; optimize the network parameters θ by minimizing the loss function. μ This enhances the ability to predict future states.

[0191] Step S3 utilizes deep reinforcement learning to iteratively solve the load fluctuation adjustment optimization model and the user distributed resource allocation optimization model based on dynamically divided peak, flat, and valley periods. By modeling with Markov decision processes and optimizing the agent's strategy using a deep deterministic policy gradient algorithm, the peak-valley coefficient and distributed resource allocation are dynamically adjusted to minimize net load fluctuation levels and total carbon emissions. Simultaneously, it enhances the absorption of distributed renewable energy by users and utilizes energy storage capacity. Ultimately, it obtains predicted generator output and the distributed renewable energy and energy storage capacity of each user, addressing the issues of low system stability and accuracy when net load power deviates from preset values.

[0192] In summary, the flexible resource scheduling decision adaptive adjustment method based on reinforcement learning according to the embodiments of the present invention has the following beneficial effects:

[0193] 1. The flexible resource scheduling decision-making adaptive adjustment method based on reinforcement learning in this invention can enhance the exploratory nature of resource scheduling and avoid premature convergence, thereby identifying better or near-optimal load fluctuation levels. By minimizing net load fluctuation levels and total system carbon emissions, it effectively smooths load fluctuations, reduces system operation risks, and improves the stability of the power system. Dynamic peak-valley time period division ensures the operating efficiency of the power system, reduces peak-valley differences, reduces the peak-valley differences in system electrical energy, and further improves the flexibility, stability, and reliability of resource scheduling.

[0194] 2. This invention utilizes Markov decision process modeling and deep reinforcement learning optimization to dynamically adjust resource scheduling strategies during load fluctuations, adapting to changes in power system structure and load characteristics without requiring manual intervention for parameter adjustments. The intelligent agent learns the optimal strategy through interaction with the environment, automatically adapting to new generator units or changes in load characteristics in the power system, reducing operational complexity and workload, and enhancing the adaptability of power system resource scheduling.

[0195] 3. This method can effectively cope with uncertainties such as sudden changes in power system load, line faults, randomness of distributed renewable energy generation, and changes in meteorological conditions, maintaining the stability and accuracy of the power system; by adjusting load fluctuation adjustment model parameters such as peak-valley coefficient, it encourages users to configure distributed renewable energy and energy storage devices, thereby achieving sustainable development of the power system and effective control of carbon emissions.

[0196] 4. This method utilizes deep reinforcement learning algorithms to enable the agent to autonomously learn optimal strategies, achieve intelligent decision-making, and improve the automation level of power system resource scheduling. Based on dynamic learning using historical and real-time data, the strategy is continuously optimized to adapt to the dynamic changes of the power system, enhancing the system's flexibility, resource scheduling decision-making, and operational efficiency; it also achieves dynamic adaptive adjustment during peak and valley periods, thereby reducing the peak-valley difference in system power energy, improving the safety and reliability of system operation, and promoting the consumption of distributed renewable energy.

[0197] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for adjusting the adaptability of a flexible resource scheduling decision based on reinforcement learning, characterized in that, Includes the following steps: Based on the net load fluctuation level and total carbon emissions of the power system, as well as the configuration of distributed renewable energy and energy storage by users, a two-layer model for load fluctuation adjustment of resource scheduling is constructed, including a load fluctuation adjustment optimization model and a user distributed resource configuration optimization model. The load fluctuation adjustment optimization model includes: an objective function aimed at minimizing the system's net load fluctuation level and the system's total carbon emissions, as well as a set of constraints; The objective function of the load fluctuation adjustment optimization model is as follows: wherein, respectively are the total system carbon emission and the system net load fluctuation level of the power system, respectively are the normalized total system carbon emission and the system net load fluctuation level, is a typical day, is the probability of the typical day occurring, is a set of typical days, is a set of time intervals, is a set of thermal power units, is the carbon emission intensity of the thermal power unit g, is the power output of the thermal power unit , is the time interval of a time interval, is the system net load average value; The set of constraints for the load fluctuation adjustment optimization model includes: peak-valley coefficient boundary limit constraints, system power balance power constraints, and minimum limit constraints for peak-valley coefficients. The peak-valley coefficient boundary limit constraints are as follows: in, for t Peak-valley coefficient for a given time period This is a collection of low-period periods. For the usual time period set, For peak hours, This is the coefficient for the trough period. The coefficient difference between the trough and the normal period. This represents the difference in coefficients between normal and peak hours; The system power balance power constraints are as follows: in, For users On a typical day Electricity consumption below For users Internet usage power consumption For thermal power units Power output, The set of users in the system affected by the peak-valley coefficient. This refers to the set of users in the system who are not affected by the peak-valley coefficient. The power output of thermal power units that are not affected by peak-valley coefficients; The minimum limit constraint for the peak-valley coefficient is as follows: in, The recovery coefficient is used to recover excess electrical energy from users. For users The initial load demand, Let g be the unit electrical energy generation cost of thermal power unit g. The cost of transmitting electricity per unit of electrical energy in a power system; The user distributed resource configuration optimization model includes: an objective function and a set of constraints aimed at minimizing the distributed renewable energy configuration capacity and energy storage configuration capacity; The objective function of the user distributed resource configuration optimization model is as follows: in, The annualized cost per unit capacity for configuring distributed renewable energy. Indicates user Renewable energy configuration capacity, The annualized cost per unit capacity of energy storage. Indicates user The energy storage configuration capacity; The set of constraints for the user distributed resource configuration optimization model includes: power balance constraints, distributed renewable energy configuration and operation constraints, and energy storage configuration and operation constraints. The power balance constraint is as follows: in, User Configure the actual power generation of distributed renewable energy sources. and These are the energy storage charging and discharging power, respectively. The configuration and operational constraints of the distributed renewable energy source are as follows: in, and It refers to the power generation efficiency and capacity factor of distributed renewable energy. Configure the maximum capacity for distributed renewable energy sources; The energy storage configuration and operational constraints are as follows: in, and These are the upper limits of the charging and discharging power of energy storage, respectively. , users respectively The current energy storage capacity at times t and t-1 on a typical day s. For users The energy storage charging / discharging efficiency, and These represent the energy storage charging and discharging power at time t-1. Configure the maximum capacity for energy storage; The net load data of a typical day within a certain historical period of the power system is obtained. Based on the dynamic peak-valley time period determination method, the net load power curve of each typical day is divided into peak-valley time periods to obtain the corresponding normal period, low-valley period and peak period. Based on the normal, off-peak, and peak periods, a two-layer model for load fluctuation adjustment in resource scheduling is iteratively solved using deep reinforcement learning. The net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration of the power system are defined as state variables in the state space of a Markov decision process. The Markov agent is optimized using an actor-critic network based on deep deterministic policy gradient reinforcement learning. When the loss function of the critic network converges, the predicted generator output and the distributed renewable energy configuration capacity and energy storage configuration capacity of each user after the load fluctuation level optimization are obtained. The step of iteratively solving the load fluctuation adjustment optimization model and the user distributed resource allocation optimization model using the DRL learning method based on the normal period, off-peak period, and peak period includes: The peak-valley coefficient is obtained by inputting the normal period, off-peak period and peak period into the load fluctuation adjustment and optimization model. Based on the peak-valley coefficient, the user distributed resource configuration optimization model obtains the corresponding distributed renewable energy configuration and energy storage configuration; Based on the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration, a Markov decision process model is performed, including: The state variables in the Markov state space include the current net load fluctuation level, peak-valley coefficient, generator output, distributed renewable energy configuration, and energy storage configuration; At each time step Inside, the state variables observed by the Markov agent ,as follows: in, Indicates the first Peak-valence coefficient of the step; Indicates the first Net load of the step; Actions of the intelligent agent The settings are as follows: During the operation of the power system, the smaller the objective function value of the load fluctuation adjustment optimization model, the higher the reward of the Markov agent; the smaller the objective function value of the user distributed resource allocation optimization model, the lower the reward of the Markov agent. The Markov agent is optimized using a deep deterministic policy gradient reinforcement learning algorithm with an actor-critic structure. The critic network is based on the state variables. The 48-dimensional vector and the action The output is a Q-value estimate of the current state-action pair, which scores the action. The actor network optimizes its parameters based on the Q-value estimate output by the critic network until the loss function of the critic network converges and the iteration stops.

2. The method according to claim 1, characterized in that, The dynamic peak-valley time period determination method is used to divide each typical day into different time periods, including: Input an array of net load data for each typical day. ; Initialize the load time-sharing fluctuation matrix dp and the minimum fluctuation segment matrix divider, set the elements in the dp matrix to ∞, and set dp[0][0] to 0; set the elements in the divider matrix to 0; Set the maximum number of segments, max_intervals, to 8; The values ​​in the dp matrix and divider matrix are calculated using three nested loops; where... The first loop iterates through the typical daily hour number h from 1 to 24; The second loop iterates through the number of segments j from 1 to min(h, max_intervals); The third loop variable k is from arrive Perform a loop to calculate the variance of the net load time series. variance ,as follows: in, This represents the net load data volume within this segment. For the first Net load value per hour, For from the first Hours The average value of net load data within the time period; judge Is it less than dp[h][j]? Where dp[h][j] represents the sum of the minimum variances of the net load time series achieved by dividing the first h hours into j segments. If so, update dp[h][j] to And update divider[h][j] to k; where divider[h][j] represents the corresponding segmentation point; Determine the segmentation points sequentially from dp[24][max_intervals] to obtain the time interval values; merge consecutive time interval values ​​with the same value. Output the normal period, off-peak period, and peak period, along with their corresponding hourly values; Among them, the power value of the net load data is less than the valley threshold, which is the valley period; it is greater than the peak threshold, which is the peak period; and the period between the valley threshold and the peak threshold is the normal period.

3. The method according to claim 1, characterized in that, reward function of Markov agent The settings are as follows: in, The sum of the user's distributed renewable energy configuration and energy storage configuration.

4. The method according to claim 1, characterized in that, The loss function of the critic network is as follows: Where MSE represents the mean squared error. These are the parameters for the critic network. and These are the state-action pair for the previous time step and the state-action pair for the next time step, respectively. The reward for the current step; The discount factor is used to determine the reward; the network parameters are optimized by minimizing the loss function. This enhances the ability to predict future states.

5. The method according to claim 1, characterized in that, Action value function of actor network The parameters of the actor network are optimized according to the following formula: in, The parameters of the actor network are optimized based on the Q-value estimate from the output of the critic network. Perform gradient ascent operation to make the action value function maximize.

Citation Information

Patent Citations

  • Dynamic power system economic dispatching method based on deep reinforcement learning

    CN112186743A

  • Distributed energy storage optimal peak-valley period division method considering net load demand

    CN115205068A